Skip to content

← All posts

Technical SEO5-min read

The checks that decide whether a page can be indexed

Before content quality enters the picture, a page has to clear a short list of technical gates. Fail any one of them and Google will not index the page, however good it is. These are the checks our diagnosis runs on every URL, in order, and what a healthy result looks like for each. The free checker runs the same pass on a single URL.

1. HTTP status

The URL you want indexed should return 200. A 301 or 302 is fine only as a single clean hop to the real page, and in that case you should point your sitemap and internal links at the destination, not the redirecting URL. A 404 or 410 is correct for a page that is gone, but it is not indexable. A 5xx during Google’s crawl gets the page dropped. Redirect chains of three or more hops, and redirect loops, are their own failure: the URL never resolves to something Google can index.

2. robots.txt

A Disallow rule that matches the path stops the crawl before the page is ever fetched. Google can still list a blocked URL with no description if enough sites link to it, but it cannot read the content or act on anything inside it. Check the longest-matching rule for both Googlebot and *, and remember that a rule blocking a directory blocks everything under it.

3. X-Robots-Tag header

This carries the same directives as the meta robots tag, but in the HTTP response instead of the HTML. That makes it easy to miss, because you will not see it in the page source. A noindex set here, often by a CDN rule or a framework default, deindexes the page just as hard as one in the <head>, and it is invisible unless you inspect response headers.

4. Meta robots

<meta name="robots" content="noindex"> in the head is the most common self-inflicted deindex we see: a staging flag that shipped to production, a CMS setting, a plugin toggle someone flipped and forgot. If the page should be indexed, this tag should be absent or say index.

5. Canonical

<link rel="canonical"> tells Google which URL is the real one. If it points at a different URL, you are asking Google to index that one instead of the page in front of you. If it points at a URL that 404s or redirects, you have broken the signal and Google will guess. A page you want indexed should have a canonical that points at itself.

How the checks combine

Our diagnose() resolves all of these deterministically first, with no model involved. Only then does it label the page: INDEXED in our vocabulary means “indexable, no technical blocker found,” alongside NOT_INDEXED and UNKNOWN. That label is our inference, kept separate from anything Google reports. If an AI summary runs on top, it only prioritises and phrases the findings. The evidence is always the deterministic checks.

Run one URL through the free checker to see every check at once. The product runs the same pass across every URL in your sitemap and re-runs it on a schedule, so a canonical that breaks next month shows up as a change instead of a slow disappearance. See how the pipeline fits together.