Excluded by 'noindex' Tag: What It Means and How to Fix It
"Excluded by 'noindex' tag" means Google crawled the URL, found a noindex directive, and obeyed it. For a login page or a 404 that is correct, so leave it. For a page you want ranking, remove the noindex from wherever it lives (a meta tag, an X-Robots-Tag header, a platform setting or a script), then click Validate Fix in Search Console. Google says validation usually finishes within about two weeks.
Google's Page Indexing report has one line that sends more founders into a panic than any other: "Excluded by 'noindex' tag," followed by a count that is often larger than the number of pages you think you have. The report is doing its job. Google found an instruction to stay out and respected it. The work is deciding which of those URLs you actually meant to hide, and tracking down where the instruction came from on the ones you didn't.
This guide is the triage. It sits under our broader walk-through of how to get your site indexed by Google, and goes deep on the one status that page could only flag in passing: what the excluded-by-noindex list means, how to sort it in ten minutes, and the four places a noindex hides on a site that was built with Lovable, v0, Cursor or Claude Code.
What "Excluded by 'noindex' tag" actually means
A noindex directive is an instruction, sent with a page, telling search engines not to include that page in results. Google's own Page Indexing documentation describes the status plainly: when Google tried to index the page, it encountered a noindex directive and therefore did not index it.
A noindex tag is a robots directive telling search engines to keep one page out of their results. Its common companion, nofollow, is a separate instruction about the page's links, and the two often travel together as noindex, nofollow. Four details about the status matter for the fix:
- Google got in. The page was crawled, so robots.txt isn't the issue here. That rules out the most common alternative culprit and narrows the search to the page's own response.
- The directive can come from two channels. Google's noindex spec accepts a
<meta name="robots" content="noindex">tag in the HTML, or anX-Robots-Tag: noindexHTTP response header. The Page Indexing report doesn't say which one it saw. - URL Inspection does say. Inspect the URL and read the "Indexing allowed?" line.
No: 'noindex' detected in 'robots' meta tagsends you to the HTML.No: 'noindex' detected in 'X-Robots-Tag' http headersends you to the server or hosting config. That one line saves most of the hunting below. - "Excluded" is only a status. Your site's standing is untouched. The page is simply absent, and it comes back once Google recrawls it without the directive.
So the report answers "which URLs carry a noindex," and leaves you to answer "which of those should."
How to fix "Excluded by 'noindex' tag"
- Open Indexing, then Pages, in Search Console and click the "Excluded by 'noindex' tag" row to see the affected URLs.
- Separate the URLs that should stay hidden (login, checkout, 404 pages) from the pages that should rank.
- Run URL Inspection on each page that should rank to learn whether the noindex comes from the robots meta tag or the X-Robots-Tag header.
- Remove the directive at that source, or restrict it to preview deployments only.
- Test the live URL, request indexing, and click Validate Fix on the report row.
Step 2 is where most of the time goes, and where most wasted fixes happen, so the next two sections are about sorting. Steps 3 and 4 get their own section after that.
Most excluded URLs are supposed to be there
The count in the report is rarely the number of broken pages. Apps generate dozens of URLs that should never rank: auth screens, account pages, checkout steps, share links, and error pages. A well-built site marks those noindex on purpose, and every one of them lands in this report.
/login, /signup, /account/*, /checkout, thank-you pages, internal search results, share and certificate links, and your 404 page.
Leave these alone. The report is confirming your setup works.
Your homepage, /pricing, /docs, /blog/*, /about, /changelog: any page a stranger should be able to find from a search.
These are leaks. Fix them first.
Our own site is a fair example. Request a URL that doesn't exist on tabkeel.com and the server answers HTTP 404 with <meta name="robots" content="noindex">; our login page answers 200 with the same tag. Both would sit in this report, and both are exactly where we want them. The homepage, /check and every blog post carry no robots meta and no X-Robots-Tag header (checked with curl on September 22, 2026).
That split is exactly how Tabkeel's search-visibility check decides whether a noindex becomes a finding. It only flags the tag on a page that answered with a 2xx status and belongs to a type that should rank: home, pricing, docs, changelog, blog or about. We learned the 2xx rule the hard way. An early version of the check flagged a noindex on an /about-style URL that didn't exist; the site had correctly routed it to a branded 404 page, and a 404 page carrying noindex is the right behavior, not a defect. A report that cries wolf on error pages trains you to ignore it, so the check now stays quiet on anything that isn't a real, rankable page.
The four-bucket noindex triage
Export the example URLs from the report (Indexing, then Pages, then the "Excluded by 'noindex' tag" row) and drop each one into a bucket. The bucket decides the action, and it keeps you from "fixing" pages that were never broken.
| Bucket | What it looks like | Verdict | Action |
|---|---|---|---|
| 1. Intended | Auth, account, checkout, thank-you, search results, 404 | Correct | Leave the noindex. If the URL is also in your sitemap, take it out: a sitemap entry asks Google to index a page the tag forbids. |
| 2. Leaked | A should-rank page with a noindex visible in view-source | Bug | Remove the tag, or make it conditional on the preview environment. Then validate. |
| 3. Hidden source | A should-rank page whose view-source is clean, yet the report still lists it | Bug, elsewhere | Check response headers, the rendered DOM, and platform settings (next section). |
| 4. Conflicted | A page with a noindex and a canonical pointing at a different URL | Mixed signal | Pick one directive. See the canonical section below. |
The report shows at most 1,000 example URLs. For a bigger list, a desktop crawler such as Screaming Frog records the meta robots and X-Robots-Tag value of every page in one pass, which makes the bucketing a spreadsheet filter. And if you remember an older error called "Submitted URL marked 'noindex'," that was bucket 1's sitemap contradiction under a louder label.
On most AI-built sites, bucket 1 holds the bulk of the list and bucket 2 holds the damage, often a single template-level tag that took the homepage with it. Bucket 3 is where people lose an afternoon, because every guide tells them to check the meta tag and the meta tag is clean.
Four places a noindex hides
If the page is a should-rank page and Google says noindex, the directive is coming from one of four sources. Check them in this order, because the first two cover almost every case on a modern JavaScript stack.
1. A meta tag in the served HTML
The classic source. Open view-source (not DevTools Elements, which shows the page after scripts ran) and search for noindex, or for content="none", which means noindex plus nofollow. On a Next.js app the tag usually comes from metadata rather than hand-written HTML: a robots: { index: false } entry in a layout's metadata export renders a noindex on every page under that layout. Put it in the root layout "for staging" and it ships on the whole site.
2. An X-Robots-Tag response header
If URL Inspection named the X-Robots-Tag header, start here. A header-level noindex never shows up in the HTML at all, which is why it survives every view-source check. Run curl -sI https://yourdomain.com | grep -i x-robots, or open DevTools, pick the document request in the Network tab, and read the response headers. On Vercel this has a specific, documented pattern: Vercel adds X-Robots-Tag: noindex to preview deployments automatically, and skips it when a custom domain is assigned to a non-production branch. Headers set in vercel.json, next.config headers(), a Netlify _headers file or middleware can all add it on production by mistake, usually copied from a staging config. On an Apache server it's typically one line in .htaccess: Header set X-Robots-Tag "noindex".
3. A tag injected by JavaScript
Google renders JavaScript, so a noindex that a script inserts after load still counts. The mirror case trips more people: a page ships with noindex in the initial HTML and a script removes it on the client. Google's JavaScript SEO documentation warns that when it encounters the tag, Google may skip rendering entirely, so removing a noindex with JavaScript may not work. If the served HTML says noindex, treat the page as noindexed, whatever the browser shows after hydration.
4. A platform, plugin or CMS setting
Some builders and CMSs expose indexing as a toggle rather than code. WordPress's "Discourage search engines from indexing this site" checkbox under Settings, then Reading, is the famous one: tick it during development, forget it at launch, and every page carries a noindex. SEO plugins add a second layer: Yoast SEO and Rank Math both keep a per-page robots setting, so one post can be noindexed while the site-wide checkbox is clear. Hosted builders put the switch in page settings, such as the "Let search engines index this page" toggle in Wix's SEO panel or the option in Squarespace's page settings to hide a page from search results. If the code is clean and the headers are clean, look for the switch.
Tabkeel's search-visibility check reads the meta robots and googlebot tags in the HTML your server sends, which is where the AI-built noindex lives in the large majority of cases. It does not yet read X-Robots-Tag headers or the post-JavaScript DOM, so sources 2 and 3 above still need the manual curl and rendered-DOM checks. A clean result from the tool rules out source 1, not all four.
A real finding: the staging default that shipped
The typical pattern on sites generated by AI coding tools: a founder launches, shares the link, gets signups from their own audience, and hears nothing from Google for weeks. Search Console shows the homepage under "Excluded by 'noindex' tag," and view-source shows a robots meta that was switched on for the preview build and never switched off.
The finding the Tabkeel exam writes for that case is ranked blocker, the top severity, with this evidence line:
the home page carries a noindex: you are telling Google to drop the WHOLE site from search, almost certainly a staging/preview leftover shipped to production
A homepage noindex is a blocker because it removes the page every other page hangs from; a noindex on a few interior pages is ranked high instead, since the rest of the site survives. The fix Tabkeel hands back is a prompt for your coding agent, since the right file depends on the framework (shown here without the quoted finding it carries):
Remove the noindex directive so Google can index the site. Search for a <meta name="robots" content="noindex"> (or "none") in the <head>: it is almost always a staging/preview default that got shipped to production. Delete it, or make it environment-conditional so it only applies to preview deploys, never to the live domain.
The environment-conditional option is the one to prefer. Deleting the tag fixes today; tying it to the environment (index only when the deployment is production, on the production domain) keeps the next preview from leaking the same way. After the fix ships, view-source on the live homepage should show no robots meta at all, or one that says index, follow.
Canonical and noindex on the same page
Bucket 4 deserves its own section, because it's the case where two correct-sounding instincts collide. A canonical tag says "this page is a copy; index that other URL instead." A noindex says "don't index this page." Put both on one URL and you've told Google to consolidate this page's signals into another URL while also telling it to drop this page, and Google has to guess which you meant.
Google's guidance has hardened over time. In 2021 John Mueller hedged that using both "maybe" forwards some signals; by 2024 his advice was to just pick one, as Search Engine Journal documented. The practical rule:
- Duplicate you want merged into a main page (a filtered listing, a UTM-tagged variant, a print view): canonical only, no noindex.
- Page you want gone from search entirely (a thin tag page, an internal utility): noindex only, and no canonical pointing elsewhere.
AI-built sites land in bucket 4 in a specific way. The template sets a canonical on every page, a developer adds noindex to a few utility routes, and those routes now carry both. The page still disappears from search, so nobody notices, but any links pointing at it are asking Google to trust a URL that is simultaneously telling Google to ignore it.
After the fix: validate, then give it time
Once the noindex is gone from production, do three things in order. First, run the homepage through URL Inspection, click Test Live URL, and confirm the live test says indexing is allowed; this proves the fix is actually deployed, not just merged. Second, click Request Indexing on that URL. Third, open the "Excluded by 'noindex' tag" row and click Validate Fix, which asks Google to recheck every listed URL as a batch.
Validation will fail for the bucket 1 URLs, because they still carry an intended noindex. That's expected. Judge the result by whether your should-rank pages move to indexed; the validation status can read Failed purely because of bucket 1.
To see recovery in clicks rather than in status labels, connect Search Console to Tabkeel's Search analytics: the 90-day pulse shows a page that was invisible start collecting impressions again, and marking the change "in test" gives you a before-and-after verdict instead of a guess. Connecting and seeing the opportunities is free.
Five mistakes that keep pages excluded
Blocking the page in robots.txt to "reinforce" the noindex. Google's spec is explicit: for noindex to work, the page must not be blocked by robots.txt. Block the crawl and Google never sees the tag, so the URL can linger in results as a bare link.
Removing the tag with client-side JavaScript. If the served HTML still says noindex, Google may never render far enough to see your script remove it. Fix it on the server.
Trying to get every excluded URL indexed. A login page in Google's index is worse than a login page out of it. Bucket 1 is supposed to stay excluded.
Checking only the HTML. A clean view-source with a persistent report entry points at a header. curl -sI takes five seconds and ends the mystery.
Fixing staging instead of production. On Vercel, a preview URL carrying X-Robots-Tag: noindex is by design. What matters is the header and HTML on the domain customers type in.
The fastest way through bucket 2 on a live site is to let a crawler do the page-by-page reading. The free tool at tabkeel.com/tools/search-visibility goes through your pages, quotes the noindex line wherever it finds one on a should-rank page, and ignores the error and auth pages that are meant to carry it. If the site was built with an AI coding tool and you want every front checked before it costs you another month of silence, run the full Tabkeel exam: search visibility is one of seven fronts, alongside whether AI crawlers can read your pages at all. And once the pages are indexed, the next lever is turning impressions into clicks, which is what improving CTR in Search Console covers.
Frequently asked questions
What does "Excluded by 'noindex' tag" mean in Google Search Console?
It means Google crawled the URL, found a noindex directive in either a robots meta tag or an X-Robots-Tag HTTP header, and left the page out of its index as instructed. Nothing about your site's standing changes, and the page returns to search after Google recrawls it without the directive.
Should I fix every URL marked "Excluded by 'noindex' tag"?
No. Login, account, checkout, thank-you, internal search and 404 pages should stay excluded, and seeing them in the report means your setup works. Only fix pages you want ranking: the homepage, pricing, docs, blog posts, about and changelog pages.
Why does Search Console say noindex when I can't find the tag in my HTML?
The directive is probably in an X-Robots-Tag response header, which never appears in the HTML. Run curl -sI on the URL and look for it. Other hidden sources are a tag injected by JavaScript after load and a platform setting such as WordPress's "Discourage search engines" checkbox.
How long does it take for a page to be indexed after removing noindex?
Google says the Validate Fix process typically takes up to about two weeks, and sometimes longer. Requesting indexing on important URLs through URL Inspection speeds up the recrawl. Low-importance pages left to natural crawling can take months to be revisited.
Can I use canonical and noindex on the same page?
You can, but Google advises picking one. A canonical asks Google to merge the page into another URL; a noindex asks Google to drop it. Use canonical alone for duplicates you want consolidated, and noindex alone for pages you want out of search.
See what AI and Google read on your site
The exam crawls your public site and returns every finding with the evidence and a paste-ready fix. Free, no signup.
Run the free examMore articles