Tabkeel
September 20, 2026·Francisco Ferreira·17 min read

How to Get Your Site Indexed by Google

Quick answer

To get your site indexed by Google, verify the domain in Google Search Console, use the URL Inspection tool to request indexing of your homepage, and submit an XML sitemap. If nothing appears after two weeks, the cause is almost never patience: it's a noindex tag, a robots.txt block, a canonical pointing at the wrong URL, or a page that renders only in the browser. Check those four before you wait any longer.

A new site takes Google anywhere from a few hours to a few weeks to index, on John Mueller's own estimate. So the first advice every founder gets is to wait. That advice is fine right up until it isn't, and the sites that are still absent after a month are usually not slow. They're blocked, by a single setting nobody chose on purpose. That setting gets shipped by the tool, not the person, which is why AI-built sites carry it so often: Lovable, v0, Cursor and Claude Code build the screen you asked for and inherit the defaults you never saw.

This page is the full walk-through of getting indexed, and then the part the generic guides skip: the four failure modes specific to a site an AI assistant generated, each with the exact check and the fix. If Google can crawl your site fine and you just want more clicks from the traffic you already earn, that's a different problem, covered in how to improve CTR in Google Search Console. This page is the gate before that one: getting into the index at all.

1 lineof robots.txt, Disallow: / under the wrong agent, can remove every page you have from Google
HTTP 200the "everything's fine" status a broken 404 wrongly returns, which tells Google to keep a page that says "not found"
2 weeksa fair wait for a healthy new site; past that, suspect a block, not slow crawling

First, check whether you're already indexed

Before fixing anything, find out what Google currently knows. Two checks, under a minute each, and they point you at the right problem instead of a guessed one.

Type site:yourdomain.com into Google. The result count is a rough census of how many of your pages are in the index. Zero results on a site that's been live for weeks is a loud signal: this is a block, not a delay. A handful of results when you expected dozens means some pages are getting in and others aren't, which usually points at a per-page problem like a scattered noindex.

Then open Search Console and run your homepage URL through the URL Inspection tool. It tells you the exact status in Google's words: "URL is on Google," or "URL is not on Google" with a reason attached. That reason line is the whole game. "Blocked by robots.txt," "Excluded by noindex tag," "Alternate page with proper canonical tag," and "Discovered, currently not indexed" each send you to a different fix, and the sections below are organized by those reasons.

How to get your site indexed by Google, step by step

For a clean site with no blocks, indexing is four steps, in this order:

  1. Verify your domain in Google Search Console (DNS verification covers every page on the domain at once).
  2. Run your homepage through the URL Inspection tool and click Request Indexing. This nudges Google to crawl, and from your homepage it discovers the pages linked from it.
  3. Generate an XML sitemap listing every page you want found, then submit it under Sitemaps in Search Console.
  4. Give it time, then re-inspect. A healthy new site starts appearing within days to a couple of weeks.

If you've done all four and pages still aren't showing up, you're not in the waiting category anymore. Something on the site is telling Google to stay out, and the next four sections are the four things it usually is, worst first.

The four ways an AI-built site blocks its own indexing

Every one of these ships silently. The page looks perfect in your browser, converts fine, and is invisible to search, and nothing in the UI ever tells you. This is the front Tabkeel's exam calls search visibility, and here's what it checks, in severity order.

1. A noindex tag shipped from staging

A noindex is one meta tag: <meta name="robots" content="noindex">. It does exactly what it says, tells Google to drop the page from search, and it exists for good reasons on staging environments where you don't want a half-built preview competing with your real site. The failure is that the tag rides along to production. The preview deploy had noindex switched on so Google wouldn't index the work-in-progress, the switch never got flipped back, and the launch went out telling Google to ignore the entire site.

Tabkeel ranks a noindex on the homepage as a blocker, the top severity, because it takes out everything at once. A noindex on a few interior pages is high severity instead: those specific pages vanish while the rest survive. To check by hand, view the page source (not the rendered DOM) and search for noindex. If it's on your homepage, that one line is why site:yourdomain.com returned nothing.

2. A robots.txt that disallows the whole site

robots.txt has a single line that ends crawling on the spot: Disallow: / under User-agent: * or User-agent: Googlebot. It tells Google not to crawl a single URL on the domain. Starter templates ship this deliberately so a scaffold doesn't get indexed before it's real, and then the placeholder becomes the launch.

Open yourdomain.com/robots.txt directly in a browser and read it. A Disallow: / with no matching Allow under * or Googlebot is a blocker. Worth knowing: robots.txt blocks crawling, not indexing, so a page blocked here can still show up in results as a bare URL with no description if Google found it linked elsewhere. Either way it can't rank on content it was never allowed to read.

3. A canonical baked to a preview URL

This is the one almost nobody checks, and the one most specific to AI-built sites. A canonical tag tells Google which URL is the real, authoritative version of a page. AI builders and preview platforms often hardcode it to the deploy URL they generated the site on: a *.vercel.app, *.netlify.app, *.pages.dev or even a localhost address. So your production homepage at your real domain carries a canonical pointing somewhere else. Google reads that as "the real version of this page lives over there," consolidates all the ranking signals onto a URL you don't control, and your actual domain never ranks.

It's insidious because the site is fully crawlable, fully indexable, no noindex, no robots block, and still doesn't rank, because you handed the credit to a preview address. View source on your homepage and find <link rel="canonical">. If the href isn't your production domain, that's the finding. Tabkeel flags it as high severity and recognizes the common preview hosts specifically, because a canonical to your-app.vercel.app is a fingerprint of a site that shipped its build URL by accident.

4. No sitemap, so new pages take forever to be found

A sitemap is a map of your pages. Without one, Google finds pages only by following links, which is slower and misses anything not linked from somewhere it already crawls. Plenty of AI builders don't generate a sitemap.xml at all, and don't reference one in robots.txt either. This won't keep an indexable page out of Google forever, which is why it's medium severity, not a blocker, but it stretches "a few days" into "a few weeks" and leaves orphaned pages undiscovered.

Check yourdomain.com/sitemap.xml in a browser. If it 404s and robots.txt has no Sitemap: line pointing to one, generate a sitemap and submit it. Every framework has a plugin or a built-in route for this, and it's ten minutes of work that shaves weeks off discovery for a growing site.

A real finding: the canonical that pointed at Vercel

Here's the pattern the exam keeps surfacing, because the canonical trap is the one people don't believe until they see it on their own page. A site looks flawless. No noindex. robots.txt is clean. The homepage is in Google's index. And it still won't rank for the founder's own brand name, the one query a live site should own outright.

The exam found a canonical tag on the production homepage pointing at the project's original *.vercel.app preview URL. Google had done exactly what the tag asked: treated the preview deploy as the master copy and quietly folded the real domain's signals into a URL that wasn't even meant to be public. The fix isn't code you paste, it's a prompt you hand your agent, because the exact file depends on the framework:

Find where the canonical link tag or metadata is set for the site (a layout file, a head component, or the framework's metadata config) and make every page's canonical resolve to the production domain, not the deploy or preview URL. Base it on the request's real host or a configured production URL, never a hardcoded vercel.app, netlify.app or localhost address. Confirm the homepage's canonical is self-referential and points at the live domain.

With the canonical pointing home again, the signals stop leaking to a preview address the founder doesn't own, and the real domain can finally accumulate its own. Nothing about the content changes. The site just stops telling Google to credit somebody else's URL.

Soft 404s: the page that says "not found" and returns 200

A soft 404 is a page that shows a "not found" message to a human while returning HTTP status 200 (the code for "OK, here's a real page") to a crawler. Google treats these as low-quality and drops them from the index, and it can drag down how it reads the rest of the site. AI-built apps produce them constantly, because a client-side router catches every unknown path, renders an error component, and never sets the response status, so every typo URL and every deleted page answers "200 OK, everything's fine" while displaying an error.

Google's own Page Indexing report calls these out under "Soft 404." The honest test is to request a URL you know doesn't exist, say yourdomain.com/this-page-is-fake, and check the actual HTTP status the server returns rather than what the page looks like. A real 404 returns status 404. A soft 404 returns 200 with an error page, which is the bug. Tabkeel catches this under its missing-states front rather than search visibility, but the effect is an indexing problem: Google won't keep pages that lie about being found.

One check the automated pass doesn't cover

A noindex can also live in an HTTP response header (X-Robots-Tag: noindex) instead of a meta tag, and some preview platforms add it to protect staging. Tabkeel's search-visibility check reads the meta tag in the served HTML, where the AI-built noindex lives almost every time, but a header-level noindex needs a look at the raw response headers (your browser's Network tab, or curl -I yourdomain.com). If the meta tag is clean and Search Console still says "Excluded by noindex tag," check the headers; the full triage for that status is in Excluded by 'noindex' tag: what it means and how to fix it.

The indexing blockers, ranked by severity

Not every problem here weighs the same. A blocker takes out the whole site today; a medium issue just slows discovery. Work top to bottom, and stop worrying about the lower rows until the top ones are clear.

SeverityBlockerWhat it doesFastest check
Blockernoindex on the homepageDrops the entire site from GoogleView source, search for noindex
Blockerrobots.txt Disallow: / under * or GooglebotGoogle can't crawl a single pageOpen /robots.txt
HighCanonical pointing at a preview or localhost URLSignals consolidate to a URL you don't control; your domain never ranksView source, read rel="canonical"
Highnoindex on individual should-rank pagesThose pages vanish; the rest surviveURL Inspection per page
MediumMissing sitemap.xmlNew and orphaned pages take much longer to be foundOpen /sitemap.xml
MediumSoft 404 (error page returning HTTP 200)Google drops the page and distrusts the patternRequest a fake URL, check the status code
LowMissing self-referential canonicalDuplicate URL variants compete against each otherView source, check for any canonical

"Discovered, currently not indexed": what Google is actually telling you

Two Search Console statuses confuse people more than any others, and they mean different things. "Discovered, currently not indexed" means Google knows the URL exists but hasn't crawled it yet, often a crawl-budget or a thin-connection issue on a new site with few internal or external links. "Crawled, currently not indexed" means Google fetched the page, looked at it, and decided not to keep it, which is almost always a quality or duplication signal, not a technical block.

The reflex is to keep hitting Request Indexing. It rarely helps for these two, because neither is a "Google didn't notice" problem. "Discovered" wants better internal linking and a few real inbound links so the page looks worth crawling. "Crawled, not indexed" wants a page that isn't a near-duplicate of another and actually says something. If you're seeing either at scale on a brand-new site, the underlying issue is often the render trap in the next section: Google crawled a mostly empty shell and had nothing worth indexing.

The render trap: a page that's empty until JavaScript runs

Google can index a JavaScript site, but it does it in two passes: it reads the raw HTML first, then comes back later to render the page and read what the scripts built. On a client-rendered app, that first pass gets a near-empty shell, a <div id="root"></div> and a script tag, with your headline, copy and pricing all assembled in the browser afterward. Google eventually renders it, but slower, with weaker signals in the gap, and answer engines that never run JavaScript get nothing at all.

The test that actually tells you the truth is to look at what the server sends before any script runs. Use view-source (not "Inspect," which shows the DOM after JavaScript executed), or run curl yourdomain.com from a terminal. If your real content isn't in that raw response, Google is working harder than it should to index you and some crawlers aren't managing it at all. The full version of this front, what a machine reads before rendering, is the subject of how to make your website readable to AI, and it's worth clearing early because it sits underneath both indexing and AI visibility.

Getting indexed is not the same as getting cited

Indexed by Google

Your pages are in the search index and can appear when someone searches. This is what every check above is about.

Visible to AI answers

ChatGPT, Perplexity and AI Overviews can read your site and name your product. A separate bar, with its own crawlers and its own failure modes.

A site can be perfectly indexed by Google and still be invisible when someone asks an AI assistant what your product does, because the answer engines use their own crawlers, most of which don't render JavaScript and several of which get blocked by the same "protect my content from AI" robots.txt line that seemed harmless. Indexing gets you into Google. Being named by an AI is the next, separate job, and it's what getting cited by ChatGPT and AI search walks through. Both matter now, and neither one covers for the other.

Mistakes that keep a site out of the index

Waiting when you should be checking. "Give it time" is right for the first two weeks and wrong after. Past that window, a healthy site is indexed, so persistent silence means a block, and no amount of waiting clears a noindex tag.

Requesting indexing over and over on a "Crawled, not indexed" page. Google already crawled it and passed. Hammering the button doesn't change its verdict; a page worth keeping does.

Checking the page in Chrome and calling it fine. Chrome runs your JavaScript and follows none of your robots rules. It's the least representative reader of your site. view-source and /robots.txt tell you what a crawler sees; the browser tells you what you see.

Fixing the sitemap while a blocker is still live. A perfect sitemap submitted under a homepage noindex is careful work spent on a site Google was told to ignore. Clear the blockers first, then help discovery.

Run the whole check at once

You can work this list by hand: site: and URL Inspection for the census, view-source for the noindex and canonical, /robots.txt and /sitemap.xml in the address bar, and a fake URL to smoke out a soft 404. Or point the free search-visibility tool at your URL and it runs the noindex, robots.txt, canonical and sitemap checks in one pass and shows each finding with the evidence attached, nothing to install. For all seven fronts of an AI-built launch, including the render trap and the missing-states check behind soft 404s, the full Tabkeel exam reads the whole site and writes each fix as a prompt you can paste into your agent, with the most severe one shown in full for free. The seven fronts and how they connect are laid out in the pre-launch checklist for AI-built sites, and the scoring behind each one is in the methodology.

Frequently asked questions

How long does it take Google to index a new site?

Anywhere from a few hours to a few weeks, by Google's own estimate, with popularity, internal linking and inbound links all affecting the speed. Two weeks is a fair wait for a healthy new site. Past that with nothing showing up, the cause is usually a block (a noindex tag, a robots.txt disallow, or a canonical pointing at the wrong URL) rather than slow crawling, and requesting indexing in Search Console won't clear a block.

How do I check if Google has indexed my site?

Search site:yourdomain.com in Google for a rough count of your indexed pages, then run your homepage through the URL Inspection tool in Google Search Console for the exact status and the reason behind it. Zero results on a site that's been live for weeks means a block, not a delay.

Why is my website not showing up on Google even though it works?

Because "works in a browser" and "readable by Google" are different things. The four common causes on a working site are a noindex meta tag left on from staging, a robots.txt that disallows the whole site, a canonical tag pointing at a preview or localhost URL, and a page that renders only in the browser so the first crawl gets an empty shell. All four are invisible in the browser and each one is checkable in under a minute.

What does "Discovered, currently not indexed" mean?

Google knows the URL exists but hasn't crawled it yet, usually a crawl-budget or thin-connection issue on a new site with few links pointing at the page. It's different from "Crawled, currently not indexed," where Google fetched the page and chose not to keep it, which points at a quality or duplication problem. Better internal linking and a few real inbound links help the first; a page that isn't a near-duplicate and says something helps the second.

Does a robots.txt block remove my site from Google?

It stops Google from crawling your pages, which means they can't rank on their content, but a blocked URL can still appear as a bare link with no description if Google found it referenced elsewhere. To truly keep a page out of search, use a noindex tag on a page Google is allowed to crawl, not a robots.txt disallow, because Google has to be able to read the noindex for it to count.

FF
Francisco Ferreira
Builds Tabkeel and runs the exam on AI-built sites every day. About

See what AI and Google read on your site

The exam crawls your public site and returns every finding with the evidence and a paste-ready fix. Free, no signup.

Run the free exam

More articles

← All articles