Website indexing is the process by which Google discovers, crawls, and stores a copy of your web pages so they can appear in search results. Without it, your pages are invisible to search — no matter how good the content. If you're serious about SEO in South Africa, understanding how Google's index works — and what can break it — is the first step to fixing it.

Google holds over 90% of South Africa's search engine market (2026 estimates), which means Googlebot is the crawler that matters most to local businesses. What many owners miss is that indexing does not happen automatically and is not assured: pages can be blocked, ignored, or actively excluded — often due to technical errors that are easy to fix once you know what to look for.

This guide explains how the indexing process works, what causes pages to stay out of Google's index, and how to diagnose and fix the most common problems — including a diagnostic table that maps every major Google Search Console status to its most likely cause and fix.

Quick Answer

The process practitioners call website indexing runs in three stages: an automated crawler visits your page URL and reads its content (crawl), the HTML and JavaScript are processed to understand what users see (render), then the page is added to the index and becomes eligible to appear in search results (index). Common reasons a page stays out of the index include a noindex tag, robots.txt blocking, thin or duplicate content, server errors, and redirect problems. You can check your status and diagnose issues for free in Google Search Console's Page Indexing report.

Not sure which of your pages Google can actually see?

Send us your domain and we'll pull your Page Indexing report and show you exactly which URLs are missing — and why.

Get a free indexing audit

How Website Indexing Works

Google website indexing follows a three-stage pipeline that runs continuously across billions of pages worldwide: crawl, render, index. Each stage can fail, and the cause of failure at each stage is different.

1. Crawl. Googlebot — specifically the Googlebot Smartphone variant, which simulates a mobile user — visits your page by following links or reading your sitemap. Google's own documentation confirms that "the vast majority of the new pages Google finds every day are through links," which makes your internal and external link structure more important to indexing than any other single factor. Google's published crawler documentation references a 2MB processing limit per HTML file — SEO practitioners treat this as the working threshold above which page content risks being cut off before Google finishes reading.

2. Render. After crawling, Google renders your page — it processes your HTML, CSS, and JavaScript to see what a user would actually see. Pages that hide important content behind JavaScript that Googlebot cannot execute may not be understood correctly. Google aims to "see the page the same way an average user does," so anything that prevents access to CSS and JS resources can reduce how well a page is understood and indexed.

3. Index. If Google considers the page sufficiently unique and valuable, it stores a copy in its index and the page becomes eligible to appear in search results. "Eligible" does not mean it will rank — that depends on relevance, authority, and hundreds of other signals — but it cannot rank at all if it's not indexed first.

Key Fact

Google's SEO Starter Guide states that for most sites, you "usually don't need to do anything except publish your site on the web" to be discovered — but discovery does not equal indexing. A page can be discovered and still be excluded from the index for quality, technical, or content-related reasons.

Why SA Businesses Cannot Afford Unindexed Pages

For website indexing South Africa presents a near-single-engine market: over 90% of local searches run through Google, which means Google's index is the only one that materially affects organic revenue. A page that is not in Google's index generates zero organic traffic, regardless of how much effort went into writing it or building the site. For businesses running content marketing, local SEO, or product pages, an unindexed URL is a dead asset. When important pages stay out of the index after you have fixed the basics, our SEO services in South Africa start with a free audit that shows which URLs Google is ignoring and why.

The mobile-first context makes this more pressing locally. Google completed its switch to mobile-first indexing in October 2023 — meaning Googlebot now uses the mobile version of your site as the primary source for what gets indexed and how. If your mobile version is slower, has less content, or is structured differently from your desktop version, what Google indexes may not reflect your best work. Given that most South African internet users access the web via smartphone, the stakes of mobile quality are double: poor mobile experience harms both real users and how your content is understood by Googlebot.

Server performance is a particularly South African concern. Sites hosted on local shared servers that experience downtime during load-shedding cycles, or that respond slowly due to constrained bandwidth, send a stressed-server signal to Googlebot. When Google detects high response latency or frequent 5xx errors, it throttles its crawl rate automatically — which means new pages take longer to reach the index, and already-indexed pages may not be refreshed as frequently.

For a practical look at how to improve your site's crawlability before worrying about indexing, that guide walks through the technical foundations that determine whether Googlebot can even reach your pages.

What Keeps Pages Out of Google's Search Results

Common website indexing problems prevent organic search traffic from reaching a site — and website indexing status is something Search Console names precisely for every excluded URL. The table below maps the most common GSC Page Indexing statuses to their most likely cause — including South African-specific triggers — and a priority level for your fix queue.

GSC StatusWhat it meansCommon SA triggerFixPriority
Crawled – currently not indexedGoogle found the page but judged the content too thin, duplicate, or low-value to add to the indexThin category pages, near-duplicate product listings, auto-generated tag pagesImprove content depth or consolidate duplicates; add noindex to low-value pages you don't need to rankHigh
Discovered – currently not indexedGoogle found the URL (via sitemap or links) but hasn't crawled it yet — it's in a queueNew site or section with weak internal linking; pages only reachable from sitemap, not from navigationAdd meaningful internal links from existing indexed pages pointing to the new URLMedium
Blocked by robots.txtYour robots.txt file is preventing Googlebot from crawling the pageStaging-environment robots.txt accidentally pushed to production during a site migration or redesignCheck robots.txt immediately; remove the disallow rule for any page you want indexedCritical
Excluded by noindex tagA noindex directive in the page's meta robots tag or HTTP header is telling Google not to index the pageDeveloper added noindex during build; tag not removed before launch; plugin setting applied site-wide by mistakeRemove the noindex directive from any page that should appear in search resultsCritical
Server error (5xx)Your server returned an error when Googlebot tried to crawl the pageLoad-shedding-related server downtime; shared hosting instability; misconfigured server after an updateResolve the underlying server issue; consider switching to hosting with a generator-backed UPS or a more reliable providerHigh
Not found (404)The page no longer exists and returns a 404 errorPage deleted without updating the sitemap; URL changed without a redirectRemove from sitemap; redirect to a relevant live page if the content exists elsewhere; or let it 404 cleanlyMedium
Duplicate, Google chose different canonicalGoogle disagrees with your canonical tag and is indexing a different version of the pageFaceted navigation (filter URLs), www vs non-www inconsistency, HTTP vs HTTPS mixed signalsStrengthen canonical signals: ensure the canonical tag, sitemap, and internal links all point to the same preferred URLMedium
Page with redirectThe URL redirects somewhere; the original URL is not indexedPost-migration redirects pointing to the homepage instead of a genuinely relevant page (treated as a soft 404)Redirect to the most relevant live equivalent; homepage redirects for specific content pages signal a soft 404 to GoogleMedium

Blocking crawling vs. blocking search results: two separate controls

These two controls do different jobs and are frequently confused. robots.txt tells Google's crawler whether it may crawl a page. A noindex meta tag (placed in the page's own HTML) or an X-Robots-Tag HTTP response header tells Google not to index it. The noindex signal belongs in the page's HTML or as an HTTP response header.

The consequence: Google's own documentation states that a page blocked in robots.txt "can still be indexed if linked to from other sites." If you need a page out of Google's index, a noindex meta tag placed in the page's own HTML is the correct tool — robots.txt blocking alone will not reliably achieve this. For pages you want excluded from search results, allow crawling and apply a noindex meta tag in the page's HTML.

For a deeper look at the two most confusing GSC statuses, see the guides on what causes Crawled – Currently Not Indexed and what causes Discovered – Currently Not Indexed — both are common and have different root causes.

How to Check Whether Your Pages Are Indexed

Google Search Console's Page Indexing report (formerly the Coverage report) is the authoritative source for your site's index status — it shows every URL Google has encountered, whether it's indexed, and if not, why not.

To access it: open Search Console → select your property → click Indexing → Pages. The report splits URLs into "Indexed" and "Not indexed" tabs, and the "Not indexed" tab gives you reason codes — exactly the statuses in the table above. Click any reason to see the affected URLs, then use the URL Inspection tool on individual pages to see the last crawl date, the verdict, and any specific issues Google found.

Efficient use of GSC: Click into the "Not indexed" tab, sort by reason, and focus on the reason code with the most high-priority URLs first. A single robots.txt blocking rule or a site-wide noindex tag can affect thousands of pages at once — fixing it requires one change and resolves immediately after Googlebot recrawls.

Avoid: Using the site: operator (e.g., site:yourdomain.co.za) as a reliable index count. Google has confirmed that the site: operator is not an accurate count of indexed pages — it's a sampling tool. The Page Indexing report in Search Console is the only reliable source.

For a complete workflow on reading and acting on Search Console data for SEO, the guide to improving your site's indexation rate covers this in more detail, including how to request indexing for individual URLs after fixing issues.

How Long Does It Take Pages to Appear in Google?

Google's own guidance is deliberately wide: changes may take "a few hours" to "several months." The actual timeline depends on your site's crawl frequency, which is determined by your authority, quality of inbound links, server reliability, and content freshness. The following ranges are practitioner heuristics — consistent with what the SEO community observes — not official Google thresholds.

Site typeTypical indexing rangeWhat drives it faster
Established, high-authority site24–48 hours (new pages)Strong internal linking to new content; frequent updates signal to Googlebot to return often
Medium-authority site (1–3 years old, some backlinks)3–7 daysSubmitting the URL to GSC using URL Inspection → Request Indexing; XML sitemap updated promptly
New or low-authority site (under 12 months, few external links)1–4 weeks (sometimes longer)Building internal links from already-indexed pages; earning even a small number of external links to the domain

Waiting longer than four weeks without seeing a page indexed — and the URL Inspection tool showing no crawl date — usually points to a technical or content problem rather than Google simply being slow. Check the Page Indexing report for an explicit reason code before assuming a timing issue.

Don't confuse indexing with ranking

A page can be indexed within days and take months to rank competitively. Indexing makes a page eligible for search results; ranking requires sustained signals of authority and relevance. For realistic timelines on when SEO effort translates to traffic, see how long SEO takes in South Africa.

How to Help Google Index Your Pages Faster

SA site owners trying to improve website indexing are almost always better served by removing friction than by adding new tactics — Google is already trying to crawl and index your pages; the job is to stop getting in the way. The following actions have the most consistent impact on how quickly new pages reach the index.

1. Internal links from indexed pages. Google's primary discovery mechanism is following links. When you publish a new page, add a contextual link to it from at least one high-traffic, already-indexed page on your site. A blog post linked from your homepage or a main category page will be discovered faster than an orphaned page linked only from your sitemap. See the post on how to find orphan pages to identify URLs currently without internal links.

2. Submit an XML sitemap. A sitemap tells Google about every URL you consider important — it does not guarantee those URLs will be crawled or indexed, but it removes the discovery bottleneck, especially for new pages with few incoming links. Submit your sitemap via Google Search Console → Sitemaps. Keep it updated: include live URLs only, and remove 404s and redirected URLs promptly.

3. Use URL Inspection → Request Indexing. For individual high-priority pages, the URL Inspection tool in Search Console has a "Request Indexing" button. It does not force indexing — Google still decides whether to index the page — but it signals to Googlebot that this URL is worth a prompt visit. Use it for pages that matter most (key product or service pages, fresh news content) rather than applying it site-wide.

4. Fix server response time. When a server takes more than three seconds to respond consistently, Googlebot throttles its crawl rate to avoid overloading the server. This reduces how many pages are crawled per visit. For South African sites on shared hosting that goes down during load-shedding, the pattern of 5xx errors can compound into a significantly reduced crawl frequency. Monitoring uptime and responding quickly to server errors keeps your crawl allocation healthy.

5. Content quality. The single biggest lever for persistent indexing problems is content quality. Practitioners widely report that "Crawled – currently not indexed" is most common on pages with thin content, near-duplicate content, or very low user value. Improving depth, specificity, and uniqueness resolves far more of these cases than any technical fix.

Crawl budget: when it matters and when it doesn't

Most South African small business sites — under a few hundred pages — do not have a crawl budget problem. Google can comfortably crawl a well-linked site of that size without running out of crawl allocation. Crawl budget management becomes meaningful on large ecommerce sites with thousands of product and filter URLs, or media sites with large archives. If your site is under 500 pages and Googlebot still isn't indexing pages, the problem is almost certainly content quality or a technical block — not crawl budget.

Pages published but not showing up in Google?

Tell us your domain and which pages concern you. No obligation — we'll get back to you within 24 hours.

Request a free page indexing check

Why South African Businesses Choose Growth Pulse Media for SEO

Growth Pulse Media was founded by Dirk van Greuning, who built and scaled a large South African ecommerce business before starting the agency — which means the indexing and SEO strategies we apply are ones that have been stress-tested in real South African operating conditions, including hosting constraints, load-shedding-related downtime patterns, and the specific behaviour of Googlebot on local sites.

All work is executed in-house. We don't outsource site audits or technical fixes to junior staff or offshore contractors. When a client's Page Indexing report shows 400 URLs stuck at "Crawled – currently not indexed," a senior practitioner investigates the actual content and technical signals — not a template report.

Our client load stays limited by design so every engagement gets meaningful senior time. If you're spending money on content that isn't reaching Google's index, the problem is solvable — but only after someone has actually looked at your site.

Who This Guide Is NOT For

You want a quick technical shortcut. There is no shortcut to earning a place in Google's index on a content-thin site. If the pages not being indexed are genuinely low-quality, the fix is improving the content — not a plugin or a sitemap resubmission.

You're managing a large site without a technical SEO strategy. Large ecommerce or media sites with tens of thousands of URLs need a deliberate crawl budget strategy, canonical architecture, and log-file monitoring — the principles here are a starting point, not a complete playbook. See the guide on how to prioritise SEO fixes for a framework suited to complex sites.

You're looking for a ranking guarantee. Indexing is the entry requirement for ranking — not the outcome. A page that gets indexed on a new site with no authority may not rank for competitive terms for months. This guide solves the indexing problem; ranking is a separate, longer effort.

You need an immediate turnaround on a brand new domain. New sites with no backlinks and limited crawl history typically take 1–4 weeks (or longer) to get initial pages indexed. There is no configuration that meaningfully accelerates first-time indexing on a domain with no external signals — publishing consistently and building links is the only path.

Know your indexing problems but not sure where to start fixing?

Share your Search Console access or a domain URL and we'll map your indexing issues against your highest-value pages — so you fix in the right order.

Book a free SEO audit

Frequently Asked Questions About Page Indexing

What is the difference between crawling and indexing?

In SEO practice, crawling refers to when an automated search engine program visits a URL and reads the page content. Indexing is when Google stores a processed copy of that page in its search index and makes it eligible to appear in results. A page can be crawled without being indexed — Google may visit a page and decide it is too thin, too similar to other pages, or not useful enough to store. You need both: crawling is the prerequisite, indexing is the goal.

Can I stop Google from indexing a specific page?

Yes — use a noindex meta robots tag or an X-Robots-Tag HTTP response header on the page. Do not rely on robots.txt to prevent indexing: Google explicitly states that a page blocked in robots.txt can still be indexed if other sites link to it, because Googlebot cannot read the noindex tag on a page it isn't allowed to crawl. The noindex tag is the correct and reliable tool for excluding a page from the index.

Why is my page indexed but not ranking?

Indexing and ranking are separate processes. A page being in Google's index means it is eligible to appear in search results — it doesn't determine where it appears. Ranking depends on relevance (how well the page matches the query), authority (the strength of your site's link profile), user experience signals, and competition from other indexed pages. A newly indexed page on a low-authority site may rank on page 10 or not at all for competitive terms.

How do I check which pages on my site are indexed?

Open Google Search Console, go to Indexing → Pages. The "Indexed" count shows how many pages Google has added to its index; the "Not indexed" tab shows which URLs are excluded and why. For individual URLs, use the URL Inspection tool — enter any URL on your site and Search Console shows the last crawl date, rendering outcome, and whether the page is indexed. This is more reliable than the site: search operator, which Google has confirmed does not return a complete or accurate page count.

Does submitting a sitemap guarantee indexing?

No — Google's own documentation is explicit on this: a sitemap "doesn't guarantee that all the items in your sitemap will be crawled and indexed." What a sitemap does is tell Google which URLs you consider important and removes the discovery bottleneck for new or orphaned pages. It's a useful signal, not a publishing guarantee. Pages in your sitemap that Google repeatedly ignores usually have a content or technical issue that needs to be resolved first.

Why does my page show "Crawled – currently not indexed"?

This status means Google visited your page but decided not to add it to its index. The most common causes are thin content (the page has too little unique, useful information), near-duplicate content (it is very similar to another page on your site or elsewhere on the web), or a soft 404 (the page returns a 200 status but the content implies "nothing to see here"). Fixing this almost always requires improving the content itself — adding depth, specificity, and genuine usefulness — rather than a technical change.

Get a Clear Picture of What Google Can Actually See

Growth Pulse Media reviews your Google Search Console Page Indexing report, identifies which of your most important pages are not indexed and why, and gives you a plain-language fix list — including any SA-specific server or hosting issues affecting your crawl rate. All work in-house, senior attention from day one. No obligation — we'll get back to you within 24 hours.

Request your free indexing audit

Dirk van Greuning — Founder, Growth Pulse Media
Dirk van Greuning Founder, Growth Pulse Media

Founder of Growth Pulse Media and a specialist in South African search dominance. Dirk translates his experience in scaling South African businesses into high-velocity digital strategies for B2B and retail leaders. He writes about SEO, lead generation, and paid media from an operator's perspective — prioritising pipeline value over impressions.

Connect on LinkedIn