An seo site architecture checklist is a structured audit framework for organising a website's URLs, navigation, internal links, and crawl controls into a logical hierarchy that search engines can discover and rank efficiently. Architecture is one of the most leveraged technical foundations in SEO for South African businesses — fix the structure once and every page on your site benefits immediately and long-term. If your pages compete with each other, disappear from the index, or take months to rank despite solid content, the problem is often structural, not the content itself.

This checklist covers the six layers that determine how Googlebot moves through your site, how authority flows between pages, and how clearly your topic clusters signal their relevance. It is written for existing websites that need an audit, not sites being built from scratch. Start with the layer that matches your biggest current gap, not layer one.

Quick Answer

An seo site architecture checklist audits six layers: URL structure, navigation depth, internal linking, crawlability, canonical and duplicate content controls, and mobile-first performance. Use it to find where Google is wasting crawl capacity, where authority is leaking from key pages, and where pages sit too deep to be indexed regularly. For most South African websites under 1,000 pages, the highest-impact fixes are navigation depth and internal link gaps — not crawl budget tuning, which only becomes meaningful at scale.

Is your site structure holding your rankings back?

Send us your URL and we will identify the top structural gaps in a free crawl depth review — no obligation, response within 24 hours.

Get a Free Structure Review

Why Your Structure Determines Your Rankings

Site structure is the map Googlebot uses to discover, crawl, and assign relative importance to every page on your domain — and a poor map means critical service or category pages sit four clicks deep, get revisited infrequently, and rank below their content quality. Three factors drive this. First, crawl depth: practitioners broadly observe that pages further from the homepage in click-terms receive crawl capacity less frequently — they sit further from the authority that flows down from a well-linked root. Second, link equity flow: internal links pass authority between pages; a page with no internal links pointing to it receives almost none, regardless of its backlink profile. Third, topical coherence: your SEO site structure — the combination of URL paths, navigation levels, and internal link clusters — signals to search engines how your topic areas relate, and Google's Search Essentials guidance notes that logical directory organisation helps search engines understand how frequently different sections of a site update, which influences how often they crawl those sections.

For South African businesses, this matters at every scale. A seven-page professional services site can suffer from URL structure problems and zero internal linking. A 500-page ecommerce store can inadvertently generate thousands of duplicate filter URLs that consume crawl capacity. The fix is always the same — audit the six layers below and work from your biggest gap.

The SEO Site Architecture Checklist: Six Layers to Audit

A complete seo site architecture checklist covers six distinct layers, each addressing a different factor that practitioners identify as influencing how search engines discover, crawl, and rank pages. Think of it as a website architecture checklist that works through each audit area in order of impact rather than alphabetical sequence. Treating all layers as equally urgent is the fastest route to wasted effort. The table below maps each layer to what it controls and the primary mechanism it affects.

LayerWhat It ControlsPrimary Google Mechanism
1. URL StructureHow URLs are formed, readable, and uniqueCrawl efficiency and duplicate content avoidance
2. Navigation and click depthClicks from homepage to key pagesCrawl frequency and authority flow to top pages
3. Internal linkingWhich pages link to which, and with what anchor textPageRank distribution and topical relevance signals
4. Crawlability and robots controlsWhich pages Googlebot can and cannot accessCrawl capacity allocation
5. Canonical and duplicate contentWhich URL is the authoritative version of each pageIndex consolidation and authority concentration
6. Mobile-first and page speedHow the site performs for mobile usersMobile-first indexing and user experience signals

Layer 1: URL Structure

A clean URL structure makes every page's topic readable to both users and crawlers, reduces the risk of duplicate content, and helps Googlebot understand how pages relate to each other. According to Google's URL structure guidance, hyphens are preferred over underscores as word separators, and URLs should use descriptive words rather than numeric IDs or query parameters wherever possible.

  • Use hyphens — not underscores — to separate words in all URLs
  • Keep URLs short, lowercase, and descriptive of the page's actual content
  • Avoid dynamic parameters where static slugs are possible (e.g. /products/red-running-shoes/ rather than /products?id=847&colour=red)
  • Group related pages into logical subdirectories — for example /services/, /blog/, /resources/ — so directory structure mirrors topic grouping
  • Ensure each page has exactly one canonical URL — resolve any trailing slash variants, www vs non-www splits, and HTTP vs HTTPS inconsistencies via 301 redirects
  • Never use URL fragments (#anchors) to load different page content — Google does not index fragment-based content variations
  • Confirm all live URLs return 200 status; fix or redirect any 404s that still receive internal links from your own pages

Click depth — the number of clicks required to reach a page from the homepage — affects how often Googlebot revisits that page and how much internal link authority flows to it. Practitioners broadly agree that every revenue-driving page should be reachable within three clicks; pages buried four or five clicks deep are crawled significantly less often and receive a smaller share of your site's internal authority.

  • Map your current hierarchy — crawl your site with a tool such as Screaming Frog or Sitebulb and export click depth per URL
  • Identify every page sitting four or more clicks from the homepage
  • Prioritise revenue pages (service pages, category pages, product pages) — move them to three clicks or fewer from the homepage
  • Ensure your main navigation covers all primary topic areas — categories missing from the nav are often also missing from the index
  • Add breadcrumbs on all pages below the first navigation tier — they create shortcut paths for Googlebot and improve usability on mobile
  • Audit orphan pages with no inbound internal links — a page Googlebot cannot reach via links is effectively invisible, regardless of how it appears in a sitemap
  • For larger ecommerce stores: confirm paginated listing pages (page 2, page 3 of a category) are reachable from the category navigation, not only via pagination controls

Layer 3: Internal Linking

Internal links are the primary mechanism through which search engines understand your site's topic relationships and distribute link authority across pages. Evaluating site architecture for SEO purposes starts here for most websites that have been publishing for more than six months, because internal linking problems accumulate quietly while content grows. A well-linked page signals importance; a page with few or no internal links pointing to it — regardless of content quality — ranks as though it exists in isolation.

  • Identify your ten most important pages (service pages, category pages, pillar content) — confirm each has multiple inbound internal links from relevant supporting pages
  • Use descriptive anchor text that reflects the target page's topic — avoid generic anchors such as "click here" or "read more"
  • Implement the pillar-cluster model: one comprehensive pillar page per topic (linking down to sub-topics) and several focused supporting pages (each linking back up to the pillar)
  • Find orphan pages using your crawl tool — filter for pages with zero inbound internal links — then add contextual links from relevant content
  • Audit existing anchor text on links pointing to your most important pages — if most say "here" or "this article", Googlebot is not receiving topical context from those links
  • Remove the nofollow attribute from internal links — nofollow stops authority passing to your own pages
  • Avoid closed link loops where a small group of pages only link to each other — these trap link equity in a circuit with no outbound flow to the rest of your site

Key Takeaway: Orphan Pages

An orphan page has content but no internal links pointing to it — search engines can find it only via a sitemap or external links, not by following your site's own structure. Our guide on how to find orphan pages walks through how to identify and resolve them systematically. For sites that have been publishing for more than a year, fixing orphan pages is typically the fastest internal linking improvement available.

Layer 4: Crawlability and Robots Controls

The technical site architecture controls in this layer determine which pages Googlebot is allowed to access, which directly affects how efficiently your crawl capacity is spent on pages that need to be indexed. Blocking the wrong pages, or leaving a large volume of low-value parameterised URLs unblocked, wastes the crawl access that should be reaching your content pages. When robots rules and site structure need a careful rework, our founder-led technical SEO starts with a free audit so nothing changes until the priorities are clear.

When crawl budget actually matters: Google's guidance on crawl budget targets sites with over one million pages updating weekly, or sites with over 10,000 pages changing daily. If your site has fewer than 1,000 pages, crawl budget is rarely the constraint — click depth and internal links are. See how to improve website indexation for the full diagnostic.

  • Review your robots.txt file — confirm you are not inadvertently blocking category pages, product pages, CSS, or JavaScript that Google needs to render your pages correctly
  • Understand the distinction: robots.txt blocks crawling, not indexing — a blocked URL can still appear in search results if other sites link to it; use noindex for true search removal
  • Identify parameter-generated URLs (filtering, sorting, search results, session IDs) — these multiply page count without adding unique content; block or canonicalise them
  • For WooCommerce and Shopify stores with faceted navigation: prevent crawling of filter-generated URLs that do not correspond to pages you want appearing in search results
  • Remove or noindex thin pages: tag archives, author archives, internal search result pages, and empty category pages
  • Verify your XML sitemap lists only canonical, indexable URLs — no redirects, no noindex pages, no URLs returning 4xx errors
  • Submit your sitemap in Google Search Console and monitor the Crawl Stats report for error spikes and unusually high server response times

Layer 5: Canonical Tags and Duplicate Content

Canonical tags tell Google which URL is the authoritative version of a page, consolidating ranking signals from duplicate or near-duplicate URLs onto a single strong page. Without them, a single product accessible via three different filter combinations splits its authority three ways instead of ranking as one well-supported page.

  • Confirm every page has a self-referencing canonical tag pointing to its own clean URL
  • For pages with URL parameter variants (sort order, session IDs, UTM tags), add canonical tags on all variants pointing to the clean URL
  • For paginated series (page 2, page 3 of a category listing), do not point all pages back to page 1 — each paginated URL should self-canonicalise
  • Treat canonical tags as strong hints, not absolute directives — Google may override a canonical if other signals (internal links, backlinks, content similarity) point strongly to a different URL
  • Check for HTTP vs HTTPS and www vs non-www duplication — each pair needs to resolve to one canonical destination via a 301 redirect, not a canonical tag alone
  • Do not canonicalise a page to a URL that itself redirects — canonical tags should point to the final live destination

Key Takeaway: Canonical vs Redirect

A canonical tag consolidates ranking signals without removing the duplicate URL from your server. A 301 redirect removes the old URL entirely. Use a redirect when you want the old URL gone permanently; use a canonical when the URL must remain accessible (for session handling, analytics, or platform constraints) but you want ranking signals consolidated. For content that duplicates across multiple URLs because of keyword targeting, see our guide on fixing keyword cannibalisation.

Layer 6: Mobile-First and Page Speed

Google indexes the mobile version of your site first, meaning a page's mobile rendering determines whether it appears in the index, how it ranks, and how Google reads its structured data. In South Africa, DataReportal's Digital 2026 report shows 98.7% of mobile connections running on broadband networks (3G, 4G, or 5G) — the practical issue for most SA businesses is not connectivity, but page weight on mid-range Android devices where 4G speeds vary by location and can be further compressed during load-shedding when users switch from fibre to mobile data.

  • Confirm your site is fully mobile-responsive — test every major template (homepage, category, product, blog post) on a physical mid-range Android device, not only a browser emulator
  • Verify your mobile HTML contains the same content and internal links as your desktop HTML — navigation links hidden on mobile may prevent Googlebot from discovering their destination pages
  • Review Core Web Vitals via Google Search Console's Core Web Vitals report — field data reflects how real users on South African networks actually experience your pages
  • Use Chrome DevTools (Performance tab) to diagnose load time issues — the standalone Web Vitals Chrome extension was retired in January 2025
  • Compress and appropriately size images — unoptimised images are the most common page speed issue on SA ecommerce sites running on mid-range devices
  • Minimise render-blocking JavaScript — if your navigation or page content requires significant JavaScript to render, Googlebot may not see your internal links until after JS execution completes
  • Test page load on a throttled 4G connection — this simulates the experience of a South African user who has switched from fibre to mobile data during load-shedding

Not sure which layer to tackle first?

Tell us your site's scale and where you are seeing gaps — we will map your structure against this checklist in a free 30-minute assessment and confirm what to fix in what order. No obligation, response within 24 hours.

Book a Free Architecture Assessment

How to Prioritise What to Fix First

Fixing all six layers simultaneously is rarely possible and rarely necessary. The table below maps site scale to the highest-impact architecture fix at each tier. These are practitioner-recommended priorities based on where structure problems cause the most ranking damage at each site size — not Google-confirmed ranking weights.

Site sizeFix this firstThenThen
Fewer than 100 pagesNavigation depth — ensure every page is reachable within 3 clicks from the homepageURL structure cleanup and consistent canonical setupInternal linking between topic-related pages
100–1,000 pagesInternal link gaps and orphan pages — find all pages with zero inbound internal linksCanonical and duplicate URL controls, especially for ecommerce filter URLsSitemap audit and Search Console monitoring
1,000+ pagesCrawl waste — thin pages, parameter URLs, and faceted navigation generating index bloatIndex audit — noindex or remove pages that add no search valueLink equity flow to revenue-generating pages

For a broader framework for triaging all your organic search tasks alongside architecture, see our guide on how to prioritise SEO fixes. For the detailed technical implementation steps for each crawlability item above, see how to improve crawlability.

Key Takeaway: Crawl Budget Is Not a Small-Site Problem

Google's guidance states that crawl budget becomes a meaningful concern for sites with over 10,000 pages that change daily, or sites with over one million pages. If your site has fewer than 1,000 pages, crawl budget is almost certainly not the constraint — click depth and internal linking are. Spending hours configuring robots.txt on a 200-page site is the wrong priority. Fix navigation depth and link orphans first.

Why South African Businesses Choose Growth Pulse Media

Fixing site structure requires someone who has worked at every layer — from writing robots.txt directives that do not break navigation, to restructuring internal links across hundreds of pages without losing existing rankings. Growth Pulse Media approaches this from an operator's perspective: founder Dirk van Greuning built and scaled a large South African ecommerce business before founding the agency, which means every architecture decision is informed by real commercial consequences, not theory.

Our approach to SEO for South African businesses treats site architecture as the foundation, not an afterthought. All execution is in-house — no junior teams running checklists — and we maintain a limited client load to ensure the depth of attention this kind of technical work demands. If your internal links are pointing to the wrong pages, your faceted navigation is generating thousands of thin URLs, or your most important pages are sitting five clicks deep, we will show you exactly what to fix and in what order.

Who This Is NOT For

Marketplace-only sellers. If your entire business runs on Takealot, Checkers Sixty60, or a similar marketplace and you have no standalone website, structure decisions are made by the platform. There is nothing here to audit.

Businesses in the middle of a full rebuild. If your site is being redeveloped and going live in the next three months, architecture decisions belong in the project brief, not a post-launch audit. The right time to implement this checklist is before the build, not after it goes live with the wrong structure already baked in.

Teams with no developer or CMS access. Several items in this checklist — robots.txt edits, canonical tag implementation, sitemap configuration, redirect setup — require access to server config or CMS templates. If you have marketing budget but no way to implement technical changes, fix the access problem before running this audit.

Websites with fewer than five pages. A homepage, two service pages, an about page, and a contact page have no meaningful structural problem. Build content to at least 20 pages before site architecture becomes a lever worth pulling.

Ready to get your site structure audited?

Book a free SEO audit and receive a prioritised architecture fix list specific to your site's current state — not a generic template. No obligation, response within 24 hours.

Book Your Free SEO Audit

Frequently Asked Questions

What is an SEO site architecture checklist?

An SEO site architecture checklist is a structured audit framework that evaluates how a website's URLs, navigation, internal links, crawl controls, and canonical tags are organised. It identifies where Googlebot is wasting crawl capacity, where link authority is leaking from important pages, and where pages sit too deep in a hierarchy to be indexed and ranked effectively. A complete checklist covers six layers: URL structure, navigation and click depth, internal linking, crawlability, canonical and duplicate content, and mobile-first performance.

How does site architecture affect SEO?

Site architecture affects organic search performance through three mechanisms: crawl depth (pages buried deep in the hierarchy are crawled less often), link equity flow (internal links distribute authority — pages with no internal links pointing to them struggle to rank regardless of content), and topical coherence (logical grouping helps search engines understand what a site covers). Google's own guidance notes that organising content into logical directory groups helps search engines understand update frequency across different parts of a site, influencing how often those sections are crawled.

How many clicks deep should pages be for SEO?

Practitioners broadly recommend keeping every revenue-driving page reachable within three clicks from the homepage. Pages at four or more clicks deep are crawled less frequently and receive less internal link authority than shallower pages. This is practitioner consensus rather than a confirmed Google threshold — but it aligns with Google's guidance on ensuring important pages are well-linked and discoverable through site navigation rather than relying on sitemaps alone.

Does site architecture matter for small South African websites?

Yes, but the priority fixes differ by site size. A small site under 100 pages does not need crawl budget tuning — Google's guidance reserves that concern for sites with 10,000 or more pages. What small SA sites most commonly get wrong is navigation depth (key pages buried in submenus or accessible only through footer links) and zero internal linking between related content. Fixing these two issues on a small site typically delivers faster ranking improvements than any additional content publication.

How do I find orphan pages on my website?

Orphan pages have no internal links pointing to them — Googlebot can find them only via a sitemap or external links, not by following your site's own structure. Use a crawl tool such as Screaming Frog, Sitebulb, or Ahrefs Site Audit to export all internal links, then compare that list against all indexed URLs in Google Search Console. Any indexed URL that receives zero internal links is an orphan. Our guide on how to find orphan pages walks through the full process and shows how to prioritise which ones to fix first.

Fix Your Structure Before You Publish Another Page

Every page added to a poorly structured website inherits the same architectural problems. Growth Pulse Media audits South African business websites against this exact checklist — identifying click depth issues, orphan pages, index bloat, and canonical failures — and delivers a prioritised fix plan built around your specific CMS and resource constraints. All work is executed in-house by senior practitioners. No obligation — we will get back to you within 24 hours.

Book Your Free SEO Audit
Dirk van Greuning — Founder, Growth Pulse Media
Dirk van Greuning Founder, Growth Pulse Media

Founder of Growth Pulse Media and a specialist in South African search dominance. Dirk translates his experience in scaling South African businesses into high-velocity digital strategies for B2B and retail leaders. He writes about SEO, lead generation, and paid media from an operator's perspective — prioritising pipeline value over impressions.

Connect on LinkedIn