An SEO log file analysis guide is a systematic examination of your web server's access logs to see exactly which URLs Googlebot requests, how frequently, and what response codes it receives — giving you bot-behaviour data that no crawler simulation can replicate. If you are building sustainable search visibility in South Africa, your server logs are the one data source that shows you real Googlebot activity rather than an approximation of it. A page that never appears in those logs was never crawled by Googlebot, which means it cannot be indexed.

For most South African WordPress and WooCommerce sites with stable content and under 1,000 pages, crawl budget is rarely the binding constraint. But once your site generates significant URL volume through product filters, pagination, or blog archives, access logs start revealing problems that Google Search Console alone will not surface: intermittent server errors, crawl waste on low-value parameter URLs, and high-priority pages that Googlebot visits far less often than session-ID duplicates. This guide walks through how to get your logs, what to look for, and what to do with what you find.

Quick Answer

A log file analysis for SEO means downloading your web server's raw access logs, filtering for Googlebot and other search-engine crawlers, then auditing which URLs were crawled, how often, and with what HTTP response code. The goal is to identify crawl waste — resources spent on low-value URLs — and crawl gaps, where important pages are not being requested at all.

As a working rule of thumb, run at least 30 days of log data; 90 days gives you enough sample to identify crawl frequency patterns. The free tier of Screaming Frog's Log File Analyser handles up to 1,000 log events — enough to audit most small SA sites without spending anything.

Not Sure Where Your Crawl Budget Is Going?

Share your site URL and we will run a crawl audit to show you which pages Googlebot is ignoring and which low-value URLs are consuming your allocation.

Request a Crawl Audit

What Log Files Reveal That No Other Tool Can

Server access logs give you four data points no other SEO tool can replicate at the source: the exact URLs Googlebot requested, the precise timestamp of each visit, the HTTP response code returned, and the crawl frequency of each URL relative to the rest of the site. Your web server writes a new line to its access log for each request from any client — a human browser, Googlebot, a social media scraper, or an AI training crawler. That raw record contains the requesting IP address, the timestamp, the URL path, the HTTP method, the response code, the bytes transferred, and the user-agent string. Treat the user-agent as a claim, not a fact: verify Googlebot with forward-confirmed reverse DNS (reverse-look up the IP, then forward-resolve the hostname and confirm it returns the same IP) or against Google's published IP ranges. No tool you install on top of your site can produce this data retrospectively; it is created at the server layer, before any JavaScript runs.

Third-party crawl tools (Screaming Frog's main crawler, Ahrefs Site Audit, Semrush) simulate what a crawler would find if it followed links through your site. They are useful for finding broken links and missing metadata. But they do not show you how Googlebot actually behaves: how often it visits each URL, whether it returned a clean 200 or hit a 503 during load, and whether it is spending the majority of its requests on your core product pages or on filter URLs that should never have been crawlable in the first place.

HTTP Response Codes in Your Logs: Quick Reference

200 — Page served successfully. Googlebot has the content.

301 — Permanent redirect. Googlebot follows it; crawl budget is spent on two requests.

302 — Temporary redirect. Googlebot follows it but may revisit the original URL repeatedly.

404 — Not found. Googlebot returned empty-handed; repeated 404 requests waste your crawl allocation.

500 / 503 — Server error or temporarily unavailable. Googlebot backs off; prolonged errors signal unreliability and reduce crawl frequency.

The table below shows at a glance when log file analysis is load-bearing versus optional for your site:

Site sizeLog analysis priorityKey question to answer
Under 1,000 pagesOptional — run when debugging a crawl problemAre my most important pages being crawled at all?
1,000–10,000 pagesQuarterly review recommendedAm I wasting crawl on parameter, filter, or session-ID URLs?
Over 10,000 pagesMonthly — treat as a standard audit taskWhere is crawl allocation going, and which strategic pages are being skipped?

Google's guidance identifies two rough thresholds where crawl budget becomes a meaningful factor: sites with 1 million or more pages whose content changes at least weekly, and sites with 10,000 or more pages that update very frequently — daily or near-daily. These are estimates, not exact cut-offs. For any site where Search Console shows a growing proportion of URLs in a "Discovered — currently not indexed" state, log analysis is worth running regardless of page count. For sites that fall below these thresholds with stable, low-volume content, keeping an accurate XML sitemap and monitoring the Coverage report is typically sufficient.

Getting Your Server Logs: SA Hosting Environments

Downloading your access logs is a one-time configuration step that most SA hosting control panels make straightforward. The exact path depends on your stack.

cPanel (Afrihost, WebAfrica, Hetzner SA shared hosting): Log in to cPanel, go to Logs & Analytics, click Raw Access, then download the gzipped (.gz) log file for your domain. The file covers the current calendar month; older months are archived below. Decompress with any standard archive tool before importing into your analysis tool.

Plesk (common on VPS plans): Navigate to Websites & Domains, select your domain, click Logs, and download the access log. Plesk records Apache error logs and NGINX proxy access logs separately — for crawl analysis, you want the access log (proxy_access_log on NGINX configurations).

Direct server access (SSH): Apache logs are typically at /var/log/apache2/access.log or /var/log/httpd/access_log. NGINX logs are usually at /var/log/nginx/access.log. Use scp or an SFTP client to pull the file to your local machine before analysis.

Shopify: Shopify does not expose raw server-level access logs to merchants. If your SA store runs on Shopify, you cannot run a traditional server log file analysis. Your alternatives are Google Search Console's crawl data (under Coverage and URL Inspection) and any CDN-level analytics your Shopify plan or Cloudflare setup exposes. This limitation is worth knowing before you spend time searching for a log file that does not exist in your hosting panel.

Key Takeaway

Collect at least 30 days of log data before analysis — as a practical heuristic — so that infrequently crawled pages have a chance to appear. Ninety days gives you a pattern rather than a snapshot. A single week of logs will show you errors but will not reveal which important pages Googlebot visits infrequently compared to low-value ones.

The Five-Step SEO Log File Analysis Guide

A structured log file analysis for SEO follows a consistent sequence: import the data, isolate the relevant crawlers, profile response codes, cross-reference against what should be crawled, and then identify the specific waste or gap driving the problem. Each step builds on the previous one.

Step 1 — Import and parse your log file. Screaming Frog's Log File Analyser (free up to 1,000 log events; £99 per year for unlimited) accepts Apache, W3C Extended (IIS and NGINX), and Amazon Elastic Load Balancing log formats. Drag the downloaded log file onto the interface and the tool parses it automatically, grouping events by URL, response code, and bot type. It runs on Windows, macOS, and Linux. For larger log files or server-side analysis, GoAccess is a free command-line alternative that generates HTML reports.

Step 2 — Filter for Googlebot only. Your log contains requests from human browsers, marketing tools, uptime monitors, and dozens of other crawlers. Filter the user-agent field to isolate Googlebot (Googlebot Desktop and Googlebot Smartphone are separate user-agents and may behave differently on the same URLs). Remove all other user-agents from this pass — you are looking at what Google's crawler specifically does, not the aggregate.

Step 3 — Audit the response code distribution. With Googlebot filtered, check the breakdown of 200s, 3XXs, 4XXs and 5XXs. A healthy site has a high proportion of 200 responses. Repeated 404s on the same returning URLs are either deleted pages that need redirecting or parameter variants to block via robots.txt. A pattern of 503s at specific time windows is worth cross-referencing against server load events — on SA hosting, this often correlates with load-shedding restarts that briefly degraded server availability. If server logs sit outside your team's comfort zone, our technical SEO support in South Africa starts with a free audit and continues on a flat monthly retainer.

Step 4 — Cross-reference crawled URLs against your sitemap. Export the list of Googlebot-crawled URLs from your log tool. Import your XML sitemap URLs as a second list. Anything in your sitemap that does not appear in 30 to 90 days of logs is a crawl gap — Googlebot either has not discovered it yet or is choosing not to return. Anything Googlebot is crawling heavily that is not in your sitemap and adds no ranking value is crawl waste. The gap between these two lists is your action list.

Step 5 — Analyse crawl frequency by page type. Group your Googlebot-crawled URLs by type: homepage, category pages, product or service pages, blog posts, pagination pages, filtered URLs. If your blog-post archive pages are receiving more Googlebot visits per month than your core service pages, crawl budget is being spent on lower-priority content. This analysis identifies where internal linking improvements, robots.txt disallow rules, or canonical tags need to be applied.

Key Takeaway

The most common finding in a first log file audit is not a technical error — it is that pagination and filter URLs are consuming a disproportionate share of crawl. If your category pages generate hundreds of ?sort= and ?filter= URL variants, those variants are competing with your core pages for Googlebot's attention. Disallowing them in robots.txt is usually the highest-return first action. Note that noindex does not save crawl budget — Google must fetch the page to see the tag — so use it for index control, and robots.txt for crawl control; they are not interchangeable.

Turning Findings Into Fixes

The most common crawl log analysis SEO findings each map to a specific fix: 404 loops resolved by 301 redirects or 410 responses, crawl-waste URLs blocked via robots.txt or canonical tags, and orphaned priority pages recovered through internal linking. The table below works through each pattern and its standard remediation.

What the logs showWhat it meansFix
Googlebot hitting 404s on deleted pages repeatedlyGooglebot remembers old URLs; crawler keeps returning301-redirect to a relevant live page, or return 410 Gone to signal permanent removal
Key service/product pages crawled infrequentlyCrawl allocation going elsewhere; internal link weight too lowIncrease internal links from high-crawl-frequency pages to priority pages; check for orphan pages
Filter and parameter URLs crawled heavilyCrawl waste on low-value duplicate contentDisallow parameter patterns in robots.txt or add rel=canonical pointing to the clean URL
503 errors in specific time windowsServer unavailability reducing crawl capacityInvestigate hosting infrastructure; move to a more reliable hosting tier; check for cron jobs causing resource spikes
Sitemap URLs with zero crawls in 90 daysCrawl gap on important pagesCheck for blocking robots.txt rules, missing internal links, or canonical conflicts; review with crawlability best practices
Redirect chains (3XX to 3XX)Crawl budget spent on unnecessary hops; link equity dilutedUpdate source to point directly to the final destination URL

Prioritise fixes using the same logic as any other technical work: pages with commercial intent that are currently uncrawled or crawled infrequently get fixed first. Session-ID cleanup and redirect consolidation come next. Low-traffic blog archive cleanup is last. If you are unsure how to stack your fixes, the process in how to prioritise SEO fixes applies directly here.

One fix worth calling out: a redirect from a deleted page to your homepage is treated by Google as a soft 404. A clean 404 or 410 is better than a misleading homepage redirect — a common trap in SA WooCommerce stores where discontinued products redirect to the homepage by default.

AI Crawlers Are in Your Logs Too

Alongside Googlebot, your access logs now contain requests from AI platform crawlers — and the pattern of those visits tells you whether your content is being accessed for training and real-time retrieval by systems like ChatGPT, Claude, and Perplexity. This is no longer an optional insight if you are trying to appear in AI-generated answers as well as standard search results.

The key AI crawler user-agents to filter for are: GPTBot (OpenAI model training), OAI-SearchBot (OpenAI real-time search indexing), ChatGPT-User (real-time retrieval when a user asks ChatGPT a question), ClaudeBot (Anthropic training), Claude-User (Anthropic real-time), PerplexityBot (Perplexity indexing), Perplexity-User (Perplexity real-time), and CCBot (Common Crawl). Screaming Frog's Log File Analyser includes presets for all of these, so you do not need to write custom filters manually.

The analysis logic is identical to the Googlebot process: which pages are these crawlers requesting, and which important pages are absent from the logs entirely? Absence from your log sample is a prompt to check, not proof of invisibility: GPTBot is OpenAI's training crawler, while OAI-SearchBot serves ChatGPT search, and a finite log window cannot show that a page was never retrieved.

If your detailed how-to articles or FAQ pages are not appearing in AI crawler logs, those are the pages to check for crawl blocks and to strengthen with additional internal linking. For a deeper look at how indexation connects to crawl frequency, see the guide on improving website indexation.

Key Takeaway

AI platform crawlers operate independently of Googlebot and have their own crawl priorities. A page blocked from Googlebot via robots.txt (using a disallow rule targeting only Googlebot) can still be accessed by GPTBot and PerplexityBot unless you add specific directives for those user-agents. If you want a page excluded from AI training, each crawler needs its own disallow rule — and your log files are how you verify those rules are working as intended.

Want to Know Which AI Crawlers Are Visiting Your Site?

Send us your domain and we will pull a log-based bot audit showing which AI platforms are crawling your content — and which strategic pages they are missing.

Request a Bot Audit

Why South African Businesses Choose Growth Pulse Media

Technical work like a server log file SEO audit sits at the intersection of server infrastructure and organic search strategy — not something that benefits from delegation to a junior account manager. Dirk van Greuning built and scaled a large South African ecommerce business before founding Growth Pulse Media, which means the crawl patterns he diagnoses are ones he encountered firsthand on SA hosting infrastructure: cPanel plans from local providers, NGINX configurations on VPS accounts, and the crawl behaviours specific to SA content architectures.

Our SEO work in South Africa covers the full technical stack: log file audits, crawl budget analysis, internal linking strategy, and the content layer that determines which pages Googlebot prioritises. All work is executed in-house, with a deliberately limited client load so that senior attention is not rationed.

If your site has reached a scale where crawl efficiency is a real lever — or you are seeing "Discovered — currently not indexed" growing in Search Console — that is the right time to apply this seo log file analysis guide to your specific site. For related technical work, see our guides on improving crawlability and crawl depth explained.

Who This Is NOT For

Sites with stable content and under 1,000 pages. Googlebot's crawl of a site this size is rarely the constraint on what gets indexed. Your time is better spent on content quality, internal linking, and backlink acquisition than on log file parsing.

Shopify merchants expecting server-level logs. Shopify does not give merchants access to raw access logs. If you are looking for crawl data on a Shopify store, you need to work with Google Search Console's URL Inspection tool and Coverage report rather than the server log analysis process described here. The methodology is different.

Sites that have not yet implemented a solid technical foundation. Log file analysis is a diagnostic tool, not a foundation-layer fix. If your site still has uncanonicalized duplicate content, missing meta tags, or no XML sitemap, resolve those issues first. Running a log audit before fixing fundamentals is like measuring water pressure before you have connected the pipes.

Businesses expecting log analysis to improve rankings directly. Log file analysis identifies problems and opportunities — it does not fix them. The insight that Googlebot is ignoring your core category pages does nothing unless you follow through with the internal linking and crawl-waste cleanup that changes the behaviour. If you are looking for a tool that automatically improves rankings, this is not it.

Not Sure Which Technical SEO Issues to Tackle First?

Tell us what you are seeing in Search Console and we will assess whether a log file audit is the right next step or whether a different technical fix would move the needle faster.

Get a Technical Assessment

Frequently Asked Questions

What is a log file analysis for SEO?

A log file analysis for SEO means examining your web server's access logs to see exactly which URLs search-engine crawlers request, how frequently they visit, and what HTTP response codes they receive. Unlike simulated crawl tools, server logs record real bot behaviour — giving you accurate data on crawl frequency, crawl waste, and pages that Googlebot is not visiting at all. The output is a prioritised list of crawl problems and gaps that drives specific fixes to improve how efficiently Googlebot discovers and indexes your important pages.

How do I get my server log files?

The method depends on your hosting setup. On cPanel (common with SA providers like Afrihost and WebAfrica), download the Raw Access file from the Logs section of your control panel; on Plesk, navigate to Websites and Domains and click Logs for your domain. If you have direct SSH access, find your Apache or NGINX access log in your server's log directory. Shopify does not expose raw server logs — Shopify merchants should use Google Search Console's Coverage report for crawl data instead.

How long should my log file sample cover?

As a practical working rule, collect at least 30 days of log data before running an analysis — shorter samples can miss infrequently crawled pages entirely. Ninety days gives you enough data to identify patterns in crawl frequency across different page types, not just a snapshot of which pages Googlebot requested in a single month. If you are investigating a specific problem (like a sudden drop in crawl frequency after a site migration), a shorter targeted window can be useful alongside the longer baseline.

What is a Googlebot log file analysis, and does it differ from checking Search Console?

A Googlebot log file analysis filters your server logs to show only requests made by Googlebot's user-agent, revealing exactly which pages it crawled, when, and with what result. Google Search Console's Coverage and URL Inspection tools show you indexation status — whether Google knows about a page and has chosen to index it — but they do not show raw crawl frequency or response code history the way logs do. Logs and Search Console answer different questions; used together, they give a complete picture of how Googlebot interacts with your site.

How does log analysis connect to crawl budget management?

Google defines crawl budget as the set of URLs it can and wants to crawl on your site, shaped by server capacity and Google's assessment of content value and freshness. Log file analysis is the only way to measure where that budget is actually going — which page types consume the most crawl requests, and which important pages are visited infrequently. Without log data, crawl budget management is guesswork; with it, you make targeted changes and verify in the next log sample whether Googlebot's behaviour has shifted.

Let Us Audit Your Site's Crawl Efficiency

Growth Pulse Media runs server log audits for South African businesses as part of our technical SEO work — covering Googlebot crawl patterns, AI crawler access, crawl waste identification, and a prioritised fix list. All work is executed in-house, with senior attention on every engagement. No obligation — we will get back to you within 24 hours.

Request Your Crawl Audit
Dirk van Greuning — Founder, Growth Pulse Media
Dirk van Greuning Founder, Growth Pulse Media

Founder of Growth Pulse Media and a specialist in South African search dominance. Dirk translates his experience in scaling South African businesses into high-velocity digital strategies for B2B and retail leaders. He writes about SEO, lead generation, and paid media from an operator's perspective — prioritising pipeline value over impressions.

Connect on LinkedIn