Log file analysis

Log file analysis reads your web server’s access logs to see what crawlers actually fetched: which URLs, how often, with what status, and how fast the server answered. It shows Googlebot’s real crawl, not a simulation. This page shows what to check, how to verify Googlebot, and what to do on a small site.

By , founder of Porteur · Updated 14 September 2026 · Markdown

What log file analysis is

Log file analysis is the review of raw web server access logs to understand crawler behaviour and server responses. You see exact hits, timestamps, requesting IPs, user agents, response codes, and response times.

This exposes wasted crawl on query parameters, chains of redirects, 404s, 500s, and slow responses. It also shows pages Googlebot never fetched, even if they exist in your sitemap.

A crawler audit can guess. Logs tell you what happened. For example, logs might show Googlebot fetching “/pricing?ref=twitter” 200 times a week, but ignoring “/guides/getting-started” entirely.

Why it matters on a small site

You do not have much crawl budget. Any waste on parameters, duplicate paths or broken pages delays useful pages being crawled and updated.

Logs let you prove that important pages are fetched. If “/features” was last crawled months ago while “/search?q=a” is hit daily, you know where to act.

How to get and read the logs

Ask your host, your platform, or your CDN for access logs. A CDN’s logs work too. Export the last 28 days if you can, then keep a rolling window.

Use a log file analyser to parse them. Tools such as the Screaming Frog Log File Analyser ingest common formats and group by user agent, status, URL, and response time. You can also query with your own scripts if you prefer.

  1. Filter to likely crawlers

    Start with user agents containing Googlebot, AdsBot, Bingbot, and known SEO tools. Then verify IPs for Googlebot as above.

  2. Group by URL path

    Bucket by folders: “/”, “/blog/”, “/docs/”, “/api/”. Spot sections that never get fetched or receive heavy crawl.

  3. Check status codes and time

    Count 2xx, 3xx, 4xx, 5xx. Note average response time for each section. Slow sections often map to templates you can fix.

  4. Spot parameters and duplicates

    List URLs with “?” or session IDs. Look for same content at both “/pricing” and “/pricing/” or with case changes.

What to look for and what to do

  • Important pages not crawled: If “/pricing” or “/signup” has zero Googlebot hits in 28 days, add internal links from the homepage and sitemap entries. Check robots.txt and canonicals.
  • Heavy crawl on parameters: If “/search?q=” or “?sort=” dominates, add URL parameter rules, canonical to the clean URL, or disallow low value patterns in robots.txt, then offer HTML links to key facets.
  • Redirect chains: If logs show many 301 then 301 then 200, replace chains with a single hop. Update internal links to the final URL.
  • Error hotspots: If a template returns 404 or 500 often, fix the route or handler. For missing assets, correct the paths in the template.
  • Slow responses: If time to first byte spikes on “/blog/”, cache the page type or optimise the query. A healthy log shows steady 2xx with low response times.
  • Bot traps and waste: If you see frequent hits on “/wp-admin/” on a non WordPress site, block known bad bots by IP or via your CDN rate limits.

A fixed page set looks like this: clean 200s on “/”, “/pricing”, “/guides/getting-started”, few or no parameter hits, and no recurring 4xx or 5xx from templates.

Common traps and limits

  • Partial logs: Many hosts rotate logs fast. Automate export so you do not lose days of data.
  • User agent spoofing: Always verify Googlebot by IP. Do not assume anything with “Googlebot” in the name is real.
  • CDN and origin gaps: If a CDN serves cached hits, your origin log may miss them. Pull CDN request logs as well.
  • Mixing KPIs: Logs show crawl, not rankings or clicks. Use Search Console for impressions and clicks, then map to crawl changes.
  • Over-blocking: A quick robots.txt disallow for parameters can hide useful filters. Prefer canonicals and crawl hints before hard blocks.

Questions

Sources

Check my site, free

Drop your URL to see where crawl is likely wasted and which sections look thin, then use Porteur’s free check to compare that with rivals in about thirty seconds.

  • Free check, no card
  • Read-only, your own accounts
  • Readable by your agent

Read next