# Caching: what to cache, for how long, and what it does for crawling

Caching decides how much of your site has to be built again for every visitor and every crawler. Two rules cover almost all of it: assets whose name changes when their content changes are cached for a year, and HTML is cached briefly and revalidated.

Updated 2026-09-15 · Source: https://porteur.ai/guides/caching-for-seo

## Two caches, two jobs

| Cache | Who it serves | What it saves |
| --- | --- | --- |
| The browser cache | One returning visitor | The whole request. Nothing travels at all |
| The CDN or edge cache | Everyone, including the first visitor and crawlers | The journey to your origin and the work the origin would do |
| The origin's own cache | Your server | The rendering or the database work behind a page |

Search cares mostly about the middle row. A crawler is usually a first-time visitor from somewhere far away, so a warm edge cache is what it experiences, and the time to first byte it measures is the one your visitors get too.

## Static assets: hash the name, cache for a year

Any file whose name contains a hash of its content can be cached as long as you like, because a change produces a new name and therefore a new URL.

```text
# /_next/static/chunks/main.6f3c9a2b.js, /assets/logo.4b21e0.svg
Cache-Control: public, max-age=31536000, immutable
```

- A year is the conventional maximum, and immutable tells the browser not to revalidate even on a reload.
- This applies to hashed JavaScript, CSS, fonts and images produced by your build.
- It does not apply to files with stable names, such as /logo.svg or /favicon.ico. Those get a shorter cache.
- Most frameworks and platforms set this correctly by default. Check rather than assume.

> If your assets are not hashed, caching them for a year means visitors keep a stale stylesheet after every deploy. Hash the filenames first, then cache hard.

## HTML: cache briefly, revalidate always

HTML has one URL forever, so it cannot be cached like an asset. The aim is to avoid rebuilding it for every request without serving yesterday's page after you publish.

```text
# A page that changes occasionally
Cache-Control: public, max-age=0, must-revalidate
ETag: "a1b2c3"

# The same page, cached at the edge and revalidated in the background
Cache-Control: public, s-maxage=600, stale-while-revalidate=86400
```

- max-age governs the browser, s-maxage governs the CDN. They can differ, and usually should.
- An ETag or a Last-Modified header lets a cache ask "has this changed?" and receive a small 304 instead of the whole page.
- must-revalidate means the cached copy may be used only after checking, which is what you want for pages that matter.
- Never send no-store for public pages. It forbids caching anywhere and every visitor pays full price.

## stale-while-revalidate: the setting that makes a small site feel static

With stale-while-revalidate the edge serves the cached page immediately and fetches a fresh copy in the background. The visitor who triggers the refresh never waits for it.

- The first visitor after the window gets the old page instantly, and the next one gets the new page.
- Your origin handles one request per window per edge location rather than one per visitor.
- For a marketing site or a library of guides, a window of minutes with a long stale period is usually right.
- Pair it with a purge on deploy so a correction is live immediately rather than at the end of the window.

This is the difference between a server-rendered site that feels slow and one that feels like static files, without turning it into static files.

## What a CDN actually removes

A CDN in front of the origin removes most of the time to first byte for visitors far from your server, because the response comes from a machine near them rather than from one continent away.

- The distance: the round trips of the handshake happen close to the visitor.
- The origin's work: on a cache hit, your server is not involved at all.
- The variance: a warm cache answers in a consistent time, which matters because field data is measured at the 75th percentile of real loads.
- Compression and modern protocols usually arrive with it: Brotli for text, HTTP/2 and HTTP/3 on the connection.

Time to first byte is not itself a Core Web Vital, but Largest Contentful Paint inherits every millisecond of it. Fixing the first byte is the cheapest way to move a page that is close to the 2.5 second threshold.

## What crawling gains from good caching

| What you send | What a crawler does | Result |
| --- | --- | --- |
| ETag or Last-Modified on HTML | Sends a conditional request on the next crawl | A 304 with no body when nothing changed |
| Fast, consistent responses | Fetches more without straining your server | More of your pages seen per visit |
| 5xx errors or timeouts under load | Slows down and retries later | Fewer pages crawled, and host errors in the Crawl stats report |
| A CDN cache hit | Gets the same page a visitor gets | What Google measures matches what people experience |

The Crawl stats report in Search Console shows the response codes, the download sizes and the average response time over ninety days. A rising response time there is usually a caching problem before it is a code problem.

## The caching mistakes that cost you

- No cache headers at all, so every visitor and every crawler pays for a full render.
- Caching a logged-in page at the edge, which serves one person's dashboard to everyone. Mark private responses private and vary correctly.
- Caching a page that varies by cookie without telling the cache, which mixes up sessions.
- Hard-caching unhashed assets, so a deploy leaves visitors on the old stylesheet.
- no-store on public marketing pages, usually copied from an authenticated route.
- No purge on deploy, so a fix you shipped is invisible for the length of the window.
- Caching a 404 or a 500 for a long time, which keeps an error alive after you repaired it.

> The dangerous one is the logged-in page at the edge. Test it before you widen any cache rule: request a private URL without a session and see what comes back.

## A worked configuration for a small product site

| What | Cache-Control | Why |
| --- | --- | --- |
| Hashed build assets | public, max-age=31536000, immutable | The name changes when the content does |
| Images and fonts with stable names | public, max-age=604800 | Rarely change, but the name does not prove it |
| Marketing and guide pages | public, s-maxage=600, stale-while-revalidate=86400 | Fast for everyone, refreshed in the background |
| The sitemap and robots.txt | public, max-age=3600 | Fetched by crawlers, changes with deploys |
| Authenticated pages | private, no-store | Never cached anywhere shared |
| API responses used by the page | Depends, but never public by default | The usual source of an accidental leak |

Check your live site rather than your config: request a page and read the Cache-Control and ETag headers that actually come back. Platforms override what you set more often than you would think.

## Questions

### Does caching affect SEO directly?

Not as a ranking factor of its own. It affects time to first byte, which every paint timing inherits, and it affects how much a crawler can fetch without straining your server. Both show up in Core Web Vitals and in the Crawl stats report.

### How long should I cache HTML?

Briefly, with revalidation. A short edge window with stale-while-revalidate gives you static-like speed and lets a correction go live on the next request, especially if you purge on deploy.

### Can I cache a page that shows the visitor's name?

Not in a shared cache. Mark it private and no-store, or split the page so the personalised part is fetched separately and the shell is cached.

### Do I need a CDN for a small site?

If your visitors are not all near your server, yes. It removes most of the distance from the first byte, absorbs traffic spikes, and usually brings Brotli, HTTP/2 and HTTP/3 with it.

### What is stale-while-revalidate?

A directive that lets a cache serve the stored copy immediately while fetching a fresh one in the background. Nobody waits for the refresh, and your origin handles far fewer requests.

## Read next

- [Time to first byte: the delay every other metric inherits](https://porteur.ai/guides/time-to-first-byte): TTFB sets the pace for every other speed metric. See what it includes, how to test it in field and lab, the causes on product sites, and the fixes.
- [HTTP/2 and HTTP/3: what they change for a site’s speed](https://porteur.ai/guides/http-2-and-http-3): What each protocol fixes, where it is switched on, how to check which one your site serves, and how much it is worth beside your other speed work.
- [Crawl waste on a small site: the fetches your real pages never get](https://porteur.ai/guides/crawl-waste): Crawl budget is not a small site's problem. Crawl waste is: parameters, filters, old redirects and dead URLs taking the fetches your pages need.
- [The Crawl stats report: how much Google fetches and where it struggles](https://porteur.ai/guides/search-console-crawl-stats-report): Find Crawl stats in Search Console Settings. Read the four charts, host status and breakdowns. Spot 5xx spikes and wasted crawls. Know what is normal.
- [Does page speed affect SEO? What Google has said and what the data does](https://porteur.ai/guides/does-page-speed-affect-seo): Yes, speed affects SEO, but only at Core Web Vitals’ Good thresholds. Here is what Google measures, what the data shows, and what to fix first.
- [Core Web Vitals for a product site: the three numbers](https://porteur.ai/guides/core-web-vitals): The product founder’s guide to LCP, CLS and INP: what Google measures, why phones decide the pass, common causes, and how to test and fix.
- [CDN](https://porteur.ai/glossary/cdn): A CDN serves your site’s files from servers near visitors. It cuts TTFB, speeds pages, reduces origin load and handles TLS, HTTP/2 and HTTP/3.

Paste your URL and the free check measures how fast your pages answer and render on a phone, in about thirty seconds, with what to fix first. Free check: https://porteur.ai/
