Your website is behind a CDN, but origin requests and bandwidth suddenly start climbing. That does not always mean the CDN has stopped working. More often, requests for the same content are being divided into many different cache objects, so edge nodes cannot reuse an existing response.
Confirm where the drop begins
Start with the time when the cache hit ratio changed. Compare it with origin request volume and origin bandwidth. Then request the same static file several times and inspect headers such as Age, Cache-Status, CF-Cache-Status, or X-Cache. Header names and values differ between providers, so interpret them using your CDN's documentation.
The goal is to answer three questions: Is the response eligible for caching? Does the second identical request hit the cache? Which request attribute creates a different cache key?
Check whether query strings create unnecessary variants
Tracking parameters such as utm_source and fbclid usually do not change page content, but they may still create separate cache objects. A popular URL with many campaign parameters can therefore produce many cold-cache requests.
Keep parameters that genuinely change the response. Parameters used only for analytics can often be ignored, normalized, or removed from the cache key. Do not ignore every query string globally: search, filtering, language, image resizing, signed URLs, and API parameters may return different content.
Check whether cookies prevent public pages from sharing cache
Analytics plugins, consent banners, and A/B testing tools may add cookies to every request. If the complete Cookie header becomes part of the cache key, two visitors requesting the same public page may receive different cache keys.
For public content, include only cookies that actually change the response. Logged-in pages, shopping carts, account areas, and personalized responses should normally bypass shared caching. Never remove authentication or personalization signals merely to improve the hit ratio.
Review the Vary response header
Vary: Accept-Encoding commonly separates compressed representations. Other values can fragment the cache much more aggressively. For example, Vary: User-Agent may create variants for a large number of browser identifiers, while Vary: * generally makes a response unsuitable for normal shared caching.
If mobile and desktop content must differ, use a small, controlled set of device categories instead of the full User-Agent string. Remove unnecessary Vary fields only after confirming that those request headers do not change the response.
Verify authorization, cache directives, and TTLs
Requests carrying Authorization often contain protected or user-specific data and require conservative cache handling. Do not force them into a public cache without a design that explicitly prevents data leakage.
Also inspect Cache-Control. Directives such as private and no-store, a very short max-age, or an edge TTL that is too low can cause frequent revalidation or expiration. Confirm whether the CDN respects origin headers, overrides them, or applies separate browser and edge TTL values.
Fix the cache key in a controlled order
First isolate the affected hostnames and paths. Compare a normal request with a miss: URL, query parameters, cookies, relevant request headers, response headers, and CDN rule matches. Confirm that the response is cacheable before changing the cache key.
Then remove one unnecessary dimension at a time. Start with tracking parameters, then review cookies, Vary, and TTL rules. A field should be excluded only when it cannot change the response. Make the narrowest possible rule instead of applying a site-wide exception.
After the change, purge only the affected URLs or directory when possible. Request a fixed test URL repeatedly and confirm that the second request hits the cache. Test language variants, logged-in and logged-out sessions, device categories, and parameters that are supposed to change content.
Monitor the result without chasing a perfect percentage
Track cache hit ratio together with origin requests, origin bandwidth, response time, and 4xx/5xx rates. A high hit ratio is useful only when shared content remains correct and private content stays private. The practical target is stable reuse for cacheable responses, not the highest possible percentage at the cost of incorrect or unsafe caching.