How a Two‑Line Deploy Broke Checkout for 40 Minutes
It was 4:40 p.m. on a Friday. The deploy was tiny — a copy change and a small tweak to how the cart total was calculated. Two-line diff on the backend, a matching change on the frontend. It sailed through review, passed CI, and went out. Ten minutes later, checkout conversion fell off a cliff, and the support queue lit up with "the page is broken" and "it says my cart is empty." This is the story of how a two-line change took down checkout for 40 minutes — and why the cause was a header nobody had touched in two years.
The Setup
A React SPA on a CDN, talking to a REST API. The frontend was a single-page app served as static assets (main.[hash].js) behind a CDN; the API was versioned but not strictly. The cart total had always been computed on the backend and returned as { total, currency, lines }. The "small tweak" changed the response to { total, currency, lines, breakdown } on the backend — additive, non-breaking — and the frontend change rendered that new breakdown.
GIF via GIPHY
Both went out in the same release. In staging it was flawless. Everyone was already thinking about the weekend.
What Went Wrong
Within minutes: users on the checkout page saw a blank cart or a JavaScript error, and conversion cratered. But — and this was the confusing part — not everyone. Some sessions were fine. New users seemed okay. Returning users were broken. Refreshing sometimes fixed it, sometimes didn't. The error tracker filled with Cannot read properties of undefined (reading 'map') from a component that had worked for two years.
GIF via GIPHY
The first instinct was to blame the code change. We reverted the frontend. Conversion... partially recovered. Still broken for a chunk of users. Now people were really confused — we reverted the thing we changed and it was still broken.
The Investigation
Someone opened the Network tab on a broken session (a colleague could reproduce it on their logged-in account). The main.js being loaded was the old hash — the pre-deploy bundle. But the API was returning the new response shape. The old JS expected the old shape and choked on the new breakdown field interacting with a code path that assumed lines was structured the old way.
Wait — why was the old JS still loading after we deployed the new one? The CDN. The HTML shell that referenced main.[hash].js was being served with an aggressive Cache-Control: max-age=86400 — cached for a day. So returning users had the old HTML (pointing at the old JS), but their API calls hit the new backend. Old frontend, new backend. Version skew.
GIF via GIPHY
The revert made it worse in a subtle way: now there were three versions in the wild — users on the old cached bundle, users who'd gotten the brief new bundle, and users getting the reverted bundle — all hitting one backend.
The Root Cause
Two things, compounding:
- The HTML shell was cached for 24 hours (
Cache-Control: max-age=86400, set years ago and never revisited), so returning users kept loading the old bundle reference long after deploy. - The frontend and backend were deployed together and assumed lockstep — but the caching meant they were never actually in lockstep for returning users. The "additive, non-breaking" API change wasn't backwards-compatible with the old frontend the way we'd assumed, because the old frontend had a latent bug the new field triggered.
GIF via GIPHY
The change was fine. The deploy model was broken: we assumed frontend and backend versions matched, and a stale cache guaranteed they didn't.
The Fix
Emergency (minutes): purge the CDN cache for the HTML shell, forcing everyone onto the current bundle. Conversion recovered immediately once the old JS stopped loading.
Real fix (the next week):
GIF via GIPHY
- The HTML shell got
Cache-Control: no-cache(always revalidate) — so a deploy is picked up immediately — while the hashed JS/CSS assets keptmax-age=31536000, immutable(they're content-hashed, so caching them forever is correct). - We stopped assuming lockstep. The backend change was made genuinely backwards-compatible (the old frontend's code path was fixed to tolerate the new field), following the additive-evolution rule: the API must work with every deployed frontend, not just the newest.
- Deploys were sequenced: ship the tolerant backend first, let it bake, then the frontend.
The Lessons
- Cache the shell briefly, cache hashed assets forever. The single most important frontend caching rule. An HTML shell cached for a day means a deploy doesn't take effect for a day — for returning users, you're always running old code against a new backend. Content-hashed assets are immutable and safe to cache forever; the shell that references them must revalidate.
- "Additive and non-breaking" is only true if the old client tolerates it. Your API is consumed by every deployed version of the frontend, not the one in your editor. A field the old frontend chokes on is a breaking change, cache be damned.
- The revert made it worse because state was already in the wild. Rolling back frontend code doesn't roll back the bundles users already cached. In a version-skew incident, cache-purge is often the real lever, not revert.
- "Works in staging" hid it because staging had no returning users with day-old cached shells. The incident lived in the gap between your test conditions and production's messy reality.
GIF via GIPHY
This was the lived version of the caching-layers decision (cache immutable assets far/forever, dynamic content close/briefly), the API-versioning decision (evolve additively, tolerate old clients), and the production-only-bug playbook (the delta between dev and prod was a cache header). We'd written all three down. It happened anyway — because a Cache-Control header set two years ago was quietly holding the whole thing hostage.
What did you think?