← All writing
Engineering·Sep 11, 2026·5 min read

The cache that served stale data for three days

The write succeeded, the database was correct, and the support ticket said the change 'isn't showing up.' It was showing up — to everyone except the cache key nobody remembered to invalidate.

H

Hammad Iqbal

Software Engineer

The database had the right value the entire time. So did every log line, every admin panel, every direct query anyone ran to check. The only thing lying was the cache sitting in front of the read path, holding a value from three days earlier because the write path updated the row and moved on without telling it anything had changed. Nobody had done anything wrong, exactly — they'd just written the write and the read as if the cache didn't exist.

In brief
  • A cache with no invalidation path isn't a performance optimization, it's a second source of truth that silently drifts from the first.
  • TTL-only caching turns every bug report into a timing question: is this stale, or is this actually wrong?
  • Invalidating on write is cheap when the write path is the one place a value changes — the cost shows up once you have two.

A cache is a promise you have to keep on every write

Reading from a cache is the easy half. The moment you add a cache in front of a value, you've taken on an obligation to update or clear that cache everywhere the underlying value can change — not just the one code path you were optimizing when you added it. Every other write path, migration, admin edit, and background job that touches the same row now has to know the cache exists, or it doesn't, and the cache quietly goes stale.

TTL is a mitigation, not a strategy

A short TTL bounds how wrong the cache can be, which is why it feels like a fix. It isn't one — it just trades an unbounded staleness bug for a bounded one, and a bounded staleness bug is still a staleness bug the next time someone needs the write to be visible immediately, like right after they made it.

  • A five-minute TTL means every 'fix' can take five minutes to confirm, which teaches people to just wait instead of reporting
  • Support tickets about stale data become indistinguishable from tickets about actually-broken data until someone checks the timestamp
  • Lowering the TTL to 'fix' staleness just moves load back onto whatever the cache was protecting

Invalidate at the write, not on a timer

The reliable version deletes or updates the cache key in the same transaction, or the same request, as the write that changed the underlying value. It costs a few extra lines at every write site, which is exactly why it gets skipped under deadline pressure — but it's the only version where 'I just saved it' and 'it shows the new value' are the same guarantee instead of a coincidence that depends on timing.

Takeaway

If you can't point to the line of code that invalidates the cache on every write path, you don't have a caching strategy — you have a TTL and a countdown until someone notices.

Have a project this kind of thinking applies to?

Tell me what you're building — I read every message myself.

Available for new projects

4+

Years exp.

8+

Projects shipped

<6h

Avg. reply time

Let's build something

Tell me about your project.

Share what you're building, your timeline, and the best way to reach you.