← All writing
Cloud & Infra·Sep 8, 2026·6 min read

The retry that turned a blip into an outage

A downstream service hiccups for four seconds and recovers on its own. Your client doesn't know that — it retries immediately, three times, with no backoff, and the four-second hiccup becomes a four-minute outage.

H

Hammad Iqbal

Software Engineer

The dependency was fine within seconds. What wasn't fine was the queue of requests that had failed during those seconds, all retrying at once, all landing on a service that was just getting back on its feet. The graph doesn't show a four-second blip — it shows a four-minute outage with a shape that looks exactly like a denial-of-service attack, because that's functionally what it was, and you were the one running it.

In brief
  • A retry with no backoff turns a brief hiccup into a self-inflicted denial-of-service.
  • Every layer of a call stack that retries independently multiplies the load hitting the failing service.
  • Jitter isn't decoration — synchronized retries from many clients recreate the exact spike that caused the failure.

The outage you cause is longer than the outage you had

A service that drops requests for four seconds under load will happily serve them again the moment the load clears. The problem is that nothing clears it: every client that failed during those four seconds is programmed to try again immediately, so the service sees the same spike a second time, fails again, and now has a bigger backlog to retry. Without anything to break the cycle, a transient blip settles into a stable failure mode that only resolves when someone kills client traffic by hand.

Retries compound at every hop, not just at the edge

A single user request rarely makes one network call. It's a client hitting an API, which calls a service, which calls a database or a third-party API. If each of those hops has its own retry logic — and they usually do, because each was written by someone who read the same advice about handling transient failures — a single downstream failure gets retried by every layer above it independently.

  • Client retries 3x, API retries 3x, API's dependency call retries 3x — one failure becomes up to 27 attempts
  • None of the layers know about the others' retries, so none of them back off proportionally
  • The layer closest to the failure sees the full multiplied load, right when it's least able to handle it

Backoff, jitter, and a breaker are the whole fix

Exponential backoff spaces retries out so the second attempt doesn't arrive at the same instant as the first failure. Jitter — a small random delay added to that backoff — matters just as much: without it, every client computes the same backoff and they all retry in lockstep anyway, just a bit later. A circuit breaker that stops calling a dependency after enough failures, and fails fast instead of queuing more retries, is what actually ends the incident instead of just spreading it out over a longer window.

Takeaway

If a retry doesn't have backoff, jitter, and a limit, it isn't resilience — it's a load multiplier waiting for the one bad day it turns into an outage.

Have a project this kind of thinking applies to?

Tell me what you're building — I read every message myself.

Available for new projects

4+

Years exp.

8+

Projects shipped

<6h

Avg. reply time

Let's build something

Tell me about your project.

Share what you're building, your timeline, and the best way to reach you.