← All writing
Shipping & Ops·Aug 24, 2026·6 min read

What actually breaks when a third-party API goes down

A payment processor, an email service, a geocoding API — every product leans on someone else's uptime. The outage that hurts isn't theirs, it's the one your own app inherits by not planning for it.

H

Hammad Iqbal

Software Engineer

Every integration demo assumes the third-party service responds in under 200ms and never returns an error. That assumption holds until the first time it doesn't, and by then the outage isn't the vendor's problem anymore — it's a support queue filling up with users who can't tell the difference between 'our servers are down' and 'a service we depend on is down.'

In brief
  • A dependency's outage becomes your outage the moment there's no timeout on the call.
  • Retrying blindly turns a blip into a cascading failure.
  • The degraded mode needs to be decided in the design, not improvised during the incident.

Set a timeout shorter than your users' patience

A request with no timeout doesn't fail when the third party has a bad day — it hangs, and every request behind it in the queue hangs with it. A slow dependency without a timeout is worse than a dependency that's fully down, because a clean error is at least something the app can handle. Pick a timeout based on what the user is waiting for, not on what feels generous to the API.

Retries need a ceiling and a backoff, not just a loop

Retrying a failed call is the right instinct and the wrong implementation if it's an unbounded loop with no delay — that's how a brief blip on the vendor's side turns into a self-inflicted denial-of-service against them, and against your own request queue. A capped number of retries with exponential backoff turns a transient failure into a transient failure, instead of an incident.

  • Cap retries at 2-3 attempts, not until it succeeds
  • Back off exponentially, with jitter, so retries don't all land in the same window
  • Circuit-break after repeated failures instead of continuing to hammer a dependency that's clearly down

Decide the degraded mode before you need it

When the geocoding API is down, does checkout block entirely, or does it fall back to a manual address field? When email delivery is delayed, does the app say so, or does it silently queue and let the user assume it worked? These are product decisions, and making them during an actual outage — under pressure, with a support queue growing — produces worse answers than making them in a design doc with nobody watching.

Takeaway

Budget for the dependency's bad day, not just its good one. The timeout, the retry ceiling, and the fallback behavior are part of the integration — not a follow-up ticket for after the first outage.

Have a project this kind of thinking applies to?

Tell me what you're building — I read every message myself.

Available for new projects

4+

Years exp.

8+

Projects shipped

<6h

Avg. reply time

Let's build something

Tell me about your project.

Share what you're building, your timeline, and the best way to reach you.