← All writing
Shipping & Ops·Sep 18, 2026·5 min read

The cron job that stopped running and nobody noticed

Alerting is built to catch a job that fails loudly. It has nothing to say about a job that simply never starts — and that's the failure mode that actually gets missed for weeks.

H

Hammad Iqbal

Software Engineer

The nightly job had been running for over a year without anyone thinking about it, which is exactly the problem: it had also stopped running three weeks before anyone noticed, and nothing paged. A container restart policy had changed during an infrastructure migration, the scheduled task didn't get re-registered on the new host, and the process that used to run every night at 2am simply never ran again. No error was thrown, because nothing tried to run and fail — it just didn't run. The alerting was wired to catch exceptions inside the job, and a job that never starts doesn't throw one.

In brief
  • Monitoring built around errors only fires when a job runs and fails — a job that never starts produces no error to catch.
  • Scheduled tasks get dropped by infrastructure changes far more often than by anyone deciding to remove them on purpose.
  • The fix is a heartbeat: something external that expects to hear from the job on schedule and pages when it goes quiet, not just when the job complains.

Alerts fire on errors, not on silence

Most job monitoring is built the same way: wrap the job body in a try/catch, send an alert on the catch. That design assumes the job at least starts. It says nothing about the case where the job never gets invoked at all — no process to throw, no stack trace to capture, no signal of any kind. From the monitoring's point of view, a job that runs and succeeds and a job that never runs look identical: silence.

How a job disappears without anyone touching it

Nobody usually decides to break a cron job. It happens as a side effect of something else changing:

  • A platform migration re-provisions the host and the crontab or scheduler config doesn't come with it
  • A container's restart policy resets on redeploy, and a task registered at runtime never re-registers
  • A typo in a cron expression gets accepted silently instead of rejected, and just never matches a time
  • An environment variable the job depends on gets renamed elsewhere, and the job now fails before it can even log that it started

Monitor for the absence, not just the failure

The pattern that actually catches this is a dead man's switch: the job pings a monitoring endpoint when it starts and again when it finishes, and that endpoint pages if it hasn't heard from the job within its expected window. The check isn't 'did the job throw' — it's 'did the job check in at all.' It's a small addition to the job itself, and it's the only design that notices silence instead of requiring an error to notice something.

Takeaway

A job that fails loudly gets fixed the same day. A job that stops running quietly can go for weeks, because the only thing watching it was built to catch errors, not absence.

Have a project this kind of thinking applies to?

Tell me what you're building — I read every message myself.

Available for new projects

4+

Years exp.

8+

Projects shipped

<6h

Avg. reply time

Let's build something

Tell me about your project.

Share what you're building, your timeline, and the best way to reach you.