Production system Case studyLive under all six systemsUnattended operations

The Reliability Layer Under a Production AI Stack

Twenty-eight automation workflows running unattended on a real business, where health is proven by fresh output rather than asserted by an exit code, and silence is treated as a failure. This is the half of AI work that does not demo well and decides everything.

28workflows in daily operation2x dailywatchdog on every jobProofof output, not exit codesFail closedsilence raises an alarm

The problem

The failure mode of production automation is almost never a crash.

  • It is a job that runs, exits zero, and quietly does nothing. For weeks. While everyone downstream assumes it is working, because it never complained.
  • Exit codes lie. A workflow that reached the end of its steps having written nothing is a failure that reports success, and every monitoring system that watches process status agrees with it.
  • Credentials expire on their own schedule. A rotating token chain that is never exercised is a chain that breaks at the worst possible moment and takes a live system down with it.
  • And nobody notices. That is the actual root cause of most automation outages. No model choice, framework, or vendor fixes it, which is why it is the first thing to build and usually the last thing anyone builds.

The solution

A job proves it is alive by leaving something dated behind.

Every recurring job writes fresh, dateable output as a condition of being considered healthy. A watchdog runs on its own independent schedule, twice a day, and compares the age of that output against the cadence the job promised.

A job that ran but produced nothing new is treated as failed, because it is one. The absence of fresh output raises an alarm rather than passing quietly, which inverts the default assumption from success to failure. For unattended work, that is the correct default.

The whole discipline in one lineHealth is proven by output, never asserted by an exit code. A job that ran and did nothing is a failure, and treating it as one is what separates automation that survives a quarter from automation that is quietly dead by week three.

How it works

Eight practices, applied to every workflow without exception.

ProofFreshness, not exit codes

Every recurring job writes output a watchdog can date. Health is the age of that output, which cannot be faked by a process that finished cleanly having done nothing.

WatchdogTwice a day, every job

An independent watchdog checks each job's output age against its promised cadence. Crucially it is not the job reporting on itself, because a broken job is exactly the wrong narrator of its own condition.

AlarmSilence is the signal

Absence of fresh output raises an alarm. The default assumption is failure until proven otherwise, which is the only safe posture for work nobody is watching.

CredentialsKept alive on purpose

Rotating credential chains are refreshed on a schedule and then spent on a real call, so the refresh is proven to have worked rather than assumed. A token that was renewed but never used is an untested token.

TimeoutsEvery network call, no exceptions

No external call may hang indefinitely. A slow third party degrades one job instead of stalling a queue behind it and taking unrelated systems with it.

PartialsA good result is never overwritten by a bad one

A partial or failed fetch does not replace a known-good result. Stale but correct beats fresh but truncated, every time, and the distinction is enforced in code.

BrittlenessNothing a rename can break

Identifiers, field names, and paths are resolved at run time rather than hardcoded, so an ordinary rename upstream does not silently empty a pipeline that continues to report success.

RecoveryBackfill is part of the design

When an upstream dependency fails, the recovery path replays what was missed rather than leaving a permanent hole in the record and moving on.

01Job runsOn its own schedule
02Writes dated outputFreshness is the health signal
03Watchdog reads ageIndependent, twice daily
04FreshHealthy. No action
05Stale or absentAlarm raised. Assumed failed
06Human investigatesA person, not a retry loop
07Backfill replaysThe gap is filled, not skipped
The watchdog is deliberately separate from the jobs it watches. A broken job is the wrong narrator.

Results and evidence

Measured on a business that depends on these workflows daily.

28
Live workflows in daily operation
2x daily
Watchdog pass over every job
0
Jobs trusted to report their own health
Same day
Third-party outage caught and fully restored
100%
Network calls carrying an explicit timeout
Backfill
Replays the gap rather than skipping it
Freshness
The single health metric, across the stack
Fail closed
Silence is treated as failure by default
The incident this was built for

A third-party vector database outage was caught and fully restored the same day, with backfill and no customer impact. The layer described on this page is the reason it was caught in hours rather than discovered in weeks by a customer.

Why the watchdog is separate

Self-reporting health checks fail in exactly the situations that matter, because the component reporting is the component that is broken. An independent observer with its own schedule is the only arrangement that survives the real failure.

The unglamorous half

Anyone can ship a demo. Keeping automation alive, unattended, on a business that depends on it, is the harder half, and it is the half that decides whether an AI program survives its first quarter or becomes the thing nobody mentions in the second one.

What it demonstrates

The skills behind the system.

Production reliability engineeringObservability for workflow automationWatchdog and heartbeat designProof-of-output health checksCredential lifecycle managementTimeout and graceful degradation policySafe overwrite and partial-result rulesBackfill and replay recoveryProduction incident responseFailure-mode engineeringUnattended operations at scale

Outcome

This is why the other systems are described as live.

Twenty-eight workflows run daily on a real business with health proven by output rather than assumed from exit codes. This layer sits underneath every system in this portfolio. It is the reason those systems are described here as live and operating, rather than as launched and hoped for.