Solution / Stuck processes

Detect Processes That Are Alive but Stuck

Separate liveness from progress so a worker that keeps checking in cannot hide a frozen processing loop.

Solution pattern; Dead Hand engine not yet implemented

Why a live process can still be failing

A process may respond to a health endpoint and send heartbeats while its useful work has stopped. It may be blocked on one record, retrying the same input, waiting on a dependency or reporting a marker that never changes.

Use a progress marker

Choose an application-defined marker that should move while work is active: a batch number, cursor, sequence, record identifier or phase. Send it with the heartbeat only when its meaning is clear. HeartbeatHook should observe the marker; the application remains the authority for what progress means.

Three different failure classes

  • process gone: no process evidence and no heartbeat;
  • missing check-in: expected heartbeat disappears;
  • alive but stuck: heartbeat continues while progress is stale.

Limitations

Unchanged progress may be normal during idle phases. Thresholds must account for active versus idle work, expected pauses and retries. This solution does not validate every output record or replace traces and metrics.

HeartbeatHook direction

Dead Hand Monitor is the HeartbeatHook name for this progress-freshness concept. The current HPH build does not yet evaluate progress markers, so the live Tier 1 path can provide the reporting foundation but not the final stuck state.

Explore Dead Hand Monitor