Make a crash loop reach somebody instead of scrolling past.
`Restart=on-failure` with no start limit is an infinite loop by definition: the unit never reaches `failed`, `systemctl status` stays active (auto-restart), and the only evidence is a journal moving at four lines a second. That is how 464 restarts over fifteen hours went unnoticed. All three units now stop after five failures in 120s and run OnFailure=cashumints-alert@%n.service. The window is 120s and not 60s because RestartSec=5s plus a process that takes a few seconds to die can spread five failures past a sixty second window, reset the counter, and loop forever anyway. cashumints-alert@.service is a oneshot that takes the failed unit's name as its instance. Configuration is /etc/cashumints/alert.env: NTFY_URL gets a plain-text body, WEBHOOK_URL gets JSON carrying `content` so one payload fits Discord and Slack-compatible endpoints. With neither set — or the file absent — it still writes to the journal at ERROR via a `<3>` syslog prefix, so `journalctl -p err -t cashumints-alert` is a complete history on a host nobody configured. It cannot become a second thing to debug: each curl is bounded at 10s, each failure falls back to a journal line, and the shell ends in `true`, so the alerter always exits 0. Verified with systemd-analyze verify and by running the ExecStart body against a local sink — the JSON parses, and every branch exits 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
9ffa53094d
commit
0ebc8ada54
@@ -3,13 +3,27 @@
|
||||
Description=cashumints.space indexer and API
|
||||
Wants=network-online.target
|
||||
After=network-online.target
|
||||
# Stop after five failures in a minute rather than restarting forever. A wrong Node on
|
||||
# PATH once produced four thousand identical crashes in the journal before anyone read
|
||||
# one of them; `failed` in `systemctl status` says the same thing in one line. Both keys
|
||||
# belong to [Unit] — under [Service] systemd only warns and ignores them.
|
||||
StartLimitIntervalSec=60
|
||||
# Stop after five failures in two minutes rather than restarting forever.
|
||||
#
|
||||
# The window is 120s and not 60s because RestartSec=5s below means five attempts take
|
||||
# a little over twenty seconds of restarts plus however long each attempt lives before
|
||||
# it dies. A process that fails *slowly* — a database that times out, a port that takes
|
||||
# four seconds to refuse — can spread five failures past a sixty second window and reset
|
||||
# the counter forever, which is the loop this is supposed to stop. 120s covers that.
|
||||
#
|
||||
# The failure this exists for: ExecStart named a .ts file, /usr/bin/node was 20, and
|
||||
# every start died in under a second. 464 restarts over fifteen hours, and because
|
||||
# Restart=on-failure without a start limit never reaches a `failed` state, nothing
|
||||
# anywhere went red. Both keys belong to [Unit] — under [Service] systemd only warns and
|
||||
# ignores them.
|
||||
StartLimitIntervalSec=120
|
||||
StartLimitBurst=5
|
||||
|
||||
# Carry a failure off the machine. `%n` is this unit's own name, so the alert says
|
||||
# which one died. cashumints-alert@.service writes to the journal at ERROR always and
|
||||
# curls NTFY_URL or WEBHOOK_URL from /etc/cashumints/alert.env when either is set.
|
||||
OnFailure=cashumints-alert@%n.service
|
||||
|
||||
[Service]
|
||||
Type=simple
|
||||
User=cashumints
|
||||
|
||||
Reference in New Issue
Block a user