Make a crash loop reach somebody instead of scrolling past.

`Restart=on-failure` with no start limit is an infinite loop by definition: the
unit never reaches `failed`, `systemctl status` stays active (auto-restart), and
the only evidence is a journal moving at four lines a second. That is how 464
restarts over fifteen hours went unnoticed.

All three units now stop after five failures in 120s and run
OnFailure=cashumints-alert@%n.service. The window is 120s and not 60s because
RestartSec=5s plus a process that takes a few seconds to die can spread five
failures past a sixty second window, reset the counter, and loop forever anyway.

cashumints-alert@.service is a oneshot that takes the failed unit's name as its
instance. Configuration is /etc/cashumints/alert.env: NTFY_URL gets a plain-text
body, WEBHOOK_URL gets JSON carrying `content` so one payload fits Discord and
Slack-compatible endpoints. With neither set — or the file absent — it still
writes to the journal at ERROR via a `<3>` syslog prefix, so `journalctl -p err -t
cashumints-alert` is a complete history on a host nobody configured.

It cannot become a second thing to debug: each curl is bounded at 10s, each
failure falls back to a journal line, and the shell ends in `true`, so the alerter
always exits 0. Verified with systemd-analyze verify and by running the ExecStart
body against a local sink — the JSON parses, and every branch exits 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
michilis
2026-08-25 16:10:01 +02:00
co-authored by Claude Opus 5
parent 9ffa53094d
commit 0ebc8ada54
6 changed files with 347 additions and 15 deletions
+19 -5
View File
@@ -3,13 +3,27 @@
Description=cashumints.space indexer and API
Wants=network-online.target
After=network-online.target
# Stop after five failures in a minute rather than restarting forever. A wrong Node on
# PATH once produced four thousand identical crashes in the journal before anyone read
# one of them; `failed` in `systemctl status` says the same thing in one line. Both keys
# belong to [Unit] — under [Service] systemd only warns and ignores them.
StartLimitIntervalSec=60
# Stop after five failures in two minutes rather than restarting forever.
#
# The window is 120s and not 60s because RestartSec=5s below means five attempts take
# a little over twenty seconds of restarts plus however long each attempt lives before
# it dies. A process that fails *slowly* — a database that times out, a port that takes
# four seconds to refuse — can spread five failures past a sixty second window and reset
# the counter forever, which is the loop this is supposed to stop. 120s covers that.
#
# The failure this exists for: ExecStart named a .ts file, /usr/bin/node was 20, and
# every start died in under a second. 464 restarts over fifteen hours, and because
# Restart=on-failure without a start limit never reaches a `failed` state, nothing
# anywhere went red. Both keys belong to [Unit] — under [Service] systemd only warns and
# ignores them.
StartLimitIntervalSec=120
StartLimitBurst=5
# Carry a failure off the machine. `%n` is this unit's own name, so the alert says
# which one died. cashumints-alert@.service writes to the journal at ERROR always and
# curls NTFY_URL or WEBHOOK_URL from /etc/cashumints/alert.env when either is set.
OnFailure=cashumints-alert@%n.service
[Service]
Type=simple
User=cashumints