Make a crash loop reach somebody instead of scrolling past.

`Restart=on-failure` with no start limit is an infinite loop by definition: the
unit never reaches `failed`, `systemctl status` stays active (auto-restart), and
the only evidence is a journal moving at four lines a second. That is how 464
restarts over fifteen hours went unnoticed.

All three units now stop after five failures in 120s and run
OnFailure=cashumints-alert@%n.service. The window is 120s and not 60s because
RestartSec=5s plus a process that takes a few seconds to die can spread five
failures past a sixty second window, reset the counter, and loop forever anyway.

cashumints-alert@.service is a oneshot that takes the failed unit's name as its
instance. Configuration is /etc/cashumints/alert.env: NTFY_URL gets a plain-text
body, WEBHOOK_URL gets JSON carrying `content` so one payload fits Discord and
Slack-compatible endpoints. With neither set — or the file absent — it still
writes to the journal at ERROR via a `<3>` syslog prefix, so `journalctl -p err -t
cashumints-alert` is a complete history on a host nobody configured.

It cannot become a second thing to debug: each curl is bounded at 10s, each
failure falls back to a journal line, and the shell ends in `true`, so the alerter
always exits 0. Verified with systemd-analyze verify and by running the ExecStart
body against a local sink — the JSON parses, and every branch exits 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
michilis
2026-08-25 16:10:01 +02:00
co-authored by Claude Opus 5
parent 9ffa53094d
commit 0ebc8ada54
6 changed files with 347 additions and 15 deletions
+11 -5
View File
@@ -12,12 +12,18 @@
Description=cashumints.space static site server
Wants=network-online.target
After=network-online.target
# Give up after five failures in a minute instead of restarting forever. A process that
# cannot start will not start on the 4000th attempt either, and `failed` in
# `systemctl status` is a far louder signal than a journal scrolling past. These two are
# [Unit] keys; systemd ignores them under [Service] with only a warning.
StartLimitIntervalSec=60
# Give up after five failures in two minutes instead of restarting forever. A process
# that cannot start will not start on the 4000th attempt either, and `failed` in
# `systemctl status` is a far louder signal than a journal scrolling past. The window
# matches cashumints.service; see the note there for why it is 120s and not 60s. These
# two are [Unit] keys; systemd ignores them under [Service] with only a warning.
StartLimitIntervalSec=120
StartLimitBurst=5
# Carry a failure off the machine. `%n` is this unit's own name, so the alert says
# which one died. cashumints-alert@.service writes to the journal at ERROR always and
# curls NTFY_URL or WEBHOOK_URL from /etc/cashumints/alert.env when either is set.
OnFailure=cashumints-alert@%n.service
# Not Requires=cashumints.service: the pages are prerendered, so the site keeps serving
# a correct-as-of-last-build copy while the API is down. Only the islands go quiet.