Make a crash loop reach somebody instead of scrolling past.
`Restart=on-failure` with no start limit is an infinite loop by definition: the unit never reaches `failed`, `systemctl status` stays active (auto-restart), and the only evidence is a journal moving at four lines a second. That is how 464 restarts over fifteen hours went unnoticed. All three units now stop after five failures in 120s and run OnFailure=cashumints-alert@%n.service. The window is 120s and not 60s because RestartSec=5s plus a process that takes a few seconds to die can spread five failures past a sixty second window, reset the counter, and loop forever anyway. cashumints-alert@.service is a oneshot that takes the failed unit's name as its instance. Configuration is /etc/cashumints/alert.env: NTFY_URL gets a plain-text body, WEBHOOK_URL gets JSON carrying `content` so one payload fits Discord and Slack-compatible endpoints. With neither set — or the file absent — it still writes to the journal at ERROR via a `<3>` syslog prefix, so `journalctl -p err -t cashumints-alert` is a complete history on a host nobody configured. It cannot become a second thing to debug: each curl is bounded at 10s, each failure falls back to a journal line, and the shell ends in `true`, so the alerter always exits 0. Verified with systemd-analyze verify and by running the ExecStart body against a local sink — the JSON parses, and every branch exits 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
9ffa53094d
commit
0ebc8ada54
@@ -12,12 +12,18 @@
|
||||
Description=cashumints.space static site server
|
||||
Wants=network-online.target
|
||||
After=network-online.target
|
||||
# Give up after five failures in a minute instead of restarting forever. A process that
|
||||
# cannot start will not start on the 4000th attempt either, and `failed` in
|
||||
# `systemctl status` is a far louder signal than a journal scrolling past. These two are
|
||||
# [Unit] keys; systemd ignores them under [Service] with only a warning.
|
||||
StartLimitIntervalSec=60
|
||||
# Give up after five failures in two minutes instead of restarting forever. A process
|
||||
# that cannot start will not start on the 4000th attempt either, and `failed` in
|
||||
# `systemctl status` is a far louder signal than a journal scrolling past. The window
|
||||
# matches cashumints.service; see the note there for why it is 120s and not 60s. These
|
||||
# two are [Unit] keys; systemd ignores them under [Service] with only a warning.
|
||||
StartLimitIntervalSec=120
|
||||
StartLimitBurst=5
|
||||
|
||||
# Carry a failure off the machine. `%n` is this unit's own name, so the alert says
|
||||
# which one died. cashumints-alert@.service writes to the journal at ERROR always and
|
||||
# curls NTFY_URL or WEBHOOK_URL from /etc/cashumints/alert.env when either is set.
|
||||
OnFailure=cashumints-alert@%n.service
|
||||
# Not Requires=cashumints.service: the pages are prerendered, so the site keeps serving
|
||||
# a correct-as-of-last-build copy while the API is down. Only the islands go quiet.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user