Make a crash loop reach somebody instead of scrolling past.

`Restart=on-failure` with no start limit is an infinite loop by definition: the
unit never reaches `failed`, `systemctl status` stays active (auto-restart), and
the only evidence is a journal moving at four lines a second. That is how 464
restarts over fifteen hours went unnoticed.

All three units now stop after five failures in 120s and run
OnFailure=cashumints-alert@%n.service. The window is 120s and not 60s because
RestartSec=5s plus a process that takes a few seconds to die can spread five
failures past a sixty second window, reset the counter, and loop forever anyway.

cashumints-alert@.service is a oneshot that takes the failed unit's name as its
instance. Configuration is /etc/cashumints/alert.env: NTFY_URL gets a plain-text
body, WEBHOOK_URL gets JSON carrying `content` so one payload fits Discord and
Slack-compatible endpoints. With neither set — or the file absent — it still
writes to the journal at ERROR via a `<3>` syslog prefix, so `journalctl -p err -t
cashumints-alert` is a complete history on a host nobody configured.

It cannot become a second thing to debug: each curl is bounded at 10s, each
failure falls back to a journal line, and the shell ends in `true`, so the alerter
always exits 0. Verified with systemd-analyze verify and by running the ExecStart
body against a local sink — the JSON parses, and every branch exits 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
michilis
2026-08-25 16:10:01 +02:00
co-authored by Claude Opus 5
parent 9ffa53094d
commit 0ebc8ada54
6 changed files with 347 additions and 15 deletions
+188 -5
View File
@@ -933,12 +933,15 @@ laptop requirement, not a server one.
Description=cashumints.space indexer and API Description=cashumints.space indexer and API
Wants=network-online.target Wants=network-online.target
After=network-online.target After=network-online.target
# Stop after five failures in a minute rather than restarting forever: a process that # Stop after five failures in two minutes rather than restarting forever: a process
# cannot start will not start on the 4000th attempt either, and `failed` in # that cannot start will not start on the 4000th attempt either, and `failed` in
# `systemctl status` is a louder signal than a journal scrolling past. Both keys belong # `systemctl status` is a louder signal than a journal scrolling past. Both keys belong
# to [Unit] — under [Service] systemd only warns and ignores them. # to [Unit] — under [Service] systemd only warns and ignores them. See "Failing loudly"
StartLimitIntervalSec=60 # for why the window is 120s and not 60s.
StartLimitIntervalSec=120
StartLimitBurst=5 StartLimitBurst=5
# And carry that `failed` off the machine. %n is this unit's own name.
OnFailure=cashumints-alert@%n.service
[Service] [Service]
Type=simple Type=simple
@@ -1009,8 +1012,9 @@ Two behaviours are worth knowing about because they are load-bearing:
Description=cashumints.space static site server Description=cashumints.space static site server
Wants=network-online.target Wants=network-online.target
After=network-online.target After=network-online.target
StartLimitIntervalSec=60 StartLimitIntervalSec=120
StartLimitBurst=5 StartLimitBurst=5
OnFailure=cashumints-alert@%n.service
# Not Requires=cashumints.service: the pages are prerendered, so the site keeps serving # Not Requires=cashumints.service: the pages are prerendered, so the site keeps serving
# a correct-as-of-last-build copy while the API is down. Only the islands go quiet. # a correct-as-of-last-build copy while the API is down. Only the islands go quiet.
@@ -1175,6 +1179,184 @@ nginx 1.25 and later want `http2 on;` on its own line and warn about the `listen
form above; Debian 12 ships 1.22, where the newer form is an unknown directive. The form form above; Debian 12 ships 1.22, where the newer form is an unknown directive. The form
above is the one that works on both. above is the one that works on both.
### Failing loudly
Two separate silences produced today's incident, and they need separate fixes.
The first is a **crash loop that never reports a failure**. `Restart=on-failure` with no
start limit is an infinite loop by definition: systemd restarts, the process dies,
systemd restarts. The unit never reaches `failed`, so `systemctl status` stays `active
(auto-restart)`, nothing sends anything anywhere, and the only evidence is a journal
scrolling past at four lines a second. The API did this 464 times over fifteen hours.
All three units now carry:
```ini
StartLimitIntervalSec=120
StartLimitBurst=5
OnFailure=cashumints-alert@%n.service
```
**Why 120 and not 60.** `RestartSec=5s` means five attempts cost a little over twenty
seconds of waiting, plus however long each attempt survives before dying. A process that
fails *slowly* — a database connection that times out, a port that takes four seconds to
refuse — spreads five failures past a sixty second window, resets the counter, and loops
forever anyway. 120s covers the slow case. Both keys belong under `[Unit]`; systemd
takes them under `[Service]` with only a warning and then ignores them.
**Why `OnFailure=` at all.** `StartLimitBurst` turns the loop into a `failed` state,
which is much better, and is still a state somebody has to go and look at. `OnFailure=`
is what makes the machine speak first. `%n` expands to the failed unit's own name, which
arrives as the template instance in `%i`.
The second silence is a **build that publishes an index it should have refused**; that
one is the `MIN_MINTS_FOR_BUILD` gate under "Rebuilds", and the `discovery_starved` flag
under "Discovery starvation" is what feeds it.
#### The alert unit
```ini
# /etc/systemd/system/cashumints-alert@.service
#
# The unit that makes a failure audible.
#
# The other three units each carry `OnFailure=cashumints-alert@%n.service`, so systemd
# starts one of these with the failed unit's name as the instance — `%i` below is
# literally `cashumints.service`, `cashumints-web.service` or `cashumints-site.service`.
#
# Why it exists: the API once crash looped 464 times over fifteen hours and nothing said
# so. `Restart=on-failure` with no start limit is an infinite loop that never reaches a
# `failed` state, so the journal filled with identical lines nobody was reading and
# every signal stayed green. The other half of the fix is StartLimitBurst= in each unit,
# which turns the loop into a failure; this is what carries that failure off the machine.
#
# Install:
# sudo install -m 0644 deploy/cashumints-alert@.service /etc/systemd/system/
# sudo install -d -m 0755 /etc/cashumints
# sudo install -m 0640 -o root -g root deploy/alert.env.example /etc/cashumints/alert.env
# sudo systemctl daemon-reload
#
# No [Install] section and never enabled: OnFailure= starts it, and a unit that also
# started at boot would page on every reboot.
[Unit]
Description=Notify that %i failed
# No OnFailure= here. An alerter that alerts about its own failure is a loop, and this
# one is written so its worst case is a journal line rather than a retry.
[Service]
Type=oneshot
# The one file an operator edits, and the only reason this unit is configurable at all.
# Absent is a supported state — the leading `-` says so — and then the ExecStart below
# still writes to the journal at ERROR, which is what `systemctl status` and
# `journalctl -p err` read. See alert.env.example.
EnvironmentFile=-/etc/cashumints/alert.env
# So `journalctl -t cashumints-alert` finds every alert, whichever unit triggered it.
SyslogIdentifier=cashumints-alert
# Everything is inside one shell so the "nothing configured" branch is reachable without
# a second unit. The pieces, in order:
#
# - `printf '<3>…'` on stdout. systemd reads that syslog prefix off a journal stream
# and files the line at priority 3, ERROR, so `journalctl -p err` is a complete
# history of failures on a host with no webhook configured at all. `<4>` is warning.
# A prefix rather than systemd-cat, so the unit needs nothing from the filesystem it
# has just sandboxed itself away from.
# - NTFY_URL is a topic URL (https://ntfy.sh/your-topic). It gets a plain-text body
# naming the failed unit, plus the header names ntfy understands.
# - WEBHOOK_URL gets a JSON POST instead, for Discord, Slack or anything that speaks
# `{"content": …}` — every key is sent, so one payload fits all of them.
# - `--max-time 10` and a `||` fallback on each: an alert that hangs would hold the
# failed unit's job open, and an alert that fails must not itself become a second
# failed unit for somebody to notice. The shell ends in `true` for the same reason.
#
# `%i` is the failed unit's name, passed as an argument rather than interpolated into
# the shell text: systemd expands specifiers before /bin/sh ever sees the line, and a
# unit name is not a thing to trust to quoting.
ExecStart=/bin/sh -c '\
UNIT="$1"; \
HOST="$(hostname)"; \
WHEN="$(date -Is)"; \
TEXT="$UNIT failed on $HOST at $WHEN"; \
printf "<3>%s\\n" "$TEXT"; \
if [ -n "$NTFY_URL" ]; then \
/usr/bin/curl -fsS --max-time 10 \
-H "Title: cashumints: $UNIT failed" \
-H "Priority: high" \
-H "Tags: rotating_light" \
-d "$TEXT" "$NTFY_URL" >/dev/null \
|| printf "<3>%s\\n" "alert: POST to NTFY_URL failed"; \
fi; \
if [ -n "$WEBHOOK_URL" ]; then \
/usr/bin/curl -fsS --max-time 10 \
-H "Content-Type: application/json" \
-d "{\\"unit\\":\\"$UNIT\\",\\"host\\":\\"$HOST\\",\\"at\\":\\"$WHEN\\",\\"text\\":\\"$TEXT\\",\\"content\\":\\"$TEXT\\"}" \
"$WEBHOOK_URL" >/dev/null \
|| printf "<3>%s\\n" "alert: POST to WEBHOOK_URL failed"; \
fi; \
if [ -z "$NTFY_URL" ] && [ -z "$WEBHOOK_URL" ]; then \
printf "<4>%s\\n" "alert: no NTFY_URL or WEBHOOK_URL in /etc/cashumints/alert.env, journal only"; \
fi; \
true' _ %i
# It sends one HTTP request and writes one line. It needs no identity of its own, and
# DynamicUser gives it a throwaway one rather than sharing `nobody` with everything else
# on the host that also could not be bothered to make a user.
DynamicUser=yes
NoNewPrivileges=true
PrivateDevices=true
ProtectSystem=strict
ProtectHome=true
ProtectKernelTunables=true
ProtectKernelModules=true
ProtectControlGroups=true
RestrictAddressFamilies=AF_UNIX AF_INET AF_INET6
RestrictSUIDSGID=true
LockPersonality=true
# An alert that cannot reach the network in ten seconds is not worth a stuck job.
TimeoutStartSec=30
```
Configuration is one optional file. With it absent, or with both values empty, a failure
still lands in the journal at `ERROR` and is readable with `journalctl -p err -t
cashumints-alert`; the unit is written so that its worst case is a log line rather than
a second thing to debug.
```ini
# /etc/cashumints/alert.env
#
# Read by cashumints-alert@.service, which systemd starts when any of the three units
# fails. Everything here is optional: with the file absent or both values empty, an
# alert is still written to the journal at ERROR priority and is readable with
#
# journalctl -p err -t cashumints-alert
#
# Set one or both to have failures leave the machine.
#
# Install it root-owned and not world-readable — a webhook URL is a capability:
# sudo install -d -m 0755 /etc/cashumints
# sudo install -m 0640 -o root -g root deploy/alert.env.example /etc/cashumints/alert.env
# sudo systemctl daemon-reload
# An ntfy topic URL. Free and public at ntfy.sh; pick a topic name nobody will guess,
# because anyone who knows it can read and post to it.
#NTFY_URL=https://ntfy.sh/cashumints-alerts-CHANGE-ME
# Anything that accepts a JSON POST. The body carries `unit`, `host`, `at`, `text` and
# `content` — the last of which is what Discord and most Slack-compatible endpoints read,
# so one payload fits all three.
#WEBHOOK_URL=https://discord.com/api/webhooks/…
```
Test it without breaking anything:
```bash
sudo systemctl start 'cashumints-alert@test.service'
journalctl -t cashumints-alert -n 5 --no-pager
```
### Rebuilds ### Rebuilds
Mint pages are prerendered, so new mints and new review counts appear at the next build. Mint pages are prerendered, so new mints and new review counts appear at the next build.
@@ -1196,6 +1378,7 @@ Description=Rebuild the cashumints.space static site
Requires=cashumints.service Requires=cashumints.service
After=cashumints.service network-online.target After=cashumints.service network-online.target
Wants=network-online.target Wants=network-online.target
OnFailure=cashumints-alert@%n.service
[Service] [Service]
Type=oneshot Type=oneshot
+23
View File
@@ -0,0 +1,23 @@
# /etc/cashumints/alert.env
#
# Read by cashumints-alert@.service, which systemd starts when any of the three units
# fails. Everything here is optional: with the file absent or both values empty, an
# alert is still written to the journal at ERROR priority and is readable with
#
# journalctl -p err -t cashumints-alert
#
# Set one or both to have failures leave the machine.
#
# Install it root-owned and not world-readable — a webhook URL is a capability:
# sudo install -d -m 0755 /etc/cashumints
# sudo install -m 0640 -o root -g root deploy/alert.env.example /etc/cashumints/alert.env
# sudo systemctl daemon-reload
# An ntfy topic URL. Free and public at ntfy.sh; pick a topic name nobody will guess,
# because anyone who knows it can read and post to it.
#NTFY_URL=https://ntfy.sh/cashumints-alerts-CHANGE-ME
# Anything that accepts a JSON POST. The body carries `unit`, `host`, `at`, `text` and
# `content` — the last of which is what Discord and most Slack-compatible endpoints read,
# so one payload fits all three.
#WEBHOOK_URL=https://discord.com/api/webhooks/…
+101
View File
@@ -0,0 +1,101 @@
# /etc/systemd/system/cashumints-alert@.service
#
# The unit that makes a failure audible.
#
# The other three units each carry `OnFailure=cashumints-alert@%n.service`, so systemd
# starts one of these with the failed unit's name as the instance — `%i` below is
# literally `cashumints.service`, `cashumints-web.service` or `cashumints-site.service`.
#
# Why it exists: the API once crash looped 464 times over fifteen hours and nothing said
# so. `Restart=on-failure` with no start limit is an infinite loop that never reaches a
# `failed` state, so the journal filled with identical lines nobody was reading and
# every signal stayed green. The other half of the fix is StartLimitBurst= in each unit,
# which turns the loop into a failure; this is what carries that failure off the machine.
#
# Install:
# sudo install -m 0644 deploy/cashumints-alert@.service /etc/systemd/system/
# sudo install -d -m 0755 /etc/cashumints
# sudo install -m 0640 -o root -g root deploy/alert.env.example /etc/cashumints/alert.env
# sudo systemctl daemon-reload
#
# No [Install] section and never enabled: OnFailure= starts it, and a unit that also
# started at boot would page on every reboot.
[Unit]
Description=Notify that %i failed
# No OnFailure= here. An alerter that alerts about its own failure is a loop, and this
# one is written so its worst case is a journal line rather than a retry.
[Service]
Type=oneshot
# The one file an operator edits, and the only reason this unit is configurable at all.
# Absent is a supported state — the leading `-` says so — and then the ExecStart below
# still writes to the journal at ERROR, which is what `systemctl status` and
# `journalctl -p err` read. See alert.env.example.
EnvironmentFile=-/etc/cashumints/alert.env
# So `journalctl -t cashumints-alert` finds every alert, whichever unit triggered it.
SyslogIdentifier=cashumints-alert
# Everything is inside one shell so the "nothing configured" branch is reachable without
# a second unit. The pieces, in order:
#
# - `printf '<3>…'` on stdout. systemd reads that syslog prefix off a journal stream
# and files the line at priority 3, ERROR, so `journalctl -p err` is a complete
# history of failures on a host with no webhook configured at all. `<4>` is warning.
# A prefix rather than systemd-cat, so the unit needs nothing from the filesystem it
# has just sandboxed itself away from.
# - NTFY_URL is a topic URL (https://ntfy.sh/your-topic). It gets a plain-text body
# naming the failed unit, plus the header names ntfy understands.
# - WEBHOOK_URL gets a JSON POST instead, for Discord, Slack or anything that speaks
# `{"content": …}` — every key is sent, so one payload fits all of them.
# - `--max-time 10` and a `||` fallback on each: an alert that hangs would hold the
# failed unit's job open, and an alert that fails must not itself become a second
# failed unit for somebody to notice. The shell ends in `true` for the same reason.
#
# `%i` is the failed unit's name, passed as an argument rather than interpolated into
# the shell text: systemd expands specifiers before /bin/sh ever sees the line, and a
# unit name is not a thing to trust to quoting.
ExecStart=/bin/sh -c '\
UNIT="$1"; \
HOST="$(hostname)"; \
WHEN="$(date -Is)"; \
TEXT="$UNIT failed on $HOST at $WHEN"; \
printf "<3>%s\\n" "$TEXT"; \
if [ -n "$NTFY_URL" ]; then \
/usr/bin/curl -fsS --max-time 10 \
-H "Title: cashumints: $UNIT failed" \
-H "Priority: high" \
-H "Tags: rotating_light" \
-d "$TEXT" "$NTFY_URL" >/dev/null \
|| printf "<3>%s\\n" "alert: POST to NTFY_URL failed"; \
fi; \
if [ -n "$WEBHOOK_URL" ]; then \
/usr/bin/curl -fsS --max-time 10 \
-H "Content-Type: application/json" \
-d "{\\"unit\\":\\"$UNIT\\",\\"host\\":\\"$HOST\\",\\"at\\":\\"$WHEN\\",\\"text\\":\\"$TEXT\\",\\"content\\":\\"$TEXT\\"}" \
"$WEBHOOK_URL" >/dev/null \
|| printf "<3>%s\\n" "alert: POST to WEBHOOK_URL failed"; \
fi; \
if [ -z "$NTFY_URL" ] && [ -z "$WEBHOOK_URL" ]; then \
printf "<4>%s\\n" "alert: no NTFY_URL or WEBHOOK_URL in /etc/cashumints/alert.env, journal only"; \
fi; \
true' _ %i
# It sends one HTTP request and writes one line. It needs no identity of its own, and
# DynamicUser gives it a throwaway one rather than sharing `nobody` with everything else
# on the host that also could not be bothered to make a user.
DynamicUser=yes
NoNewPrivileges=true
PrivateDevices=true
ProtectSystem=strict
ProtectHome=true
ProtectKernelTunables=true
ProtectKernelModules=true
ProtectControlGroups=true
RestrictAddressFamilies=AF_UNIX AF_INET AF_INET6
RestrictSUIDSGID=true
LockPersonality=true
# An alert that cannot reach the network in ten seconds is not worth a stuck job.
TimeoutStartSec=30
+11 -5
View File
@@ -12,12 +12,18 @@
Description=cashumints.space static site server Description=cashumints.space static site server
Wants=network-online.target Wants=network-online.target
After=network-online.target After=network-online.target
# Give up after five failures in a minute instead of restarting forever. A process that # Give up after five failures in two minutes instead of restarting forever. A process
# cannot start will not start on the 4000th attempt either, and `failed` in # that cannot start will not start on the 4000th attempt either, and `failed` in
# `systemctl status` is a far louder signal than a journal scrolling past. These two are # `systemctl status` is a far louder signal than a journal scrolling past. The window
# [Unit] keys; systemd ignores them under [Service] with only a warning. # matches cashumints.service; see the note there for why it is 120s and not 60s. These
StartLimitIntervalSec=60 # two are [Unit] keys; systemd ignores them under [Service] with only a warning.
StartLimitIntervalSec=120
StartLimitBurst=5 StartLimitBurst=5
# Carry a failure off the machine. `%n` is this unit's own name, so the alert says
# which one died. cashumints-alert@.service writes to the journal at ERROR always and
# curls NTFY_URL or WEBHOOK_URL from /etc/cashumints/alert.env when either is set.
OnFailure=cashumints-alert@%n.service
# Not Requires=cashumints.service: the pages are prerendered, so the site keeps serving # Not Requires=cashumints.service: the pages are prerendered, so the site keeps serving
# a correct-as-of-last-build copy while the API is down. Only the islands go quiet. # a correct-as-of-last-build copy while the API is down. Only the islands go quiet.
+5
View File
@@ -22,6 +22,11 @@ Requires=cashumints.service
After=cashumints.service network-online.target After=cashumints.service network-online.target
Wants=network-online.target Wants=network-online.target
# Carry a failure off the machine. `%n` is this unit's own name, so the alert says
# which one died. cashumints-alert@.service writes to the journal at ERROR always and
# curls NTFY_URL or WEBHOOK_URL from /etc/cashumints/alert.env when either is set.
OnFailure=cashumints-alert@%n.service
[Service] [Service]
Type=oneshot Type=oneshot
User=cashumints User=cashumints
+19 -5
View File
@@ -3,13 +3,27 @@
Description=cashumints.space indexer and API Description=cashumints.space indexer and API
Wants=network-online.target Wants=network-online.target
After=network-online.target After=network-online.target
# Stop after five failures in a minute rather than restarting forever. A wrong Node on # Stop after five failures in two minutes rather than restarting forever.
# PATH once produced four thousand identical crashes in the journal before anyone read #
# one of them; `failed` in `systemctl status` says the same thing in one line. Both keys # The window is 120s and not 60s because RestartSec=5s below means five attempts take
# belong to [Unit] — under [Service] systemd only warns and ignores them. # a little over twenty seconds of restarts plus however long each attempt lives before
StartLimitIntervalSec=60 # it dies. A process that fails *slowly* — a database that times out, a port that takes
# four seconds to refuse — can spread five failures past a sixty second window and reset
# the counter forever, which is the loop this is supposed to stop. 120s covers that.
#
# The failure this exists for: ExecStart named a .ts file, /usr/bin/node was 20, and
# every start died in under a second. 464 restarts over fifteen hours, and because
# Restart=on-failure without a start limit never reaches a `failed` state, nothing
# anywhere went red. Both keys belong to [Unit] — under [Service] systemd only warns and
# ignores them.
StartLimitIntervalSec=120
StartLimitBurst=5 StartLimitBurst=5
# Carry a failure off the machine. `%n` is this unit's own name, so the alert says
# which one died. cashumints-alert@.service writes to the journal at ERROR always and
# curls NTFY_URL or WEBHOOK_URL from /etc/cashumints/alert.env when either is set.
OnFailure=cashumints-alert@%n.service
[Service] [Service]
Type=simple Type=simple
User=cashumints User=cashumints