diff --git a/README.md b/README.md index 0b66e34..f29dbc3 100644 --- a/README.md +++ b/README.md @@ -933,12 +933,15 @@ laptop requirement, not a server one. Description=cashumints.space indexer and API Wants=network-online.target After=network-online.target -# Stop after five failures in a minute rather than restarting forever: a process that -# cannot start will not start on the 4000th attempt either, and `failed` in +# Stop after five failures in two minutes rather than restarting forever: a process +# that cannot start will not start on the 4000th attempt either, and `failed` in # `systemctl status` is a louder signal than a journal scrolling past. Both keys belong -# to [Unit] — under [Service] systemd only warns and ignores them. -StartLimitIntervalSec=60 +# to [Unit] — under [Service] systemd only warns and ignores them. See "Failing loudly" +# for why the window is 120s and not 60s. +StartLimitIntervalSec=120 StartLimitBurst=5 +# And carry that `failed` off the machine. %n is this unit's own name. +OnFailure=cashumints-alert@%n.service [Service] Type=simple @@ -1009,8 +1012,9 @@ Two behaviours are worth knowing about because they are load-bearing: Description=cashumints.space static site server Wants=network-online.target After=network-online.target -StartLimitIntervalSec=60 +StartLimitIntervalSec=120 StartLimitBurst=5 +OnFailure=cashumints-alert@%n.service # Not Requires=cashumints.service: the pages are prerendered, so the site keeps serving # a correct-as-of-last-build copy while the API is down. Only the islands go quiet. @@ -1175,6 +1179,184 @@ nginx 1.25 and later want `http2 on;` on its own line and warn about the `listen form above; Debian 12 ships 1.22, where the newer form is an unknown directive. The form above is the one that works on both. +### Failing loudly + +Two separate silences produced today's incident, and they need separate fixes. + +The first is a **crash loop that never reports a failure**. `Restart=on-failure` with no +start limit is an infinite loop by definition: systemd restarts, the process dies, +systemd restarts. The unit never reaches `failed`, so `systemctl status` stays `active +(auto-restart)`, nothing sends anything anywhere, and the only evidence is a journal +scrolling past at four lines a second. The API did this 464 times over fifteen hours. + +All three units now carry: + +```ini +StartLimitIntervalSec=120 +StartLimitBurst=5 +OnFailure=cashumints-alert@%n.service +``` + +**Why 120 and not 60.** `RestartSec=5s` means five attempts cost a little over twenty +seconds of waiting, plus however long each attempt survives before dying. A process that +fails *slowly* — a database connection that times out, a port that takes four seconds to +refuse — spreads five failures past a sixty second window, resets the counter, and loops +forever anyway. 120s covers the slow case. Both keys belong under `[Unit]`; systemd +takes them under `[Service]` with only a warning and then ignores them. + +**Why `OnFailure=` at all.** `StartLimitBurst` turns the loop into a `failed` state, +which is much better, and is still a state somebody has to go and look at. `OnFailure=` +is what makes the machine speak first. `%n` expands to the failed unit's own name, which +arrives as the template instance in `%i`. + +The second silence is a **build that publishes an index it should have refused**; that +one is the `MIN_MINTS_FOR_BUILD` gate under "Rebuilds", and the `discovery_starved` flag +under "Discovery starvation" is what feeds it. + +#### The alert unit + +```ini +# /etc/systemd/system/cashumints-alert@.service +# +# The unit that makes a failure audible. +# +# The other three units each carry `OnFailure=cashumints-alert@%n.service`, so systemd +# starts one of these with the failed unit's name as the instance — `%i` below is +# literally `cashumints.service`, `cashumints-web.service` or `cashumints-site.service`. +# +# Why it exists: the API once crash looped 464 times over fifteen hours and nothing said +# so. `Restart=on-failure` with no start limit is an infinite loop that never reaches a +# `failed` state, so the journal filled with identical lines nobody was reading and +# every signal stayed green. The other half of the fix is StartLimitBurst= in each unit, +# which turns the loop into a failure; this is what carries that failure off the machine. +# +# Install: +# sudo install -m 0644 deploy/cashumints-alert@.service /etc/systemd/system/ +# sudo install -d -m 0755 /etc/cashumints +# sudo install -m 0640 -o root -g root deploy/alert.env.example /etc/cashumints/alert.env +# sudo systemctl daemon-reload +# +# No [Install] section and never enabled: OnFailure= starts it, and a unit that also +# started at boot would page on every reboot. + +[Unit] +Description=Notify that %i failed +# No OnFailure= here. An alerter that alerts about its own failure is a loop, and this +# one is written so its worst case is a journal line rather than a retry. + +[Service] +Type=oneshot + +# The one file an operator edits, and the only reason this unit is configurable at all. +# Absent is a supported state — the leading `-` says so — and then the ExecStart below +# still writes to the journal at ERROR, which is what `systemctl status` and +# `journalctl -p err` read. See alert.env.example. +EnvironmentFile=-/etc/cashumints/alert.env + +# So `journalctl -t cashumints-alert` finds every alert, whichever unit triggered it. +SyslogIdentifier=cashumints-alert + +# Everything is inside one shell so the "nothing configured" branch is reachable without +# a second unit. The pieces, in order: +# +# - `printf '<3>…'` on stdout. systemd reads that syslog prefix off a journal stream +# and files the line at priority 3, ERROR, so `journalctl -p err` is a complete +# history of failures on a host with no webhook configured at all. `<4>` is warning. +# A prefix rather than systemd-cat, so the unit needs nothing from the filesystem it +# has just sandboxed itself away from. +# - NTFY_URL is a topic URL (https://ntfy.sh/your-topic). It gets a plain-text body +# naming the failed unit, plus the header names ntfy understands. +# - WEBHOOK_URL gets a JSON POST instead, for Discord, Slack or anything that speaks +# `{"content": …}` — every key is sent, so one payload fits all of them. +# - `--max-time 10` and a `||` fallback on each: an alert that hangs would hold the +# failed unit's job open, and an alert that fails must not itself become a second +# failed unit for somebody to notice. The shell ends in `true` for the same reason. +# +# `%i` is the failed unit's name, passed as an argument rather than interpolated into +# the shell text: systemd expands specifiers before /bin/sh ever sees the line, and a +# unit name is not a thing to trust to quoting. +ExecStart=/bin/sh -c '\ + UNIT="$1"; \ + HOST="$(hostname)"; \ + WHEN="$(date -Is)"; \ + TEXT="$UNIT failed on $HOST at $WHEN"; \ + printf "<3>%s\\n" "$TEXT"; \ + if [ -n "$NTFY_URL" ]; then \ + /usr/bin/curl -fsS --max-time 10 \ + -H "Title: cashumints: $UNIT failed" \ + -H "Priority: high" \ + -H "Tags: rotating_light" \ + -d "$TEXT" "$NTFY_URL" >/dev/null \ + || printf "<3>%s\\n" "alert: POST to NTFY_URL failed"; \ + fi; \ + if [ -n "$WEBHOOK_URL" ]; then \ + /usr/bin/curl -fsS --max-time 10 \ + -H "Content-Type: application/json" \ + -d "{\\"unit\\":\\"$UNIT\\",\\"host\\":\\"$HOST\\",\\"at\\":\\"$WHEN\\",\\"text\\":\\"$TEXT\\",\\"content\\":\\"$TEXT\\"}" \ + "$WEBHOOK_URL" >/dev/null \ + || printf "<3>%s\\n" "alert: POST to WEBHOOK_URL failed"; \ + fi; \ + if [ -z "$NTFY_URL" ] && [ -z "$WEBHOOK_URL" ]; then \ + printf "<4>%s\\n" "alert: no NTFY_URL or WEBHOOK_URL in /etc/cashumints/alert.env, journal only"; \ + fi; \ + true' _ %i + +# It sends one HTTP request and writes one line. It needs no identity of its own, and +# DynamicUser gives it a throwaway one rather than sharing `nobody` with everything else +# on the host that also could not be bothered to make a user. +DynamicUser=yes +NoNewPrivileges=true +PrivateDevices=true +ProtectSystem=strict +ProtectHome=true +ProtectKernelTunables=true +ProtectKernelModules=true +ProtectControlGroups=true +RestrictAddressFamilies=AF_UNIX AF_INET AF_INET6 +RestrictSUIDSGID=true +LockPersonality=true +# An alert that cannot reach the network in ten seconds is not worth a stuck job. +TimeoutStartSec=30 +``` + +Configuration is one optional file. With it absent, or with both values empty, a failure +still lands in the journal at `ERROR` and is readable with `journalctl -p err -t +cashumints-alert`; the unit is written so that its worst case is a log line rather than +a second thing to debug. + +```ini +# /etc/cashumints/alert.env +# +# Read by cashumints-alert@.service, which systemd starts when any of the three units +# fails. Everything here is optional: with the file absent or both values empty, an +# alert is still written to the journal at ERROR priority and is readable with +# +# journalctl -p err -t cashumints-alert +# +# Set one or both to have failures leave the machine. +# +# Install it root-owned and not world-readable — a webhook URL is a capability: +# sudo install -d -m 0755 /etc/cashumints +# sudo install -m 0640 -o root -g root deploy/alert.env.example /etc/cashumints/alert.env +# sudo systemctl daemon-reload + +# An ntfy topic URL. Free and public at ntfy.sh; pick a topic name nobody will guess, +# because anyone who knows it can read and post to it. +#NTFY_URL=https://ntfy.sh/cashumints-alerts-CHANGE-ME + +# Anything that accepts a JSON POST. The body carries `unit`, `host`, `at`, `text` and +# `content` — the last of which is what Discord and most Slack-compatible endpoints read, +# so one payload fits all three. +#WEBHOOK_URL=https://discord.com/api/webhooks/… +``` + +Test it without breaking anything: + +```bash +sudo systemctl start 'cashumints-alert@test.service' +journalctl -t cashumints-alert -n 5 --no-pager +``` + ### Rebuilds Mint pages are prerendered, so new mints and new review counts appear at the next build. @@ -1196,6 +1378,7 @@ Description=Rebuild the cashumints.space static site Requires=cashumints.service After=cashumints.service network-online.target Wants=network-online.target +OnFailure=cashumints-alert@%n.service [Service] Type=oneshot diff --git a/deploy/alert.env.example b/deploy/alert.env.example new file mode 100644 index 0000000..3f6b638 --- /dev/null +++ b/deploy/alert.env.example @@ -0,0 +1,23 @@ +# /etc/cashumints/alert.env +# +# Read by cashumints-alert@.service, which systemd starts when any of the three units +# fails. Everything here is optional: with the file absent or both values empty, an +# alert is still written to the journal at ERROR priority and is readable with +# +# journalctl -p err -t cashumints-alert +# +# Set one or both to have failures leave the machine. +# +# Install it root-owned and not world-readable — a webhook URL is a capability: +# sudo install -d -m 0755 /etc/cashumints +# sudo install -m 0640 -o root -g root deploy/alert.env.example /etc/cashumints/alert.env +# sudo systemctl daemon-reload + +# An ntfy topic URL. Free and public at ntfy.sh; pick a topic name nobody will guess, +# because anyone who knows it can read and post to it. +#NTFY_URL=https://ntfy.sh/cashumints-alerts-CHANGE-ME + +# Anything that accepts a JSON POST. The body carries `unit`, `host`, `at`, `text` and +# `content` — the last of which is what Discord and most Slack-compatible endpoints read, +# so one payload fits all three. +#WEBHOOK_URL=https://discord.com/api/webhooks/… diff --git a/deploy/cashumints-alert@.service b/deploy/cashumints-alert@.service new file mode 100644 index 0000000..fb9a1bd --- /dev/null +++ b/deploy/cashumints-alert@.service @@ -0,0 +1,101 @@ +# /etc/systemd/system/cashumints-alert@.service +# +# The unit that makes a failure audible. +# +# The other three units each carry `OnFailure=cashumints-alert@%n.service`, so systemd +# starts one of these with the failed unit's name as the instance — `%i` below is +# literally `cashumints.service`, `cashumints-web.service` or `cashumints-site.service`. +# +# Why it exists: the API once crash looped 464 times over fifteen hours and nothing said +# so. `Restart=on-failure` with no start limit is an infinite loop that never reaches a +# `failed` state, so the journal filled with identical lines nobody was reading and +# every signal stayed green. The other half of the fix is StartLimitBurst= in each unit, +# which turns the loop into a failure; this is what carries that failure off the machine. +# +# Install: +# sudo install -m 0644 deploy/cashumints-alert@.service /etc/systemd/system/ +# sudo install -d -m 0755 /etc/cashumints +# sudo install -m 0640 -o root -g root deploy/alert.env.example /etc/cashumints/alert.env +# sudo systemctl daemon-reload +# +# No [Install] section and never enabled: OnFailure= starts it, and a unit that also +# started at boot would page on every reboot. + +[Unit] +Description=Notify that %i failed +# No OnFailure= here. An alerter that alerts about its own failure is a loop, and this +# one is written so its worst case is a journal line rather than a retry. + +[Service] +Type=oneshot + +# The one file an operator edits, and the only reason this unit is configurable at all. +# Absent is a supported state — the leading `-` says so — and then the ExecStart below +# still writes to the journal at ERROR, which is what `systemctl status` and +# `journalctl -p err` read. See alert.env.example. +EnvironmentFile=-/etc/cashumints/alert.env + +# So `journalctl -t cashumints-alert` finds every alert, whichever unit triggered it. +SyslogIdentifier=cashumints-alert + +# Everything is inside one shell so the "nothing configured" branch is reachable without +# a second unit. The pieces, in order: +# +# - `printf '<3>…'` on stdout. systemd reads that syslog prefix off a journal stream +# and files the line at priority 3, ERROR, so `journalctl -p err` is a complete +# history of failures on a host with no webhook configured at all. `<4>` is warning. +# A prefix rather than systemd-cat, so the unit needs nothing from the filesystem it +# has just sandboxed itself away from. +# - NTFY_URL is a topic URL (https://ntfy.sh/your-topic). It gets a plain-text body +# naming the failed unit, plus the header names ntfy understands. +# - WEBHOOK_URL gets a JSON POST instead, for Discord, Slack or anything that speaks +# `{"content": …}` — every key is sent, so one payload fits all of them. +# - `--max-time 10` and a `||` fallback on each: an alert that hangs would hold the +# failed unit's job open, and an alert that fails must not itself become a second +# failed unit for somebody to notice. The shell ends in `true` for the same reason. +# +# `%i` is the failed unit's name, passed as an argument rather than interpolated into +# the shell text: systemd expands specifiers before /bin/sh ever sees the line, and a +# unit name is not a thing to trust to quoting. +ExecStart=/bin/sh -c '\ + UNIT="$1"; \ + HOST="$(hostname)"; \ + WHEN="$(date -Is)"; \ + TEXT="$UNIT failed on $HOST at $WHEN"; \ + printf "<3>%s\\n" "$TEXT"; \ + if [ -n "$NTFY_URL" ]; then \ + /usr/bin/curl -fsS --max-time 10 \ + -H "Title: cashumints: $UNIT failed" \ + -H "Priority: high" \ + -H "Tags: rotating_light" \ + -d "$TEXT" "$NTFY_URL" >/dev/null \ + || printf "<3>%s\\n" "alert: POST to NTFY_URL failed"; \ + fi; \ + if [ -n "$WEBHOOK_URL" ]; then \ + /usr/bin/curl -fsS --max-time 10 \ + -H "Content-Type: application/json" \ + -d "{\\"unit\\":\\"$UNIT\\",\\"host\\":\\"$HOST\\",\\"at\\":\\"$WHEN\\",\\"text\\":\\"$TEXT\\",\\"content\\":\\"$TEXT\\"}" \ + "$WEBHOOK_URL" >/dev/null \ + || printf "<3>%s\\n" "alert: POST to WEBHOOK_URL failed"; \ + fi; \ + if [ -z "$NTFY_URL" ] && [ -z "$WEBHOOK_URL" ]; then \ + printf "<4>%s\\n" "alert: no NTFY_URL or WEBHOOK_URL in /etc/cashumints/alert.env, journal only"; \ + fi; \ + true' _ %i + +# It sends one HTTP request and writes one line. It needs no identity of its own, and +# DynamicUser gives it a throwaway one rather than sharing `nobody` with everything else +# on the host that also could not be bothered to make a user. +DynamicUser=yes +NoNewPrivileges=true +PrivateDevices=true +ProtectSystem=strict +ProtectHome=true +ProtectKernelTunables=true +ProtectKernelModules=true +ProtectControlGroups=true +RestrictAddressFamilies=AF_UNIX AF_INET AF_INET6 +RestrictSUIDSGID=true +LockPersonality=true +# An alert that cannot reach the network in ten seconds is not worth a stuck job. +TimeoutStartSec=30 diff --git a/deploy/cashumints-site.service b/deploy/cashumints-site.service index 8a8a380..131e321 100644 --- a/deploy/cashumints-site.service +++ b/deploy/cashumints-site.service @@ -12,12 +12,18 @@ Description=cashumints.space static site server Wants=network-online.target After=network-online.target -# Give up after five failures in a minute instead of restarting forever. A process that -# cannot start will not start on the 4000th attempt either, and `failed` in -# `systemctl status` is a far louder signal than a journal scrolling past. These two are -# [Unit] keys; systemd ignores them under [Service] with only a warning. -StartLimitIntervalSec=60 +# Give up after five failures in two minutes instead of restarting forever. A process +# that cannot start will not start on the 4000th attempt either, and `failed` in +# `systemctl status` is a far louder signal than a journal scrolling past. The window +# matches cashumints.service; see the note there for why it is 120s and not 60s. These +# two are [Unit] keys; systemd ignores them under [Service] with only a warning. +StartLimitIntervalSec=120 StartLimitBurst=5 + +# Carry a failure off the machine. `%n` is this unit's own name, so the alert says +# which one died. cashumints-alert@.service writes to the journal at ERROR always and +# curls NTFY_URL or WEBHOOK_URL from /etc/cashumints/alert.env when either is set. +OnFailure=cashumints-alert@%n.service # Not Requires=cashumints.service: the pages are prerendered, so the site keeps serving # a correct-as-of-last-build copy while the API is down. Only the islands go quiet. diff --git a/deploy/cashumints-web.service b/deploy/cashumints-web.service index 3718b0e..efe9020 100644 --- a/deploy/cashumints-web.service +++ b/deploy/cashumints-web.service @@ -22,6 +22,11 @@ Requires=cashumints.service After=cashumints.service network-online.target Wants=network-online.target +# Carry a failure off the machine. `%n` is this unit's own name, so the alert says +# which one died. cashumints-alert@.service writes to the journal at ERROR always and +# curls NTFY_URL or WEBHOOK_URL from /etc/cashumints/alert.env when either is set. +OnFailure=cashumints-alert@%n.service + [Service] Type=oneshot User=cashumints diff --git a/deploy/cashumints.service b/deploy/cashumints.service index 3048edd..495ebb9 100644 --- a/deploy/cashumints.service +++ b/deploy/cashumints.service @@ -3,13 +3,27 @@ Description=cashumints.space indexer and API Wants=network-online.target After=network-online.target -# Stop after five failures in a minute rather than restarting forever. A wrong Node on -# PATH once produced four thousand identical crashes in the journal before anyone read -# one of them; `failed` in `systemctl status` says the same thing in one line. Both keys -# belong to [Unit] — under [Service] systemd only warns and ignores them. -StartLimitIntervalSec=60 +# Stop after five failures in two minutes rather than restarting forever. +# +# The window is 120s and not 60s because RestartSec=5s below means five attempts take +# a little over twenty seconds of restarts plus however long each attempt lives before +# it dies. A process that fails *slowly* — a database that times out, a port that takes +# four seconds to refuse — can spread five failures past a sixty second window and reset +# the counter forever, which is the loop this is supposed to stop. 120s covers that. +# +# The failure this exists for: ExecStart named a .ts file, /usr/bin/node was 20, and +# every start died in under a second. 464 restarts over fifteen hours, and because +# Restart=on-failure without a start limit never reaches a `failed` state, nothing +# anywhere went red. Both keys belong to [Unit] — under [Service] systemd only warns and +# ignores them. +StartLimitIntervalSec=120 StartLimitBurst=5 +# Carry a failure off the machine. `%n` is this unit's own name, so the alert says +# which one died. cashumints-alert@.service writes to the journal at ERROR always and +# curls NTFY_URL or WEBHOOK_URL from /etc/cashumints/alert.env when either is set. +OnFailure=cashumints-alert@%n.service + [Service] Type=simple User=cashumints