Files
CashuMints.space/deploy/cashumints-alert@.service
T
michilisandClaude Opus 5 0ebc8ada54 Make a crash loop reach somebody instead of scrolling past.
`Restart=on-failure` with no start limit is an infinite loop by definition: the
unit never reaches `failed`, `systemctl status` stays active (auto-restart), and
the only evidence is a journal moving at four lines a second. That is how 464
restarts over fifteen hours went unnoticed.

All three units now stop after five failures in 120s and run
OnFailure=cashumints-alert@%n.service. The window is 120s and not 60s because
RestartSec=5s plus a process that takes a few seconds to die can spread five
failures past a sixty second window, reset the counter, and loop forever anyway.

cashumints-alert@.service is a oneshot that takes the failed unit's name as its
instance. Configuration is /etc/cashumints/alert.env: NTFY_URL gets a plain-text
body, WEBHOOK_URL gets JSON carrying `content` so one payload fits Discord and
Slack-compatible endpoints. With neither set — or the file absent — it still
writes to the journal at ERROR via a `<3>` syslog prefix, so `journalctl -p err -t
cashumints-alert` is a complete history on a host nobody configured.

It cannot become a second thing to debug: each curl is bounded at 10s, each
failure falls back to a journal line, and the shell ends in `true`, so the alerter
always exits 0. Verified with systemd-analyze verify and by running the ExecStart
body against a local sink — the JSON parses, and every branch exits 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 16:10:01 +02:00

102 lines
4.7 KiB
Desktop File

# /etc/systemd/system/cashumints-alert@.service
#
# The unit that makes a failure audible.
#
# The other three units each carry `OnFailure=cashumints-alert@%n.service`, so systemd
# starts one of these with the failed unit's name as the instance — `%i` below is
# literally `cashumints.service`, `cashumints-web.service` or `cashumints-site.service`.
#
# Why it exists: the API once crash looped 464 times over fifteen hours and nothing said
# so. `Restart=on-failure` with no start limit is an infinite loop that never reaches a
# `failed` state, so the journal filled with identical lines nobody was reading and
# every signal stayed green. The other half of the fix is StartLimitBurst= in each unit,
# which turns the loop into a failure; this is what carries that failure off the machine.
#
# Install:
# sudo install -m 0644 deploy/cashumints-alert@.service /etc/systemd/system/
# sudo install -d -m 0755 /etc/cashumints
# sudo install -m 0640 -o root -g root deploy/alert.env.example /etc/cashumints/alert.env
# sudo systemctl daemon-reload
#
# No [Install] section and never enabled: OnFailure= starts it, and a unit that also
# started at boot would page on every reboot.
[Unit]
Description=Notify that %i failed
# No OnFailure= here. An alerter that alerts about its own failure is a loop, and this
# one is written so its worst case is a journal line rather than a retry.
[Service]
Type=oneshot
# The one file an operator edits, and the only reason this unit is configurable at all.
# Absent is a supported state — the leading `-` says so — and then the ExecStart below
# still writes to the journal at ERROR, which is what `systemctl status` and
# `journalctl -p err` read. See alert.env.example.
EnvironmentFile=-/etc/cashumints/alert.env
# So `journalctl -t cashumints-alert` finds every alert, whichever unit triggered it.
SyslogIdentifier=cashumints-alert
# Everything is inside one shell so the "nothing configured" branch is reachable without
# a second unit. The pieces, in order:
#
# - `printf '<3>…'` on stdout. systemd reads that syslog prefix off a journal stream
# and files the line at priority 3, ERROR, so `journalctl -p err` is a complete
# history of failures on a host with no webhook configured at all. `<4>` is warning.
# A prefix rather than systemd-cat, so the unit needs nothing from the filesystem it
# has just sandboxed itself away from.
# - NTFY_URL is a topic URL (https://ntfy.sh/your-topic). It gets a plain-text body
# naming the failed unit, plus the header names ntfy understands.
# - WEBHOOK_URL gets a JSON POST instead, for Discord, Slack or anything that speaks
# `{"content": …}` — every key is sent, so one payload fits all of them.
# - `--max-time 10` and a `||` fallback on each: an alert that hangs would hold the
# failed unit's job open, and an alert that fails must not itself become a second
# failed unit for somebody to notice. The shell ends in `true` for the same reason.
#
# `%i` is the failed unit's name, passed as an argument rather than interpolated into
# the shell text: systemd expands specifiers before /bin/sh ever sees the line, and a
# unit name is not a thing to trust to quoting.
ExecStart=/bin/sh -c '\
UNIT="$1"; \
HOST="$(hostname)"; \
WHEN="$(date -Is)"; \
TEXT="$UNIT failed on $HOST at $WHEN"; \
printf "<3>%s\\n" "$TEXT"; \
if [ -n "$NTFY_URL" ]; then \
/usr/bin/curl -fsS --max-time 10 \
-H "Title: cashumints: $UNIT failed" \
-H "Priority: high" \
-H "Tags: rotating_light" \
-d "$TEXT" "$NTFY_URL" >/dev/null \
|| printf "<3>%s\\n" "alert: POST to NTFY_URL failed"; \
fi; \
if [ -n "$WEBHOOK_URL" ]; then \
/usr/bin/curl -fsS --max-time 10 \
-H "Content-Type: application/json" \
-d "{\\"unit\\":\\"$UNIT\\",\\"host\\":\\"$HOST\\",\\"at\\":\\"$WHEN\\",\\"text\\":\\"$TEXT\\",\\"content\\":\\"$TEXT\\"}" \
"$WEBHOOK_URL" >/dev/null \
|| printf "<3>%s\\n" "alert: POST to WEBHOOK_URL failed"; \
fi; \
if [ -z "$NTFY_URL" ] && [ -z "$WEBHOOK_URL" ]; then \
printf "<4>%s\\n" "alert: no NTFY_URL or WEBHOOK_URL in /etc/cashumints/alert.env, journal only"; \
fi; \
true' _ %i
# It sends one HTTP request and writes one line. It needs no identity of its own, and
# DynamicUser gives it a throwaway one rather than sharing `nobody` with everything else
# on the host that also could not be bothered to make a user.
DynamicUser=yes
NoNewPrivileges=true
PrivateDevices=true
ProtectSystem=strict
ProtectHome=true
ProtectKernelTunables=true
ProtectKernelModules=true
ProtectControlGroups=true
RestrictAddressFamilies=AF_UNIX AF_INET AF_INET6
RestrictSUIDSGID=true
LockPersonality=true
# An alert that cannot reach the network in ten seconds is not worth a stuck job.
TimeoutStartSec=30