Dev #8

Merged
Michilis merged 7 commits from dev into main 2026-08-25 14:40:03 +00:00
6 changed files with 347 additions and 15 deletions
Showing only changes of commit 0ebc8ada54 - Show all commits
+188 -5
View File
@@ -933,12 +933,15 @@ laptop requirement, not a server one.
Description=cashumints.space indexer and API
Wants=network-online.target
After=network-online.target
# Stop after five failures in a minute rather than restarting forever: a process that
# cannot start will not start on the 4000th attempt either, and `failed` in
# Stop after five failures in two minutes rather than restarting forever: a process
# that cannot start will not start on the 4000th attempt either, and `failed` in
# `systemctl status` is a louder signal than a journal scrolling past. Both keys belong
# to [Unit] — under [Service] systemd only warns and ignores them.
StartLimitIntervalSec=60
# to [Unit] — under [Service] systemd only warns and ignores them. See "Failing loudly"
# for why the window is 120s and not 60s.
StartLimitIntervalSec=120
StartLimitBurst=5
# And carry that `failed` off the machine. %n is this unit's own name.
OnFailure=cashumints-alert@%n.service
[Service]
Type=simple
@@ -1009,8 +1012,9 @@ Two behaviours are worth knowing about because they are load-bearing:
Description=cashumints.space static site server
Wants=network-online.target
After=network-online.target
StartLimitIntervalSec=60
StartLimitIntervalSec=120
StartLimitBurst=5
OnFailure=cashumints-alert@%n.service
# Not Requires=cashumints.service: the pages are prerendered, so the site keeps serving
# a correct-as-of-last-build copy while the API is down. Only the islands go quiet.
@@ -1175,6 +1179,184 @@ nginx 1.25 and later want `http2 on;` on its own line and warn about the `listen
form above; Debian 12 ships 1.22, where the newer form is an unknown directive. The form
above is the one that works on both.
### Failing loudly
Two separate silences produced today's incident, and they need separate fixes.
The first is a **crash loop that never reports a failure**. `Restart=on-failure` with no
start limit is an infinite loop by definition: systemd restarts, the process dies,
systemd restarts. The unit never reaches `failed`, so `systemctl status` stays `active
(auto-restart)`, nothing sends anything anywhere, and the only evidence is a journal
scrolling past at four lines a second. The API did this 464 times over fifteen hours.
All three units now carry:
```ini
StartLimitIntervalSec=120
StartLimitBurst=5
OnFailure=cashumints-alert@%n.service
```
**Why 120 and not 60.** `RestartSec=5s` means five attempts cost a little over twenty
seconds of waiting, plus however long each attempt survives before dying. A process that
fails *slowly* — a database connection that times out, a port that takes four seconds to
refuse — spreads five failures past a sixty second window, resets the counter, and loops
forever anyway. 120s covers the slow case. Both keys belong under `[Unit]`; systemd
takes them under `[Service]` with only a warning and then ignores them.
**Why `OnFailure=` at all.** `StartLimitBurst` turns the loop into a `failed` state,
which is much better, and is still a state somebody has to go and look at. `OnFailure=`
is what makes the machine speak first. `%n` expands to the failed unit's own name, which
arrives as the template instance in `%i`.
The second silence is a **build that publishes an index it should have refused**; that
one is the `MIN_MINTS_FOR_BUILD` gate under "Rebuilds", and the `discovery_starved` flag
under "Discovery starvation" is what feeds it.
#### The alert unit
```ini
# /etc/systemd/system/cashumints-alert@.service
#
# The unit that makes a failure audible.
#
# The other three units each carry `OnFailure=cashumints-alert@%n.service`, so systemd
# starts one of these with the failed unit's name as the instance — `%i` below is
# literally `cashumints.service`, `cashumints-web.service` or `cashumints-site.service`.
#
# Why it exists: the API once crash looped 464 times over fifteen hours and nothing said
# so. `Restart=on-failure` with no start limit is an infinite loop that never reaches a
# `failed` state, so the journal filled with identical lines nobody was reading and
# every signal stayed green. The other half of the fix is StartLimitBurst= in each unit,
# which turns the loop into a failure; this is what carries that failure off the machine.
#
# Install:
# sudo install -m 0644 deploy/cashumints-alert@.service /etc/systemd/system/
# sudo install -d -m 0755 /etc/cashumints
# sudo install -m 0640 -o root -g root deploy/alert.env.example /etc/cashumints/alert.env
# sudo systemctl daemon-reload
#
# No [Install] section and never enabled: OnFailure= starts it, and a unit that also
# started at boot would page on every reboot.
[Unit]
Description=Notify that %i failed
# No OnFailure= here. An alerter that alerts about its own failure is a loop, and this
# one is written so its worst case is a journal line rather than a retry.
[Service]
Type=oneshot
# The one file an operator edits, and the only reason this unit is configurable at all.
# Absent is a supported state — the leading `-` says so — and then the ExecStart below
# still writes to the journal at ERROR, which is what `systemctl status` and
# `journalctl -p err` read. See alert.env.example.
EnvironmentFile=-/etc/cashumints/alert.env
# So `journalctl -t cashumints-alert` finds every alert, whichever unit triggered it.
SyslogIdentifier=cashumints-alert
# Everything is inside one shell so the "nothing configured" branch is reachable without
# a second unit. The pieces, in order:
#
# - `printf '<3>…'` on stdout. systemd reads that syslog prefix off a journal stream
# and files the line at priority 3, ERROR, so `journalctl -p err` is a complete
# history of failures on a host with no webhook configured at all. `<4>` is warning.
# A prefix rather than systemd-cat, so the unit needs nothing from the filesystem it
# has just sandboxed itself away from.
# - NTFY_URL is a topic URL (https://ntfy.sh/your-topic). It gets a plain-text body
# naming the failed unit, plus the header names ntfy understands.
# - WEBHOOK_URL gets a JSON POST instead, for Discord, Slack or anything that speaks
# `{"content": …}` — every key is sent, so one payload fits all of them.
# - `--max-time 10` and a `||` fallback on each: an alert that hangs would hold the
# failed unit's job open, and an alert that fails must not itself become a second
# failed unit for somebody to notice. The shell ends in `true` for the same reason.
#
# `%i` is the failed unit's name, passed as an argument rather than interpolated into
# the shell text: systemd expands specifiers before /bin/sh ever sees the line, and a
# unit name is not a thing to trust to quoting.
ExecStart=/bin/sh -c '\
UNIT="$1"; \
HOST="$(hostname)"; \
WHEN="$(date -Is)"; \
TEXT="$UNIT failed on $HOST at $WHEN"; \
printf "<3>%s\\n" "$TEXT"; \
if [ -n "$NTFY_URL" ]; then \
/usr/bin/curl -fsS --max-time 10 \
-H "Title: cashumints: $UNIT failed" \
-H "Priority: high" \
-H "Tags: rotating_light" \
-d "$TEXT" "$NTFY_URL" >/dev/null \
|| printf "<3>%s\\n" "alert: POST to NTFY_URL failed"; \
fi; \
if [ -n "$WEBHOOK_URL" ]; then \
/usr/bin/curl -fsS --max-time 10 \
-H "Content-Type: application/json" \
-d "{\\"unit\\":\\"$UNIT\\",\\"host\\":\\"$HOST\\",\\"at\\":\\"$WHEN\\",\\"text\\":\\"$TEXT\\",\\"content\\":\\"$TEXT\\"}" \
"$WEBHOOK_URL" >/dev/null \
|| printf "<3>%s\\n" "alert: POST to WEBHOOK_URL failed"; \
fi; \
if [ -z "$NTFY_URL" ] && [ -z "$WEBHOOK_URL" ]; then \
printf "<4>%s\\n" "alert: no NTFY_URL or WEBHOOK_URL in /etc/cashumints/alert.env, journal only"; \
fi; \
true' _ %i
# It sends one HTTP request and writes one line. It needs no identity of its own, and
# DynamicUser gives it a throwaway one rather than sharing `nobody` with everything else
# on the host that also could not be bothered to make a user.
DynamicUser=yes
NoNewPrivileges=true
PrivateDevices=true
ProtectSystem=strict
ProtectHome=true
ProtectKernelTunables=true
ProtectKernelModules=true
ProtectControlGroups=true
RestrictAddressFamilies=AF_UNIX AF_INET AF_INET6
RestrictSUIDSGID=true
LockPersonality=true
# An alert that cannot reach the network in ten seconds is not worth a stuck job.
TimeoutStartSec=30
```
Configuration is one optional file. With it absent, or with both values empty, a failure
still lands in the journal at `ERROR` and is readable with `journalctl -p err -t
cashumints-alert`; the unit is written so that its worst case is a log line rather than
a second thing to debug.
```ini
# /etc/cashumints/alert.env
#
# Read by cashumints-alert@.service, which systemd starts when any of the three units
# fails. Everything here is optional: with the file absent or both values empty, an
# alert is still written to the journal at ERROR priority and is readable with
#
# journalctl -p err -t cashumints-alert
#
# Set one or both to have failures leave the machine.
#
# Install it root-owned and not world-readable — a webhook URL is a capability:
# sudo install -d -m 0755 /etc/cashumints
# sudo install -m 0640 -o root -g root deploy/alert.env.example /etc/cashumints/alert.env
# sudo systemctl daemon-reload
# An ntfy topic URL. Free and public at ntfy.sh; pick a topic name nobody will guess,
# because anyone who knows it can read and post to it.
#NTFY_URL=https://ntfy.sh/cashumints-alerts-CHANGE-ME
# Anything that accepts a JSON POST. The body carries `unit`, `host`, `at`, `text` and
# `content` — the last of which is what Discord and most Slack-compatible endpoints read,
# so one payload fits all three.
#WEBHOOK_URL=https://discord.com/api/webhooks/…
```
Test it without breaking anything:
```bash
sudo systemctl start 'cashumints-alert@test.service'
journalctl -t cashumints-alert -n 5 --no-pager
```
### Rebuilds
Mint pages are prerendered, so new mints and new review counts appear at the next build.
@@ -1196,6 +1378,7 @@ Description=Rebuild the cashumints.space static site
Requires=cashumints.service
After=cashumints.service network-online.target
Wants=network-online.target
OnFailure=cashumints-alert@%n.service
[Service]
Type=oneshot
+23
View File
@@ -0,0 +1,23 @@
# /etc/cashumints/alert.env
#
# Read by cashumints-alert@.service, which systemd starts when any of the three units
# fails. Everything here is optional: with the file absent or both values empty, an
# alert is still written to the journal at ERROR priority and is readable with
#
# journalctl -p err -t cashumints-alert
#
# Set one or both to have failures leave the machine.
#
# Install it root-owned and not world-readable — a webhook URL is a capability:
# sudo install -d -m 0755 /etc/cashumints
# sudo install -m 0640 -o root -g root deploy/alert.env.example /etc/cashumints/alert.env
# sudo systemctl daemon-reload
# An ntfy topic URL. Free and public at ntfy.sh; pick a topic name nobody will guess,
# because anyone who knows it can read and post to it.
#NTFY_URL=https://ntfy.sh/cashumints-alerts-CHANGE-ME
# Anything that accepts a JSON POST. The body carries `unit`, `host`, `at`, `text` and
# `content` — the last of which is what Discord and most Slack-compatible endpoints read,
# so one payload fits all three.
#WEBHOOK_URL=https://discord.com/api/webhooks/…
+101
View File
@@ -0,0 +1,101 @@
# /etc/systemd/system/cashumints-alert@.service
#
# The unit that makes a failure audible.
#
# The other three units each carry `OnFailure=cashumints-alert@%n.service`, so systemd
# starts one of these with the failed unit's name as the instance — `%i` below is
# literally `cashumints.service`, `cashumints-web.service` or `cashumints-site.service`.
#
# Why it exists: the API once crash looped 464 times over fifteen hours and nothing said
# so. `Restart=on-failure` with no start limit is an infinite loop that never reaches a
# `failed` state, so the journal filled with identical lines nobody was reading and
# every signal stayed green. The other half of the fix is StartLimitBurst= in each unit,
# which turns the loop into a failure; this is what carries that failure off the machine.
#
# Install:
# sudo install -m 0644 deploy/cashumints-alert@.service /etc/systemd/system/
# sudo install -d -m 0755 /etc/cashumints
# sudo install -m 0640 -o root -g root deploy/alert.env.example /etc/cashumints/alert.env
# sudo systemctl daemon-reload
#
# No [Install] section and never enabled: OnFailure= starts it, and a unit that also
# started at boot would page on every reboot.
[Unit]
Description=Notify that %i failed
# No OnFailure= here. An alerter that alerts about its own failure is a loop, and this
# one is written so its worst case is a journal line rather than a retry.
[Service]
Type=oneshot
# The one file an operator edits, and the only reason this unit is configurable at all.
# Absent is a supported state — the leading `-` says so — and then the ExecStart below
# still writes to the journal at ERROR, which is what `systemctl status` and
# `journalctl -p err` read. See alert.env.example.
EnvironmentFile=-/etc/cashumints/alert.env
# So `journalctl -t cashumints-alert` finds every alert, whichever unit triggered it.
SyslogIdentifier=cashumints-alert
# Everything is inside one shell so the "nothing configured" branch is reachable without
# a second unit. The pieces, in order:
#
# - `printf '<3>…'` on stdout. systemd reads that syslog prefix off a journal stream
# and files the line at priority 3, ERROR, so `journalctl -p err` is a complete
# history of failures on a host with no webhook configured at all. `<4>` is warning.
# A prefix rather than systemd-cat, so the unit needs nothing from the filesystem it
# has just sandboxed itself away from.
# - NTFY_URL is a topic URL (https://ntfy.sh/your-topic). It gets a plain-text body
# naming the failed unit, plus the header names ntfy understands.
# - WEBHOOK_URL gets a JSON POST instead, for Discord, Slack or anything that speaks
# `{"content": …}` — every key is sent, so one payload fits all of them.
# - `--max-time 10` and a `||` fallback on each: an alert that hangs would hold the
# failed unit's job open, and an alert that fails must not itself become a second
# failed unit for somebody to notice. The shell ends in `true` for the same reason.
#
# `%i` is the failed unit's name, passed as an argument rather than interpolated into
# the shell text: systemd expands specifiers before /bin/sh ever sees the line, and a
# unit name is not a thing to trust to quoting.
ExecStart=/bin/sh -c '\
UNIT="$1"; \
HOST="$(hostname)"; \
WHEN="$(date -Is)"; \
TEXT="$UNIT failed on $HOST at $WHEN"; \
printf "<3>%s\\n" "$TEXT"; \
if [ -n "$NTFY_URL" ]; then \
/usr/bin/curl -fsS --max-time 10 \
-H "Title: cashumints: $UNIT failed" \
-H "Priority: high" \
-H "Tags: rotating_light" \
-d "$TEXT" "$NTFY_URL" >/dev/null \
|| printf "<3>%s\\n" "alert: POST to NTFY_URL failed"; \
fi; \
if [ -n "$WEBHOOK_URL" ]; then \
/usr/bin/curl -fsS --max-time 10 \
-H "Content-Type: application/json" \
-d "{\\"unit\\":\\"$UNIT\\",\\"host\\":\\"$HOST\\",\\"at\\":\\"$WHEN\\",\\"text\\":\\"$TEXT\\",\\"content\\":\\"$TEXT\\"}" \
"$WEBHOOK_URL" >/dev/null \
|| printf "<3>%s\\n" "alert: POST to WEBHOOK_URL failed"; \
fi; \
if [ -z "$NTFY_URL" ] && [ -z "$WEBHOOK_URL" ]; then \
printf "<4>%s\\n" "alert: no NTFY_URL or WEBHOOK_URL in /etc/cashumints/alert.env, journal only"; \
fi; \
true' _ %i
# It sends one HTTP request and writes one line. It needs no identity of its own, and
# DynamicUser gives it a throwaway one rather than sharing `nobody` with everything else
# on the host that also could not be bothered to make a user.
DynamicUser=yes
NoNewPrivileges=true
PrivateDevices=true
ProtectSystem=strict
ProtectHome=true
ProtectKernelTunables=true
ProtectKernelModules=true
ProtectControlGroups=true
RestrictAddressFamilies=AF_UNIX AF_INET AF_INET6
RestrictSUIDSGID=true
LockPersonality=true
# An alert that cannot reach the network in ten seconds is not worth a stuck job.
TimeoutStartSec=30
+11 -5
View File
@@ -12,12 +12,18 @@
Description=cashumints.space static site server
Wants=network-online.target
After=network-online.target
# Give up after five failures in a minute instead of restarting forever. A process that
# cannot start will not start on the 4000th attempt either, and `failed` in
# `systemctl status` is a far louder signal than a journal scrolling past. These two are
# [Unit] keys; systemd ignores them under [Service] with only a warning.
StartLimitIntervalSec=60
# Give up after five failures in two minutes instead of restarting forever. A process
# that cannot start will not start on the 4000th attempt either, and `failed` in
# `systemctl status` is a far louder signal than a journal scrolling past. The window
# matches cashumints.service; see the note there for why it is 120s and not 60s. These
# two are [Unit] keys; systemd ignores them under [Service] with only a warning.
StartLimitIntervalSec=120
StartLimitBurst=5
# Carry a failure off the machine. `%n` is this unit's own name, so the alert says
# which one died. cashumints-alert@.service writes to the journal at ERROR always and
# curls NTFY_URL or WEBHOOK_URL from /etc/cashumints/alert.env when either is set.
OnFailure=cashumints-alert@%n.service
# Not Requires=cashumints.service: the pages are prerendered, so the site keeps serving
# a correct-as-of-last-build copy while the API is down. Only the islands go quiet.
+5
View File
@@ -22,6 +22,11 @@ Requires=cashumints.service
After=cashumints.service network-online.target
Wants=network-online.target
# Carry a failure off the machine. `%n` is this unit's own name, so the alert says
# which one died. cashumints-alert@.service writes to the journal at ERROR always and
# curls NTFY_URL or WEBHOOK_URL from /etc/cashumints/alert.env when either is set.
OnFailure=cashumints-alert@%n.service
[Service]
Type=oneshot
User=cashumints
+19 -5
View File
@@ -3,13 +3,27 @@
Description=cashumints.space indexer and API
Wants=network-online.target
After=network-online.target
# Stop after five failures in a minute rather than restarting forever. A wrong Node on
# PATH once produced four thousand identical crashes in the journal before anyone read
# one of them; `failed` in `systemctl status` says the same thing in one line. Both keys
# belong to [Unit] — under [Service] systemd only warns and ignores them.
StartLimitIntervalSec=60
# Stop after five failures in two minutes rather than restarting forever.
#
# The window is 120s and not 60s because RestartSec=5s below means five attempts take
# a little over twenty seconds of restarts plus however long each attempt lives before
# it dies. A process that fails *slowly* — a database that times out, a port that takes
# four seconds to refuse — can spread five failures past a sixty second window and reset
# the counter forever, which is the loop this is supposed to stop. 120s covers that.
#
# The failure this exists for: ExecStart named a .ts file, /usr/bin/node was 20, and
# every start died in under a second. 464 restarts over fifteen hours, and because
# Restart=on-failure without a start limit never reaches a `failed` state, nothing
# anywhere went red. Both keys belong to [Unit] — under [Service] systemd only warns and
# ignores them.
StartLimitIntervalSec=120
StartLimitBurst=5
# Carry a failure off the machine. `%n` is this unit's own name, so the alert says
# which one died. cashumints-alert@.service writes to the journal at ERROR always and
# curls NTFY_URL or WEBHOOK_URL from /etc/cashumints/alert.env when either is set.
OnFailure=cashumints-alert@%n.service
[Service]
Type=simple
User=cashumints