Make a crash loop reach somebody instead of scrolling past.
`Restart=on-failure` with no start limit is an infinite loop by definition: the unit never reaches `failed`, `systemctl status` stays active (auto-restart), and the only evidence is a journal moving at four lines a second. That is how 464 restarts over fifteen hours went unnoticed. All three units now stop after five failures in 120s and run OnFailure=cashumints-alert@%n.service. The window is 120s and not 60s because RestartSec=5s plus a process that takes a few seconds to die can spread five failures past a sixty second window, reset the counter, and loop forever anyway. cashumints-alert@.service is a oneshot that takes the failed unit's name as its instance. Configuration is /etc/cashumints/alert.env: NTFY_URL gets a plain-text body, WEBHOOK_URL gets JSON carrying `content` so one payload fits Discord and Slack-compatible endpoints. With neither set — or the file absent — it still writes to the journal at ERROR via a `<3>` syslog prefix, so `journalctl -p err -t cashumints-alert` is a complete history on a host nobody configured. It cannot become a second thing to debug: each curl is bounded at 10s, each failure falls back to a journal line, and the shell ends in `true`, so the alerter always exits 0. Verified with systemd-analyze verify and by running the ExecStart body against a local sink — the JSON parses, and every branch exits 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
9ffa53094d
commit
0ebc8ada54
@@ -933,12 +933,15 @@ laptop requirement, not a server one.
|
||||
Description=cashumints.space indexer and API
|
||||
Wants=network-online.target
|
||||
After=network-online.target
|
||||
# Stop after five failures in a minute rather than restarting forever: a process that
|
||||
# cannot start will not start on the 4000th attempt either, and `failed` in
|
||||
# Stop after five failures in two minutes rather than restarting forever: a process
|
||||
# that cannot start will not start on the 4000th attempt either, and `failed` in
|
||||
# `systemctl status` is a louder signal than a journal scrolling past. Both keys belong
|
||||
# to [Unit] — under [Service] systemd only warns and ignores them.
|
||||
StartLimitIntervalSec=60
|
||||
# to [Unit] — under [Service] systemd only warns and ignores them. See "Failing loudly"
|
||||
# for why the window is 120s and not 60s.
|
||||
StartLimitIntervalSec=120
|
||||
StartLimitBurst=5
|
||||
# And carry that `failed` off the machine. %n is this unit's own name.
|
||||
OnFailure=cashumints-alert@%n.service
|
||||
|
||||
[Service]
|
||||
Type=simple
|
||||
@@ -1009,8 +1012,9 @@ Two behaviours are worth knowing about because they are load-bearing:
|
||||
Description=cashumints.space static site server
|
||||
Wants=network-online.target
|
||||
After=network-online.target
|
||||
StartLimitIntervalSec=60
|
||||
StartLimitIntervalSec=120
|
||||
StartLimitBurst=5
|
||||
OnFailure=cashumints-alert@%n.service
|
||||
# Not Requires=cashumints.service: the pages are prerendered, so the site keeps serving
|
||||
# a correct-as-of-last-build copy while the API is down. Only the islands go quiet.
|
||||
|
||||
@@ -1175,6 +1179,184 @@ nginx 1.25 and later want `http2 on;` on its own line and warn about the `listen
|
||||
form above; Debian 12 ships 1.22, where the newer form is an unknown directive. The form
|
||||
above is the one that works on both.
|
||||
|
||||
### Failing loudly
|
||||
|
||||
Two separate silences produced today's incident, and they need separate fixes.
|
||||
|
||||
The first is a **crash loop that never reports a failure**. `Restart=on-failure` with no
|
||||
start limit is an infinite loop by definition: systemd restarts, the process dies,
|
||||
systemd restarts. The unit never reaches `failed`, so `systemctl status` stays `active
|
||||
(auto-restart)`, nothing sends anything anywhere, and the only evidence is a journal
|
||||
scrolling past at four lines a second. The API did this 464 times over fifteen hours.
|
||||
|
||||
All three units now carry:
|
||||
|
||||
```ini
|
||||
StartLimitIntervalSec=120
|
||||
StartLimitBurst=5
|
||||
OnFailure=cashumints-alert@%n.service
|
||||
```
|
||||
|
||||
**Why 120 and not 60.** `RestartSec=5s` means five attempts cost a little over twenty
|
||||
seconds of waiting, plus however long each attempt survives before dying. A process that
|
||||
fails *slowly* — a database connection that times out, a port that takes four seconds to
|
||||
refuse — spreads five failures past a sixty second window, resets the counter, and loops
|
||||
forever anyway. 120s covers the slow case. Both keys belong under `[Unit]`; systemd
|
||||
takes them under `[Service]` with only a warning and then ignores them.
|
||||
|
||||
**Why `OnFailure=` at all.** `StartLimitBurst` turns the loop into a `failed` state,
|
||||
which is much better, and is still a state somebody has to go and look at. `OnFailure=`
|
||||
is what makes the machine speak first. `%n` expands to the failed unit's own name, which
|
||||
arrives as the template instance in `%i`.
|
||||
|
||||
The second silence is a **build that publishes an index it should have refused**; that
|
||||
one is the `MIN_MINTS_FOR_BUILD` gate under "Rebuilds", and the `discovery_starved` flag
|
||||
under "Discovery starvation" is what feeds it.
|
||||
|
||||
#### The alert unit
|
||||
|
||||
```ini
|
||||
# /etc/systemd/system/cashumints-alert@.service
|
||||
#
|
||||
# The unit that makes a failure audible.
|
||||
#
|
||||
# The other three units each carry `OnFailure=cashumints-alert@%n.service`, so systemd
|
||||
# starts one of these with the failed unit's name as the instance — `%i` below is
|
||||
# literally `cashumints.service`, `cashumints-web.service` or `cashumints-site.service`.
|
||||
#
|
||||
# Why it exists: the API once crash looped 464 times over fifteen hours and nothing said
|
||||
# so. `Restart=on-failure` with no start limit is an infinite loop that never reaches a
|
||||
# `failed` state, so the journal filled with identical lines nobody was reading and
|
||||
# every signal stayed green. The other half of the fix is StartLimitBurst= in each unit,
|
||||
# which turns the loop into a failure; this is what carries that failure off the machine.
|
||||
#
|
||||
# Install:
|
||||
# sudo install -m 0644 deploy/cashumints-alert@.service /etc/systemd/system/
|
||||
# sudo install -d -m 0755 /etc/cashumints
|
||||
# sudo install -m 0640 -o root -g root deploy/alert.env.example /etc/cashumints/alert.env
|
||||
# sudo systemctl daemon-reload
|
||||
#
|
||||
# No [Install] section and never enabled: OnFailure= starts it, and a unit that also
|
||||
# started at boot would page on every reboot.
|
||||
|
||||
[Unit]
|
||||
Description=Notify that %i failed
|
||||
# No OnFailure= here. An alerter that alerts about its own failure is a loop, and this
|
||||
# one is written so its worst case is a journal line rather than a retry.
|
||||
|
||||
[Service]
|
||||
Type=oneshot
|
||||
|
||||
# The one file an operator edits, and the only reason this unit is configurable at all.
|
||||
# Absent is a supported state — the leading `-` says so — and then the ExecStart below
|
||||
# still writes to the journal at ERROR, which is what `systemctl status` and
|
||||
# `journalctl -p err` read. See alert.env.example.
|
||||
EnvironmentFile=-/etc/cashumints/alert.env
|
||||
|
||||
# So `journalctl -t cashumints-alert` finds every alert, whichever unit triggered it.
|
||||
SyslogIdentifier=cashumints-alert
|
||||
|
||||
# Everything is inside one shell so the "nothing configured" branch is reachable without
|
||||
# a second unit. The pieces, in order:
|
||||
#
|
||||
# - `printf '<3>…'` on stdout. systemd reads that syslog prefix off a journal stream
|
||||
# and files the line at priority 3, ERROR, so `journalctl -p err` is a complete
|
||||
# history of failures on a host with no webhook configured at all. `<4>` is warning.
|
||||
# A prefix rather than systemd-cat, so the unit needs nothing from the filesystem it
|
||||
# has just sandboxed itself away from.
|
||||
# - NTFY_URL is a topic URL (https://ntfy.sh/your-topic). It gets a plain-text body
|
||||
# naming the failed unit, plus the header names ntfy understands.
|
||||
# - WEBHOOK_URL gets a JSON POST instead, for Discord, Slack or anything that speaks
|
||||
# `{"content": …}` — every key is sent, so one payload fits all of them.
|
||||
# - `--max-time 10` and a `||` fallback on each: an alert that hangs would hold the
|
||||
# failed unit's job open, and an alert that fails must not itself become a second
|
||||
# failed unit for somebody to notice. The shell ends in `true` for the same reason.
|
||||
#
|
||||
# `%i` is the failed unit's name, passed as an argument rather than interpolated into
|
||||
# the shell text: systemd expands specifiers before /bin/sh ever sees the line, and a
|
||||
# unit name is not a thing to trust to quoting.
|
||||
ExecStart=/bin/sh -c '\
|
||||
UNIT="$1"; \
|
||||
HOST="$(hostname)"; \
|
||||
WHEN="$(date -Is)"; \
|
||||
TEXT="$UNIT failed on $HOST at $WHEN"; \
|
||||
printf "<3>%s\\n" "$TEXT"; \
|
||||
if [ -n "$NTFY_URL" ]; then \
|
||||
/usr/bin/curl -fsS --max-time 10 \
|
||||
-H "Title: cashumints: $UNIT failed" \
|
||||
-H "Priority: high" \
|
||||
-H "Tags: rotating_light" \
|
||||
-d "$TEXT" "$NTFY_URL" >/dev/null \
|
||||
|| printf "<3>%s\\n" "alert: POST to NTFY_URL failed"; \
|
||||
fi; \
|
||||
if [ -n "$WEBHOOK_URL" ]; then \
|
||||
/usr/bin/curl -fsS --max-time 10 \
|
||||
-H "Content-Type: application/json" \
|
||||
-d "{\\"unit\\":\\"$UNIT\\",\\"host\\":\\"$HOST\\",\\"at\\":\\"$WHEN\\",\\"text\\":\\"$TEXT\\",\\"content\\":\\"$TEXT\\"}" \
|
||||
"$WEBHOOK_URL" >/dev/null \
|
||||
|| printf "<3>%s\\n" "alert: POST to WEBHOOK_URL failed"; \
|
||||
fi; \
|
||||
if [ -z "$NTFY_URL" ] && [ -z "$WEBHOOK_URL" ]; then \
|
||||
printf "<4>%s\\n" "alert: no NTFY_URL or WEBHOOK_URL in /etc/cashumints/alert.env, journal only"; \
|
||||
fi; \
|
||||
true' _ %i
|
||||
|
||||
# It sends one HTTP request and writes one line. It needs no identity of its own, and
|
||||
# DynamicUser gives it a throwaway one rather than sharing `nobody` with everything else
|
||||
# on the host that also could not be bothered to make a user.
|
||||
DynamicUser=yes
|
||||
NoNewPrivileges=true
|
||||
PrivateDevices=true
|
||||
ProtectSystem=strict
|
||||
ProtectHome=true
|
||||
ProtectKernelTunables=true
|
||||
ProtectKernelModules=true
|
||||
ProtectControlGroups=true
|
||||
RestrictAddressFamilies=AF_UNIX AF_INET AF_INET6
|
||||
RestrictSUIDSGID=true
|
||||
LockPersonality=true
|
||||
# An alert that cannot reach the network in ten seconds is not worth a stuck job.
|
||||
TimeoutStartSec=30
|
||||
```
|
||||
|
||||
Configuration is one optional file. With it absent, or with both values empty, a failure
|
||||
still lands in the journal at `ERROR` and is readable with `journalctl -p err -t
|
||||
cashumints-alert`; the unit is written so that its worst case is a log line rather than
|
||||
a second thing to debug.
|
||||
|
||||
```ini
|
||||
# /etc/cashumints/alert.env
|
||||
#
|
||||
# Read by cashumints-alert@.service, which systemd starts when any of the three units
|
||||
# fails. Everything here is optional: with the file absent or both values empty, an
|
||||
# alert is still written to the journal at ERROR priority and is readable with
|
||||
#
|
||||
# journalctl -p err -t cashumints-alert
|
||||
#
|
||||
# Set one or both to have failures leave the machine.
|
||||
#
|
||||
# Install it root-owned and not world-readable — a webhook URL is a capability:
|
||||
# sudo install -d -m 0755 /etc/cashumints
|
||||
# sudo install -m 0640 -o root -g root deploy/alert.env.example /etc/cashumints/alert.env
|
||||
# sudo systemctl daemon-reload
|
||||
|
||||
# An ntfy topic URL. Free and public at ntfy.sh; pick a topic name nobody will guess,
|
||||
# because anyone who knows it can read and post to it.
|
||||
#NTFY_URL=https://ntfy.sh/cashumints-alerts-CHANGE-ME
|
||||
|
||||
# Anything that accepts a JSON POST. The body carries `unit`, `host`, `at`, `text` and
|
||||
# `content` — the last of which is what Discord and most Slack-compatible endpoints read,
|
||||
# so one payload fits all three.
|
||||
#WEBHOOK_URL=https://discord.com/api/webhooks/…
|
||||
```
|
||||
|
||||
Test it without breaking anything:
|
||||
|
||||
```bash
|
||||
sudo systemctl start 'cashumints-alert@test.service'
|
||||
journalctl -t cashumints-alert -n 5 --no-pager
|
||||
```
|
||||
|
||||
### Rebuilds
|
||||
|
||||
Mint pages are prerendered, so new mints and new review counts appear at the next build.
|
||||
@@ -1196,6 +1378,7 @@ Description=Rebuild the cashumints.space static site
|
||||
Requires=cashumints.service
|
||||
After=cashumints.service network-online.target
|
||||
Wants=network-online.target
|
||||
OnFailure=cashumints-alert@%n.service
|
||||
|
||||
[Service]
|
||||
Type=oneshot
|
||||
|
||||
@@ -0,0 +1,23 @@
|
||||
# /etc/cashumints/alert.env
|
||||
#
|
||||
# Read by cashumints-alert@.service, which systemd starts when any of the three units
|
||||
# fails. Everything here is optional: with the file absent or both values empty, an
|
||||
# alert is still written to the journal at ERROR priority and is readable with
|
||||
#
|
||||
# journalctl -p err -t cashumints-alert
|
||||
#
|
||||
# Set one or both to have failures leave the machine.
|
||||
#
|
||||
# Install it root-owned and not world-readable — a webhook URL is a capability:
|
||||
# sudo install -d -m 0755 /etc/cashumints
|
||||
# sudo install -m 0640 -o root -g root deploy/alert.env.example /etc/cashumints/alert.env
|
||||
# sudo systemctl daemon-reload
|
||||
|
||||
# An ntfy topic URL. Free and public at ntfy.sh; pick a topic name nobody will guess,
|
||||
# because anyone who knows it can read and post to it.
|
||||
#NTFY_URL=https://ntfy.sh/cashumints-alerts-CHANGE-ME
|
||||
|
||||
# Anything that accepts a JSON POST. The body carries `unit`, `host`, `at`, `text` and
|
||||
# `content` — the last of which is what Discord and most Slack-compatible endpoints read,
|
||||
# so one payload fits all three.
|
||||
#WEBHOOK_URL=https://discord.com/api/webhooks/…
|
||||
@@ -0,0 +1,101 @@
|
||||
# /etc/systemd/system/cashumints-alert@.service
|
||||
#
|
||||
# The unit that makes a failure audible.
|
||||
#
|
||||
# The other three units each carry `OnFailure=cashumints-alert@%n.service`, so systemd
|
||||
# starts one of these with the failed unit's name as the instance — `%i` below is
|
||||
# literally `cashumints.service`, `cashumints-web.service` or `cashumints-site.service`.
|
||||
#
|
||||
# Why it exists: the API once crash looped 464 times over fifteen hours and nothing said
|
||||
# so. `Restart=on-failure` with no start limit is an infinite loop that never reaches a
|
||||
# `failed` state, so the journal filled with identical lines nobody was reading and
|
||||
# every signal stayed green. The other half of the fix is StartLimitBurst= in each unit,
|
||||
# which turns the loop into a failure; this is what carries that failure off the machine.
|
||||
#
|
||||
# Install:
|
||||
# sudo install -m 0644 deploy/cashumints-alert@.service /etc/systemd/system/
|
||||
# sudo install -d -m 0755 /etc/cashumints
|
||||
# sudo install -m 0640 -o root -g root deploy/alert.env.example /etc/cashumints/alert.env
|
||||
# sudo systemctl daemon-reload
|
||||
#
|
||||
# No [Install] section and never enabled: OnFailure= starts it, and a unit that also
|
||||
# started at boot would page on every reboot.
|
||||
|
||||
[Unit]
|
||||
Description=Notify that %i failed
|
||||
# No OnFailure= here. An alerter that alerts about its own failure is a loop, and this
|
||||
# one is written so its worst case is a journal line rather than a retry.
|
||||
|
||||
[Service]
|
||||
Type=oneshot
|
||||
|
||||
# The one file an operator edits, and the only reason this unit is configurable at all.
|
||||
# Absent is a supported state — the leading `-` says so — and then the ExecStart below
|
||||
# still writes to the journal at ERROR, which is what `systemctl status` and
|
||||
# `journalctl -p err` read. See alert.env.example.
|
||||
EnvironmentFile=-/etc/cashumints/alert.env
|
||||
|
||||
# So `journalctl -t cashumints-alert` finds every alert, whichever unit triggered it.
|
||||
SyslogIdentifier=cashumints-alert
|
||||
|
||||
# Everything is inside one shell so the "nothing configured" branch is reachable without
|
||||
# a second unit. The pieces, in order:
|
||||
#
|
||||
# - `printf '<3>…'` on stdout. systemd reads that syslog prefix off a journal stream
|
||||
# and files the line at priority 3, ERROR, so `journalctl -p err` is a complete
|
||||
# history of failures on a host with no webhook configured at all. `<4>` is warning.
|
||||
# A prefix rather than systemd-cat, so the unit needs nothing from the filesystem it
|
||||
# has just sandboxed itself away from.
|
||||
# - NTFY_URL is a topic URL (https://ntfy.sh/your-topic). It gets a plain-text body
|
||||
# naming the failed unit, plus the header names ntfy understands.
|
||||
# - WEBHOOK_URL gets a JSON POST instead, for Discord, Slack or anything that speaks
|
||||
# `{"content": …}` — every key is sent, so one payload fits all of them.
|
||||
# - `--max-time 10` and a `||` fallback on each: an alert that hangs would hold the
|
||||
# failed unit's job open, and an alert that fails must not itself become a second
|
||||
# failed unit for somebody to notice. The shell ends in `true` for the same reason.
|
||||
#
|
||||
# `%i` is the failed unit's name, passed as an argument rather than interpolated into
|
||||
# the shell text: systemd expands specifiers before /bin/sh ever sees the line, and a
|
||||
# unit name is not a thing to trust to quoting.
|
||||
ExecStart=/bin/sh -c '\
|
||||
UNIT="$1"; \
|
||||
HOST="$(hostname)"; \
|
||||
WHEN="$(date -Is)"; \
|
||||
TEXT="$UNIT failed on $HOST at $WHEN"; \
|
||||
printf "<3>%s\\n" "$TEXT"; \
|
||||
if [ -n "$NTFY_URL" ]; then \
|
||||
/usr/bin/curl -fsS --max-time 10 \
|
||||
-H "Title: cashumints: $UNIT failed" \
|
||||
-H "Priority: high" \
|
||||
-H "Tags: rotating_light" \
|
||||
-d "$TEXT" "$NTFY_URL" >/dev/null \
|
||||
|| printf "<3>%s\\n" "alert: POST to NTFY_URL failed"; \
|
||||
fi; \
|
||||
if [ -n "$WEBHOOK_URL" ]; then \
|
||||
/usr/bin/curl -fsS --max-time 10 \
|
||||
-H "Content-Type: application/json" \
|
||||
-d "{\\"unit\\":\\"$UNIT\\",\\"host\\":\\"$HOST\\",\\"at\\":\\"$WHEN\\",\\"text\\":\\"$TEXT\\",\\"content\\":\\"$TEXT\\"}" \
|
||||
"$WEBHOOK_URL" >/dev/null \
|
||||
|| printf "<3>%s\\n" "alert: POST to WEBHOOK_URL failed"; \
|
||||
fi; \
|
||||
if [ -z "$NTFY_URL" ] && [ -z "$WEBHOOK_URL" ]; then \
|
||||
printf "<4>%s\\n" "alert: no NTFY_URL or WEBHOOK_URL in /etc/cashumints/alert.env, journal only"; \
|
||||
fi; \
|
||||
true' _ %i
|
||||
|
||||
# It sends one HTTP request and writes one line. It needs no identity of its own, and
|
||||
# DynamicUser gives it a throwaway one rather than sharing `nobody` with everything else
|
||||
# on the host that also could not be bothered to make a user.
|
||||
DynamicUser=yes
|
||||
NoNewPrivileges=true
|
||||
PrivateDevices=true
|
||||
ProtectSystem=strict
|
||||
ProtectHome=true
|
||||
ProtectKernelTunables=true
|
||||
ProtectKernelModules=true
|
||||
ProtectControlGroups=true
|
||||
RestrictAddressFamilies=AF_UNIX AF_INET AF_INET6
|
||||
RestrictSUIDSGID=true
|
||||
LockPersonality=true
|
||||
# An alert that cannot reach the network in ten seconds is not worth a stuck job.
|
||||
TimeoutStartSec=30
|
||||
@@ -12,12 +12,18 @@
|
||||
Description=cashumints.space static site server
|
||||
Wants=network-online.target
|
||||
After=network-online.target
|
||||
# Give up after five failures in a minute instead of restarting forever. A process that
|
||||
# cannot start will not start on the 4000th attempt either, and `failed` in
|
||||
# `systemctl status` is a far louder signal than a journal scrolling past. These two are
|
||||
# [Unit] keys; systemd ignores them under [Service] with only a warning.
|
||||
StartLimitIntervalSec=60
|
||||
# Give up after five failures in two minutes instead of restarting forever. A process
|
||||
# that cannot start will not start on the 4000th attempt either, and `failed` in
|
||||
# `systemctl status` is a far louder signal than a journal scrolling past. The window
|
||||
# matches cashumints.service; see the note there for why it is 120s and not 60s. These
|
||||
# two are [Unit] keys; systemd ignores them under [Service] with only a warning.
|
||||
StartLimitIntervalSec=120
|
||||
StartLimitBurst=5
|
||||
|
||||
# Carry a failure off the machine. `%n` is this unit's own name, so the alert says
|
||||
# which one died. cashumints-alert@.service writes to the journal at ERROR always and
|
||||
# curls NTFY_URL or WEBHOOK_URL from /etc/cashumints/alert.env when either is set.
|
||||
OnFailure=cashumints-alert@%n.service
|
||||
# Not Requires=cashumints.service: the pages are prerendered, so the site keeps serving
|
||||
# a correct-as-of-last-build copy while the API is down. Only the islands go quiet.
|
||||
|
||||
|
||||
@@ -22,6 +22,11 @@ Requires=cashumints.service
|
||||
After=cashumints.service network-online.target
|
||||
Wants=network-online.target
|
||||
|
||||
# Carry a failure off the machine. `%n` is this unit's own name, so the alert says
|
||||
# which one died. cashumints-alert@.service writes to the journal at ERROR always and
|
||||
# curls NTFY_URL or WEBHOOK_URL from /etc/cashumints/alert.env when either is set.
|
||||
OnFailure=cashumints-alert@%n.service
|
||||
|
||||
[Service]
|
||||
Type=oneshot
|
||||
User=cashumints
|
||||
|
||||
@@ -3,13 +3,27 @@
|
||||
Description=cashumints.space indexer and API
|
||||
Wants=network-online.target
|
||||
After=network-online.target
|
||||
# Stop after five failures in a minute rather than restarting forever. A wrong Node on
|
||||
# PATH once produced four thousand identical crashes in the journal before anyone read
|
||||
# one of them; `failed` in `systemctl status` says the same thing in one line. Both keys
|
||||
# belong to [Unit] — under [Service] systemd only warns and ignores them.
|
||||
StartLimitIntervalSec=60
|
||||
# Stop after five failures in two minutes rather than restarting forever.
|
||||
#
|
||||
# The window is 120s and not 60s because RestartSec=5s below means five attempts take
|
||||
# a little over twenty seconds of restarts plus however long each attempt lives before
|
||||
# it dies. A process that fails *slowly* — a database that times out, a port that takes
|
||||
# four seconds to refuse — can spread five failures past a sixty second window and reset
|
||||
# the counter forever, which is the loop this is supposed to stop. 120s covers that.
|
||||
#
|
||||
# The failure this exists for: ExecStart named a .ts file, /usr/bin/node was 20, and
|
||||
# every start died in under a second. 464 restarts over fifteen hours, and because
|
||||
# Restart=on-failure without a start limit never reaches a `failed` state, nothing
|
||||
# anywhere went red. Both keys belong to [Unit] — under [Service] systemd only warns and
|
||||
# ignores them.
|
||||
StartLimitIntervalSec=120
|
||||
StartLimitBurst=5
|
||||
|
||||
# Carry a failure off the machine. `%n` is this unit's own name, so the alert says
|
||||
# which one died. cashumints-alert@.service writes to the journal at ERROR always and
|
||||
# curls NTFY_URL or WEBHOOK_URL from /etc/cashumints/alert.env when either is set.
|
||||
OnFailure=cashumints-alert@%n.service
|
||||
|
||||
[Service]
|
||||
Type=simple
|
||||
User=cashumints
|
||||
|
||||
Reference in New Issue
Block a user