Health answering 200 and the index being complete are different claims. A year of
~31-event backfills left a perfectly healthy API serving a real, correct, complete
list of eight mints. A build against that succeeds — it prerenders eight cards —
and rsync --delete-after then replaces fifty-five with eight.
cashumints-web.service gains a second ExecStartPre after the health wait: count
/api/mints, and exit non-zero below MIN_MINTS_FOR_BUILD (default 20, overridable
with `systemctl edit`). A refusal aborts the unit before `pnpm build`, and
publishing is ExecStartPost, so the previously published site is untouched; the
OnFailure alert added in the last commit says why.
Counted by the "host": key rather than by counting braces, because the list
payload is about to carry a nested object per mint. A curl that fails at all
counts as zero, which is below every floor — so an API that fell over between the
health check and this line refuses the build instead of sailing through it.
Verified against three live APIs: 73 mints passes, a doctored 8-mint database
fails with the reason, and a dead port fails.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`Restart=on-failure` with no start limit is an infinite loop by definition: the
unit never reaches `failed`, `systemctl status` stays active (auto-restart), and
the only evidence is a journal moving at four lines a second. That is how 464
restarts over fifteen hours went unnoticed.
All three units now stop after five failures in 120s and run
OnFailure=cashumints-alert@%n.service. The window is 120s and not 60s because
RestartSec=5s plus a process that takes a few seconds to die can spread five
failures past a sixty second window, reset the counter, and loop forever anyway.
cashumints-alert@.service is a oneshot that takes the failed unit's name as its
instance. Configuration is /etc/cashumints/alert.env: NTFY_URL gets a plain-text
body, WEBHOOK_URL gets JSON carrying `content` so one payload fits Discord and
Slack-compatible endpoints. With neither set — or the file absent — it still
writes to the journal at ERROR via a `<3>` syslog prefix, so `journalctl -p err -t
cashumints-alert` is a complete history on a host nobody configured.
It cannot become a second thing to debug: each curl is bounded at 10s, each
failure falls back to a journal line, and the shell ends in `true`, so the alerter
always exits 0. Verified with systemd-analyze verify and by running the ExecStart
body against a local sink — the JSON parses, and every branch exits 0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
For about a year the production RELAYS list did not include the relay carrying
the kind 38000/38172 archive. Every backfill read about thirty events, wrote them
faithfully, reported ok=true, and the nightly build republished an index of eight
mints. Nothing measured the difference between "the cycle completed" and "the
cycle read anything", so nothing went red.
Three signals now do:
- Per-relay attribution. queryRelays() replaces pool.querySync(), which merges
every relay into one deduplicated array and throws away who sent what. It
keeps one subscription per relay over the pool's existing sockets and shares
a single alreadyHaveEvent across them, so an event five relays carry is still
verified once; receivedEvent fires before that check, which is what makes the
per-relay count mean "what this relay contributed". The deadline moved out of
each Subscription's own EOSE timer so `eose` means a frame arrived rather than
something timed out.
- A WARN naming any relay that will not connect, on every cycle, and any relay
that connected and sent nothing, on backfills only. An incremental cycle is
supposed to come back empty.
- BACKFILL_MIN_EVENTS, default 200. Under it, ERROR discovery starvation
suspected and a flag health reports as discovery_starved, forcing 503. Sticky
across incremental cycles so an hourly cycle finding four events cannot clear
what a backfill diagnosed; stored in the database so a restart cannot either.
A fresh database is starved until its first backfill lands. That is intended: it
holds the build's health gate rather than publishing a site made from nothing.
Verified against the live relay set — 1528 events, five relays connected, EOSE on
all five, health 200 — and against an unreachable list, which produces the two
WARN lines, the ERROR, and 503.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The unit's ExecStart named src/index.ts, so every start depended on the host
having Node 22.18 or newer for native type stripping. A deploy onto a host with
Node 20 met ERR_UNKNOWN_FILE_EXTENSION, exited in under a second, and was
restarted 464 times over fifteen hours with nothing anywhere going red.
api/tsconfig.json now emits to api/dist. The source keeps its explicit .ts import
specifiers, which is what makes `node --watch src/index.ts` work in development;
rewriteRelativeImportExtensions turns them into .js on the way out, so what runs
in production is ordinary ESM that any Node from 20.18 up will start.
`pnpm build` builds shared, then api, then web. `pnpm dev` is unchanged.
deploy/ is tracked rather than ignored: the unit files are the thing an operator
copies to /etc/systemd/system, and the alert unit added next has to live
somewhere a deploy can find it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Gate login, reveal the form after a rating, and wire inline Rate this mint; drop --experimental-strip-types.
Co-authored-by: Cursor <cursoragent@cursor.com>
Add lnurl list/detail routes, OG fixtures, i18n strings, and the write/
index client flows so the site surfaces the new mint type end to end.
Co-authored-by: Cursor <cursoragent@cursor.com>
Add SearchAction and optional Product JSON-LD, per-page OG images/alts, and document the og build and cache headers.
Co-authored-by: Cursor <cursoragent@cursor.com>