michilisandClaude Opus 5 0ebc8ada54 Make a crash loop reach somebody instead of scrolling past.
`Restart=on-failure` with no start limit is an infinite loop by definition: the
unit never reaches `failed`, `systemctl status` stays active (auto-restart), and
the only evidence is a journal moving at four lines a second. That is how 464
restarts over fifteen hours went unnoticed.

All three units now stop after five failures in 120s and run
OnFailure=cashumints-alert@%n.service. The window is 120s and not 60s because
RestartSec=5s plus a process that takes a few seconds to die can spread five
failures past a sixty second window, reset the counter, and loop forever anyway.

cashumints-alert@.service is a oneshot that takes the failed unit's name as its
instance. Configuration is /etc/cashumints/alert.env: NTFY_URL gets a plain-text
body, WEBHOOK_URL gets JSON carrying `content` so one payload fits Discord and
Slack-compatible endpoints. With neither set — or the file absent — it still
writes to the journal at ERROR via a `<3>` syslog prefix, so `journalctl -p err -t
cashumints-alert` is a complete history on a host nobody configured.

It cannot become a second thing to debug: each curl is bounded at 10s, each
failure falls back to a journal line, and the shell ends in `true`, so the alerter
always exits 0. Verified with systemd-analyze verify and by running the ExecStart
body against a local sink — the JSON parses, and every branch exits 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 16:10:01 +02:00
2026-08-20 22:41:25 +02:00
2026-08-20 22:41:25 +02:00

cashumints.space

An ecash explorer and review site. Lists every Cashu mint and every Fedimint federation discoverable on the Nostr network, shows each one's metadata, and surfaces community reviews published as NIP-87 events.

Made by Azzamo.

  Nostr relays            Cashu mints      Fedimint federations
 (NIP-87 events)           (/v1/info)      (no status endpoint;
  38172 / 38173 / 38000                     see "Ecosystems")
       |                       |                    |
       v  discovery (hourly)   v  probe (10 min)    v  check (10 min)
 +--------------------------------------------------------------+
 |  api/   Node + Hono + SQLite                                  |
 |  last-known-good metadata, 4 REST endpoints                   |
 +--------------------------------------------------------------+
       |                          |
  build-time fetch          runtime fetch (islands)
       v                          v
 +--------------------------------------------------------------+
 |  web/   Astro, static, prerendered                            |
 +--------------------------------------------------------------+
                                  |
        browser <-> relays (review text, signing via NIP-07 or NIP-46)

The backend is a memory layer, not an authority. Its job is to remember what a mint looked like when it was last reachable, so an offline mint still has a page people can read and review. Review content is never stored server-side.

Identity is the same story: logging in is a browser-side affair the API never sees. A session is a pubkey plus, for a remote signer, the NIP-46 connection details, kept in localStorage. No private key is ever asked for, accepted or stored, by any code path in this repository.

Layout

Path What it is
api/ Indexer and REST API. Hono, SQLite or Postgres, nostr-tools.
web/ Astro static site with vanilla TypeScript islands.
shared/ Types, NIP-87 constants, URL normalization, scoring, NUT and module names.

The site is in English, Spanish and Dutch. web/src/i18n/ holds the message catalogs and the locale table; web/src/i18n/GLOSSARY.md holds the term decisions a translator needs before touching either. See Languages.

NOTES.md records what was found in the previous codebase: event kinds, tags, relays, the rating encoding, and the bugs this rebuild fixes.

Requirements

  • Node 20.18 or newer to build and to run what a build produces
  • Node 22.18 or newer to develop: pnpm dev, pnpm seed and the api test scripts run src/*.ts through node directly, which needs native type stripping
  • pnpm 9 or newer

engines.node is the first of those, not the second, on purpose: it is the floor a deployment has to clear, and a production host should never be told it needs a newer Node than the compiled service actually runs on.

Setup

pnpm install
pnpm --filter ./shared build

shared/ compiles to dist/, which both other packages import. Run it once after install and again whenever you change something in shared/src.

Seed the database

Ingests the initial mint list, probes every mint, runs a full discovery backfill against the relays, and prints a summary table. Takes about a minute against the live network.

pnpm seed

Database

Everything the API serves is read from its own database, never from a mint or a relay at request time: mint metadata, /v1/info payloads, reviews, probe history and the discovery cursor. That is what lets a mint page render in full while the mint is offline, and it is why a restart loses nothing — the process keeps no state of its own beyond a 60 second cache in front of /api/stats. Icons are files under ICON_DIR and survive alongside it.

Two backends, chosen by DATABASE_URL:

DATABASE_URL Backend
unset SQLite at DB_PATH. The default.
postgres://user:pw@host:5432/cashumints Postgres
sqlite:/var/lib/cashumints/cashumints.db SQLite at that path

SQLite is the right answer for a single API process, which is the shape this service has: one indexer, one writer, reads served from the page cache. Reach for Postgres when you need something SQLite cannot give you — the database on a different host from the API, more than one API process, or your existing backup and replication setup.

Neither needs a setup step. Both create their tables on first connection, so pointing the API at an empty Postgres database is the whole installation:

createdb cashumints

Both backends run the same SQL. One statement is written once and sent to either, so there is no dialect-specific query path to drift. pnpm --filter ./api test runs the review-dedupe statement against a real database, and CHECK_DB_URL runs it against Postgres too. See the header of api/src/db-schema.ts for the rules that keep a statement portable.

Moving between them

migrate copies every table from one database to the other. --from defaults to whatever the current configuration points at, so the usual direction needs only --to:

pnpm --filter ./api migrate --to postgres://user:pw@localhost:5432/cashumints

Then set DATABASE_URL to the same value and restart the API. It works in both directions and between two databases of the same kind:

pnpm --filter ./api migrate --from postgres://localhost/cashumints --to ./data/cashumints.db

Rows are upserted on their primary key, so an interrupted run can just be repeated, and the migrator reads the counts back from the target and fails if any table came up short. probes is the exception — an append-only log with no unique key, so a target that already has probe rows is refused unless you pass --force, which replaces them. Add --dry-run to see the row counts without writing anything.

Migrating while the API is running will copy a moving target. Stop it first.

Icons are files, not rows: migrate does not touch ICON_DIR, so copy that directory yourself if the new database lives on a different host.

Development

Runs the API on :8787 and the Astro dev server on :4321. Both ports come from .env — see Configuration.

pnpm dev

Run them separately if you prefer:

pnpm dev:api
pnpm dev:web

The web app reads the API at build time, so the API must be running before you build or start the dev server.

Skeleton loading

Two regions fetch their content at runtime and get a skeleton while they wait: the reviews panel body on a mint page (reviews come from Nostr relays) and the prerender miss summary on /mint/{host} for a mint added since the last build. Nothing prerendered is ever covered by a skeleton, so the mint header, verdict strip, sidebar, the reviews panel's own header and filter counts, the pulse ticker and every grid stay on screen as they are.

The bones are captured from the real rendered components by Boneyard and committed as web/src/bones/*.bones.json, so they ship with the island and draw immediately without a network round trip.

Regenerate them with the dev server running:

pnpm bones

That visits a real mint page and a deliberately unprerendered /mint/ URL at 360, 768 and 1280px wide, and rewrites both bone files. Both islands render fixture content (web/src/lib/skeleton-fixtures.ts) while the capture browser is looking, so the result does not depend on a relay answering. The fixtures load only during a capture and are not part of the bundle a visitor downloads.

Re-run pnpm bones after changing the layout of a review card or of the prerender miss summary, including CSS-only changes. Stale bones are not a build error, they just stop matching the content that replaces them.

Capturing needs a Chromium for Playwright, once per machine:

pnpm --filter ./web exec playwright install --with-deps chromium

Static assets

web/public/ ships as-is: the fonts, the social card and the icon set.

Fonts are self-hosted from web/src/assets/fonts/ — Space Grotesk, Inter and JetBrains Mono, as the latin and latin-ext woff2 subsets Google Fonts serves, with the same unicode-range declarations, so rendering is identical to the CDN it replaced. All three are variable fonts, so six files cover all fourteen faces. The site makes no third-party request for them, which is both one less thing in the critical path and one less entry in the list on /privacy. They are SIL Open Font License 1.1; the notice ships at /fonts/OFL.txt.

They sit in src/assets/ rather than public/ so the build fingerprints them into /_astro/, which is what makes Cache-Control: immutable honest and puts them under a cache rule web/server.mjs already has. Unhashed and uncached they are re-fetched on every client-side navigation — the router re-inserts their preload links on each swap — and the typefaces visibly reload from page to page.

The social card and icons (og.png, favicon.ico, favicon.svg, apple-touch-icon.png, icon-192.png, icon-512.png) are generated once and committed:

node web/scripts/make-assets.mjs

That draws the 1200x630 card in a headless Chromium using the site's own fonts and tokens, then downsamples one icon master into the rest. Re-run it after a brand change. It is deliberately not part of pnpm build: the site has to build on a machine with no browser, and these files change roughly never. It needs the same Chromium pnpm bones does.

Tests

Unit checks over URL normalization, rating parsing, scoring and the review-dedupe SQL:

pnpm --filter ./api test

The warning banners. Fixture-driven checks over getMintWarnings, the one function that decides whether a mint page shows "melt only", "withdrawals disabled", "mint frozen" or one of the three offline tiers. Fixtures are real /v1/info payloads (or a real one with a single flag flipped, noted in the file) under shared/fixtures/warnings:

pnpm --filter ./api test:warnings

The offline acceptance check. Copies the seeded database, points a healthy mint at a dead URL, probes it until it is marked offline, and asserts the mint page still serves its full cached metadata and reviews:

pnpm --filter ./api test:offline

The copy is made with the migrator, so the check runs against whichever backend holds the real data and exercises the migration path every time. Point CHECK_DB_URL at a scratch Postgres database to run the whole thing there — it is emptied first, so give it one of its own:

CHECK_DB_URL=postgres://localhost/cashumints_test pnpm --filter ./api test:offline

CHECK_DB_URL does the same for pnpm --filter ./api test, which then runs the review-dedupe SQL against both backends instead of just SQLite.

On-demand indexing. The address rules that keep POST /api/index from being an SSRF hole, the redirect hops, the slug collapse that stops one mint becoming two rows, the invite-code decoder and the rate limiter. Nothing here touches the network: the resolver and the fetch are both injected, because a rule that can only be exercised against the real internet stops being exercised the first time CI runs offline.

pnpm --filter ./api test:index

Type checking across the workspace:

pnpm typecheck

Build the site

pnpm build

Three packages in order, and the order is a dependency chain rather than a habit: shared emits the types and the warning copy both other packages import, api compiles api/src to api/dist, and web prerenders against a running API.

Output lands in api/dist/ and web/dist/. Every page is prerendered once per language, so ~55 mints and 9 static routes come out as ~200 pages, each with real titles, meta descriptions, OpenGraph and Twitter tags, a social card, a self-referencing canonical, a full hreflang set and a JSON-LD graph. sitemap.xml lists every indexable one with its xhtml:link alternates, and robots.txt sits beside it.

Social cards

pnpm build starts with pnpm og (web/scripts/og/build-og.mjs), which renders one 1200x630 PNG per mint and federation into web/public/og/ from the same API data the pages use — satori lays the card out and turns the text into glyph paths (the six static TTFs under web/scripts/og/fonts/ are the same faces the site uses), @resvg/resvg-js rasterises it. No browser involved; a full run of ~70 cards takes a few seconds, and a rerun with unchanged data renders nothing: each card's inputs are hashed into its filename (mint.example.com.a1b2c3d4e5.png) and web/src/generated/og-manifest.json records what is already on disk. The hashed name is also the cache-busting story — link-preview scrapers cache an og:image by URL, so a mint whose rating moved or that went offline gets a new URL on the next scheduled rebuild, while the stable-named default.png (brand plus network stats, used by every non-mint page) is served with a one-hour cache instead — the one exception in cacheControl() in web/server.mjs.

The images are one English render shared by every locale; titles, descriptions and alt text translate per page. Relative times stay out of the PNGs on purpose — a "3d ago" would go stale inside a static file — so the cards carry absolute month/year stamps, and the one exception, the "Offline {n}d" chip, is derived from a day-granular clock so it regenerates at most once a day.

pnpm og:fixtures renders the six edge cases in web/scripts/og/fixtures.mjs (30-character name, no icon, zero reviews, offline, melt only, announced-only federation) into web/og-fixtures/ at a pinned timestamp; those snapshots are committed, so eyeball them after any template change.

What is indexed, and what is not

One rule, in web/src/lib/seo.ts, decides it, and it covers both ecosystems: a page is left out of the index when the site has never once reached the thing it is about and nobody has reviewed it. Such a page has no name, no description, no version and no reviews, because all of those come from a mint that answered, a federation a check confirmed, or a person who wrote something — so every one of them is the same page as the next. It stays listed on /mints or /fedimints, stays linked, stays searchable on the site and stays reviewable; it just carries noindex, follow and is absent from the sitemap, until something answers or someone reviews it.

That rule is why an announced-only federation is usually unindexed: nothing has confirmed it, so until it collects a review its page says no more than the announcement did.

The two detail pages and the sitemap import that one predicate rather than each testing for it, and pnpm check:hreflang verifies from the built output that they still agree — a page saying noindex while the sitemap advertises it is a contradiction, and it fails the build.

Structured data is assembled in web/src/layouts/Base.astro and nowhere else, from the builders in web/src/lib/schema.ts, so a document cannot end up with two ld+json blocks or two conflicting @ids. Every page carries Organization, WebSite and WebPage; mint pages add BreadcrumbList and a Service node whose aggregateRating appears only when reviews actually carried ratings; /mints and /wallets add an ItemList. Mint names and descriptions are written by mint operators, so serializeGraph escapes them for the one context that matters — nothing that could close or reopen a script element survives it.

Check the build for broken internal links, links to routes that were never emitted, and anchors pointing at ids no page has:

pnpm check:links

It exits non-zero on a problem, so it can gate a deploy. LIST_TARGETS=1 prints every internal target and which pages link to it. NOTES-PAGES.md holds the last full audit.

Check the i18n and indexing side of the build: <html lang> against the URL's locale, a self-referencing canonical on every page, a complete and reciprocal hreflang set including x-default, and a sitemap that lists every indexable page with its alternates and no page that asked not to be indexed:

pnpm check:hreflang

The catalogs are checked before every build, and separately with:

pnpm check:i18n

Both gate a deploy. See Languages for what they look for.

Languages

Twenty-four, listed in web/src/i18n/config.ts. English is the site as it was; the other twenty-three are the same site, prerendered again, 84 pages each.

English, Spanish and Dutch are hand-written. The other twenty-one began as bulk machine translation, and they are in two states.

German, Danish, Swedish, Indonesian, Vietnamese and Turkish have had a full terminology pass: every string that names a mint, and every label, title and meta description on the home page, the mint list and a mint page. The machine had translated the site's central noun into the local word for a coin factory (Münzprägeanstalt, mincovna, darphane) or, in the Germanic languages, into the sweet: Danish and Swedish shipped Cashu-pastiller, "Cashu lozenges", and Swedish offered to sort your minttabletter. Indonesian had permen, Vietnamese cây mint, the mint plant. Those six now say mint, and sats, ecash and Lightning survive untranslated in them as the glossary requires.

The other fifteen still carry that damage in their body copy, in the same shapes: Czech mincovny, Polish mennice, Romanian monetării, Greek νομισματοκοπείο. Their chrome, titles and meta descriptions are repaired, so what a reader meets first is right, but the prose underneath is not. Roughly 400 strings, concentrated in the four inflected languages where a word-level fix needs case-correct edits inside running sentences, which is a native speaker's job rather than a careful search and replace.

web/src/i18n/GLOSSARY.md is where a pass starts, and the house rules at the top of it hold in every language, particularly the first row of the table.

Routing

Path-prefix locales, with English at the root:

Language Home A mint page
English / /mint/kashu.me
Spanish /es /es/mint/kashu.me
Dutch /nl /nl/mint/kashu.me

Every existing English URL stays exactly where it was, which matters more here than tidiness: those URLs are indexed, linked, and pasted into wallets.

Route slugs stay English in every language: /es/mints, never /es/mentas. This is a deliberate trade. Translated slugs read marginally better in an address bar and cost a permanent, growing table of slug-per-route-per-language that every internal link, the sitemap, the link checker and the switcher have to agree on forever. Stable URLs are worth more here, and the hreflang set is what actually tells a crawler these are the same page.

Nothing redirects by language. No Accept-Language sniffing, no IP geolocation. A crawler asking for /mint/kashu.me gets /mint/kashu.me, and a reader who deliberately opened the English page is not overruled by a browser setting they configured years ago. The language control in the header is how you change language, and it lands on the page you were already reading.

One file per route emits every locale: pages live under web/src/pages/[...locale]/, and localePaths() in web/src/i18n/paths.ts is their getStaticPaths. There is one copy of each page's markup, translated by t(), rather than twenty-four to keep in step.

What is and is not translated

Translated: every piece of site copy. Navigation, buttons, labels, filters, empty states, error states, form copy, tooltips, aria-labels, the static pages, the warning banners, page titles and meta descriptions, and the plain-language name beside a NUT number.

Never translated, because it is data rather than copy:

  • review text, and the name, NIP-05 and npub of whoever wrote it
  • mint names, mint descriptions and MOTDs, in whatever language the operator wrote them
  • URLs, hosts, npubs, pubkeys, version strings
  • protocol terms: NUT-04, NIP-87, bunker://, Nostr, Lightning, ecash, sat

A mint's own description is the first sentence of its page's meta description, verbatim. Only when a mint has published none does the site write that sentence itself.

Adding a language

Five steps. Four of them the build tells you about; the fifth it cannot, which is why it has a test of its own.

  1. web/src/i18n/GLOSSARY.md — add the column and decide the terms first. This is the step people skip, and it is the one that costs later: a reader who meets two words for "review" has to work out whether they are the same thing.

  2. web/src/i18n/xx.json — copy en.json and translate it. Flat, dotted keys. A key with a count uses .one / .other; a language with more plural categories may add .few, .many, .zero, and Intl.PluralRules picks between them. Nothing else in the codebase has to know.

  3. web/src/i18n/config.ts — add the entry to LOCALES:

    { code: 'pt', label: 'Português', intl: 'pt-BR', og: 'pt_BR' },
    

    label is the language's own name for itself, and is what the switcher shows. intl is the tag Intl gets, region included, because number and date formatting differ by region even when the language does not.

  4. web/src/i18n/locales.mjs — add the code to LOCALE_CODES. This is the same list in plain JavaScript, for astro.config.mjs and the checker, both of which run before TypeScript exists. check-i18n compares the two lists and fails if they disagree, so forgetting this is a build error rather than a mystery.

  5. web/src/i18n/index.ts — import pt from './pt.json' and add pt: pt as Catalog to CATALOGS. This is the step to get right, because it is the only one nothing downstream complains about: catalogFor() falls back to English by design, so a locale listed in LOCALES with no catalog behind it does not fail the build. It prerenders the whole English site under the new lang, the new URL prefix and a full reciprocal hreflang set, and tells every crawler those pages are a different language. Twenty-one locales shipped in exactly that state once. pnpm test now fails on it: see web/test/i18n-wiring.test.mjs.

That is the whole change. The Astro i18n config, the routes, the switcher, the hreflang sets, the og:locale:alternate list and the sitemap all read from LOCALES and pick the new language up on the next build.

Two things you do not have to do: dates, times and numbers come from Intl and are right for the new locale immediately, and relative times fall back to Intl.RelativeTimeFormat for any unit the new catalog does not tighten with its own time.ago.* string.

What the build checks

pnpm check:i18n runs before every build (from astro.config.mjs) and on its own. Four questions:

missing a key English has and this locale does not warning, listed by name
unknown a key a locale has and English does not error
undefined a t('...') in the source no catalog defines error
client an island reaching for a key outside CLIENT_NAMESPACES error

Missing keys are a warning on purpose: they fall back to English, and half a language is better than no language. Everything else is an error, because each one puts a wrong string in front of a reader. A key that resolves nowhere at all renders as humanised words rather than as mint.warnings.meltOnly, so even the unreachable case is not a raw key on screen.

It also diffs the warning-banner copy against shared/src/warnings.ts. The decision about which banner a mint gets lives in shared/ and has to stay language-free, so the English sentences live there too (the API's fixture tests assert on them) and the catalogs carry the same keys for translation. Two copies of safety copy is exactly the sort of thing that drifts, so the two are compared on every build.

pnpm check:hreflang runs over dist/ after a build: <html lang> against the URL, self-referencing canonicals, complete and reciprocal alternate sets, x-default pointing at English, every alternate resolving to a page that was actually emitted, and a sitemap entry with alternates for each. The 404 pages are the one exception and are checked for the opposite: no canonical, no alternates, and a noindex.

How the strings reach an island

The prerendered half of the site is translated at build time and costs a visitor nothing. The islands need their strings in the browser, and the requirement is that a Spanish page ships Spanish and nothing else.

Base.astro inlines the current locale's client-facing keys as one <script type="application/json" data-i18n> in the body, already resolved against English. Islands read it through useI18n() in web/src/i18n/client.ts.

The alternative, importing es.json from an island, does not work here: island chunks are shared across every locale, so anything imported into one is downloaded by all of them. A dynamic import() keyed on locale would avoid that but costs a round trip on the critical path of every island and leaves every catalog sitting in dist/_astro/. Inlining travels in HTML the page was sending anyway.

It is in the body rather than the head because the view transition router replaces the body on every navigation. An island reading it after a swap reads the page it is actually on; a head script would be kept and would keep answering in the language the reader arrived in.

CLIENT_NAMESPACES in config.ts decides which namespaces are inlined, and check-i18n fails the build if an island reaches outside that list.

Configuration

Every variable has a working default, so nothing has to be set to run the site locally. To change any of them, copy the example file and edit it:

cp .env.example .env

One .env at the repo root serves all three packages. The API loads it through node's own --env-file-if-exists, and the web side through web/scripts/load-env.mjs, imported at the top of astro.config.mjs. In both, a real environment variable wins over the file, so a systemd Environment= line or a one-off PORT=9000 pnpm dev:api still overrides it. .env is gitignored; .env.example is the documented copy.

API (api/)

Variable Default Meaning
PORT 8787 HTTP port
DATABASE_URL empty (SQLite at DB_PATH) postgres://… or sqlite:…. See Database.
DB_PATH api/data/cashumints.db SQLite file. Ignored when DATABASE_URL is set.
DB_POOL_MAX 10 Postgres connections held open. Unused by SQLite.
ICON_DIR api/data/icons Cached mint icons, served at /icons/*
RELAYS see shared/src/nostr.ts Comma separated relay list
BACKFILL_MIN_EVENTS 200 Events a backfill has to read before it counts as one. Under it, discovery logs ERROR discovery starvation suspected and health goes 503. See Starvation.
PROBE_INTERVAL_MIN 10 Minutes between probe cycles
DISCOVERY_INTERVAL_MIN 60 Minutes between discovery cycles
PROBE_CONCURRENCY 8 Mints probed in parallel
PROBE_TIMEOUT_MS 5000 Per-mint request timeout
FEDIMINT_OBSERVER_URL https://observer.fedimint.org/api/federations Where federation health is read from. Empty disables the lookup, and every federation stays announced. See Ecosystems.
SCORE_PRIOR_MEAN 3 Bayesian prior. See "Ranking" below before changing.
INDEX_RATE_LIMIT 10 POST /api/index submissions allowed per address per hour

Web (web/)

Variable Default Meaning
WEB_PORT 4321 Dev server and preview port
API_URL http://127.0.0.1:8787 API origin used at build time
PUBLIC_API_URL empty (same origin) API origin baked into markup and islands
SITE_URL https://cashumints.space Canonical origin for meta tags
PLAUSIBLE_URL https://analytics.azzamo.net/js/script.js Analytics script URL; empty disables analytics
PLAUSIBLE_DOMAIN cashumints.space Domain reported to analytics; empty disables analytics
RELAYS see shared/src/nostr.ts Relays for the build-time review fetch
BONES_URL http://localhost:$WEB_PORT Dev server pnpm bones captures against

PORT and API_URL have to agree: the build and the dev proxy reach the API at API_URL, so moving the API off 8787 means changing both.

PUBLIC_API_URL deliberately does not fall back to API_URL: that is a build-machine address, and a visitor's browser cannot reach 127.0.0.1. Left empty, icon src attributes and island fetches are same-origin paths (/icons/..., /api/...), which the dev server proxies to API_URL and nginx forwards in production — see "nginx" under Deployment below.

Set it only when the API answers on its own origin:

API_URL=http://127.0.0.1:8787 PUBLIC_API_URL=https://api.cashumints.space pnpm build

Ecosystems

Two things are listed: Cashu mints at /mints and /mint/{host}, and Fedimint federations at /fedimints and /fedimint/{slug}. They share one table, one review pipeline, one review card and one write-review dialog; they differ in how they are discovered, how they are checked, and what their page can honestly say.

Cashu Fedimint
mints.type cashu fedimint
Announcement kind:38172 kind:38173
Review k tag 38172 38173
Identity (d) the pubkey from /v1/info the federation id
Address (u) the mint URL the invite code (fed11…)
Row key (mints.url) the mint URL fedimint:<federation id>
Routing slug the hostname fed- + the first 16 characters of the id
Check GET {url}/v1/info, every 10 min see below
Capability panel supported NUTs, from /v1/info modules, from the announcement
Statuses online / degraded / offline / unknown online / offline / announced

Checking a federation

A Cashu mint answers GET /v1/info over HTTPS, which is why probing one is fifteen lines. A federation has no such endpoint: its guardians speak a JSON-RPC dialect over websockets, their addresses are bech32m-encoded inside the invite code, and confirming one is up means being a Fedimint client — decoding the code, opening sockets to a quorum and agreeing a consensus session. Half of that would produce a status less trustworthy than saying nothing.

So the federation check reads fedimint.observer, which already keeps those client connections open and publishes the result at /api/federations as { id, name, invite, health }. Two consequences, and the site carries both rather than hiding them:

  • It is somebody else's check. A federation whose status came from there stores status_source: "fedimint.observer" and its page prints that beside the status, so nothing implies this site opened a socket itself. Point FEDIMINT_OBSERVER_URL somewhere else, or set it empty to disable the lookup entirely.
  • It does not cover everything. A federation the observer does not track gets the announced status: Nostr says it exists, nothing says it runs. It is never online, never offline, has no last_online, no uptime figure and no sparkline. announced sorts between the confirmed-up rows and the confirmed-down ones, because not knowing is not the same as knowing otherwise.

When the observer itself is unreachable, nothing is written at all: a federation confirmed up an hour ago is not demoted because a third party had a bad minute.

What a federation page will not say

There is no Fedimint counterpart to the "melt only", "withdrawals disabled" or "mint frozen" banners, and shared/src/warnings.ts cannot produce one. Those are read out of a mint's own NUT-04 and NUT-05 switches; a federation's modules tag says which parts it runs and nothing about whether any of them is accepting deposits, so the Modules panel lists them and stops there. A federation gets two banners at most: a real check reported its guardians down, or it has been announced for a month and nothing has ever confirmed it.

Adding a third ecosystem

The extension points, in the order you would touch them. Nothing in the review pipeline, the card, the dialog or the feed needs an edit: they are already driven by the values below rather than by a test for Cashu.

  1. A type value and its announcement kind. One entry in ANNOUNCEMENT_KINDS (shared/src/nostr.ts). That alone puts the kind on discovery's subscription list, teaches the review resolver and the /reviews feed which k tag belongs to it, and adds its pill to the ecosystem filter. mints.type is TEXT with no CHECK constraint, so no migration is involved.
  2. Discovery: how to read its announcement. A parser beside parseFedimintAnnouncement and an upsert… beside upsertFedimint, called from runDiscovery. Whatever is type-specific goes in ecosystem_json, which GET /api/mints/:host spreads across the detail payload; the shared columns (name, description, icon_url, status) are filled the same way for everyone.
  3. A probe strategy. A branch in probeAll (api/src/probe.ts) and a checker beside fedimint-observer.ts. If there is no reliable public check, use a pseudo-status like announced and record why — do not infer a status from the announcement.
  4. Pages. A list page and a detail page under web/src/pages/[...locale]/, a …Subject() builder in web/src/lib/review-subject.ts (which is what makes the reviews panel work unchanged), a route in targetPath (web/src/lib/feed-resolve.ts), nav entries in Topbar.astro and Footer.astro, and sitemap entries in src/pages/sitemap.xml.ts.
  5. Copy. A namespace in en.json and in every other catalog, the namespace added to CLIENT_NAMESPACES if an island renders any of it, and a row in web/src/i18n/GLOSSARY.md for every term the ecosystem introduces. pnpm check:i18n warns per locale until every catalog has the keys; until then they fall back to English.

Not in scope, deliberately: LNURL. Nothing in the routes, event kinds, schema values or copy refers to it, and it is planned as a later stage rather than half-built now.

API

Five endpoints, CORS open, no auth. Four read; the fifth writes.

Endpoint Notes
GET /api/health Never cached. 503 when probes are stale, discovery failed, or discovery is starved. Carries the last cycle's per-relay outcome.
GET /api/stats Network counters, memoized 60s in process.
GET /api/mints Everything listed, online first then score descending. ?limit=, ?type=.
GET /api/mints/:host One listing plus its ecosystem's own fields, distribution, uptime and probe history.
POST /api/index Index a mint nobody has announced yet. Rate limited. See below.

/icons/* serves the cached mint icons.

Discovery starvation

For about a year, GET /api/health said ok, every discovery cycle reported ok=true, and the index sat at eight mints. The production RELAYS list did not include the relay carrying the kind 38000/38172 archive, so each backfill read about thirty events, wrote them faithfully, and the nightly build republished the result. Nothing was broken in a way anything measured.

What was missing is that "the cycle completed" and "the cycle read anything" are different claims, and only the first one was being made. Three things now make the second one:

Per-relay attribution. A cycle records, for every relay in RELAYS, whether a socket opened, how many events it sent, and whether it ended in a real EOSE. Counts are taken before cross-relay deduplication, so they say what each relay contributed rather than what happened to be new because of it. The one-line cycle log carries the lot:

INFO discovery cycle mode=backfill events=1528 … starved=false \
  relays=wss://relay.cashumints.space=12 wss://nos.lol=1566 wss://relay.azzamo.net=10 \
          wss://relay.snort.social=18 wss://relay.primal.net=1

Read that line before changing RELAYS. It is also how you find out that most of this network's archive currently sits behind one relay.

A WARN per relay, naming it. A relay that would not connect is warned about on every cycle. A relay that connected and sent nothing is warned about on backfills only — an incremental cycle asking for one interval is supposed to come back empty, and an hourly warning about that would train everyone to skip the line that eventually matters.

WARN discovery relay unreachable relay=wss://relay.example.invalid mode=backfill
WARN discovery relay returned no events relay=wss://relay.azzamo.net mode=backfill

A floor. BACKFILL_MIN_EVENTS, 200 by default. A backfill asks five relays for the entire history of four kinds; on a working relay set that is thousands of events. Under the floor:

ERROR discovery starvation suspected events=31 floor=200 relays=5 silent=4 unreachable=0

and a flag is set that GET /api/health reports as discovery_starved, which forces status to degraded and the response to 503. The flag is sticky across incremental cycles: an hourly cycle finding four events must not clear a starvation a backfill diagnosed. Only the next backfill clears it.

GET /api/health grew four fields for this:

{
  "status": "degraded",
  "discovery_relays": [
    { "url": "wss://relay.cashumints.space", "connected": true, "events": 12, "eose": true },
    { "url": "wss://relay.example.invalid",  "connected": false, "events": 0, "eose": false }
  ],
  "last_discovery_events": 31,
  "last_discovery_mode": "backfill",
  "discovery_starved": true,
  "backfill_min_events": 200
}

A fresh database reports 503 until its first backfill finishes, and that is correct. Before any backfill has run, nothing has confirmed that this deployment's relay list reads anything at all, and answering ok would be the original bug in miniature. In practice it holds cashumints-web.service at its health gate — which is the point: a first deploy should not publish a site built from an empty index. The state is stored in the database rather than in memory for the same reason, so a restart cannot launder a starvation into "no cycle yet".

To silence it deliberately on a deployment that genuinely has less history than this — a private relay, a test rig — set BACKFILL_MIN_EVENTS=1.

Indexing on demand

curl -X POST https://cashumints.space/api/index \
  -H 'content-type: application/json' \
  -d '{"type":"lnurl","input":"mint.600.wtf"}'

type is cashu, fedimint or lnurl; input is a URL, or an invite code for a federation. The site itself calls this from two places: the 404 page, when somebody opens /mint/…, /fedimint/… or /lnurl-mint/… for something this build has never heard of, and the "Write a review" dialog on the three index pages.

What it does, in order, stopping at the first step that settles it:

  1. Already indexed — 200 with the ordinary GET /api/mints/:host payload plus "existing": true. Nothing is probed; a submission is not a reason to re-probe.
  2. Answers as what it claims — one request, the standard PROBE_TIMEOUT_MS. The row is written by the same functions the probe cycle uses, so it is indistinguishable from one the loop created. 201, same payload shape, plus "indexed_from": "probe".
  3. Answers as something else — 422 with error: "wrong_type" and detected_type, which is what lets the dialog offer "this is a Cashu mint, review it there" as a button. A federation's invite code is decoded rather than fetched: there is no endpoint to probe.
  4. Nothing answers — one bounded relay lookup (3s) for a NIP-87 announcement or any review naming the address. Found: the row is written from the announcement with status = offline, which is the mint that rugged last week and is exactly the one somebody wants to review. Nothing anywhere: 404, error: "unverifiable".

Everything indexed this way is an ordinary row afterwards: the probe loop owns it from the next cycle.

Because it fetches an address a stranger chose, it is hardened accordingly (api/src/safe-fetch.ts): https only, DNS resolved and every returned address checked against the private, loopback, link-local, CGNAT and multicast ranges before a socket opens, redirects followed by hand and capped at two with every hop re-checked, a 256KB response cap, and a per-IP limit of INDEX_RATE_LIMIT an hour with a clean 429. Two simultaneous submissions of one address share a single probe. pnpm --filter ./api test:index covers all of it without a network.

The limit needs to be able to tell two visitors apart, which is what include proxy_params; on every proxied location in the nginx block below is for — it is what sets X-Forwarded-For, and without it every submission arrives from the loopback peer and shares one budget. The last entry of that header is the one used, because it is the one the proxy vouched for; the header is ignored entirely when the peer is not loopback.

Ranking

Bayesian weighted rating:

score = (v / (v + m)) * R + (m / (v + m)) * C

v is the review count, R the mint's mean rating, m = 5, and C the prior mean. Offline mints are multiplied by 0.5 and always sorted below every online mint; mints with no review in 180 days are multiplied by 0.9.

C defaults to 3, not the observed global mean. BACKEND.md specifies the global mean, but almost every Cashu review is five stars, so that mean sits near 4.7. Shrinking two above-prior means toward a prior that high cannot reorder them, it only compresses them, and the result is that a single five star review outranks 39 considered ones. The neutral prior is what makes the formula behave as intended. Set SCORE_PRIOR_MEAN=global for the literal specified behaviour. shared/src/score.ts carries the arithmetic, and pnpm --filter ./api test asserts both cases.

Deployment

The unit files and the nginx block quoted below are checked in under deploy/. Those are the copies to edit; what is quoted here is the same text, for reading in context.

Three units and an nginx block. Everything this project runs listens on loopback and runs as the same unprivileged user; nginx terminates TLS and proxies to it, and opens no file belonging to the project.

        :443  nginx  ──►  127.0.0.1:8789   cashumints-site.service   the prerendered site
                     └─►  127.0.0.1:8788   cashumints.service        /api/* and /icons/*

                          cashumints-web.service   oneshot: build, then publish to
                                                   /var/lib/cashumints/web

nginx used to point a root at web/dist instead of proxying. That is the arrangement to avoid, and the reason is worth stating because the failure mode is so quiet: nginx runs as www-data, everything here runs as cashumints, and serving files across that boundary requires every directory from / down to dist to be traversable by a user with no other business in the tree. A home directory left at its default 0700 breaks the whole site, and try_files reports a permission error as a plain miss — so the symptom is a blanket 404, or an internal-redirect loop that ends in a 500, with the cause named nowhere. Proxying removes the boundary: the process that serves the files is the one that built them.

Node

Node 20.18 or newer is enough. Nothing systemd starts reads a .ts file: pnpm build compiles api/src to api/dist, the site server is plain .mjs, and both units run /usr/bin/node against ordinary JavaScript. 20.18 rather than 20.0 only because both ExecStart lines pass --env-file-if-exists, which landed there.

/usr/bin/node --version

That is the version that matters, and it is not the same question as node -v in your shell: a version manager puts its node on the interactive PATH only, while systemd resolves the absolute path in ExecStart. Install Node system-wide rather than pointing ExecStart at a version manager's path, which breaks at the next upgrade and is invisible to ProtectHome.

Why this section used to say 22.18. The API ran src/index.ts directly, on native type stripping, so the host's Node version was a runtime dependency of the service. A deploy onto a host with Node 20 met ERR_UNKNOWN_FILE_EXTENSION, exited in under a second, and was restarted by systemd 464 times over fifteen hours. Every dashboard was green throughout, because there was no dashboard: Restart=on-failure with no start limit is an infinite loop that never reports a failure. Two things changed. The service is compiled, so the host's Node version cannot break it in that way again; and the units now stop after five failures in two minutes and run an OnFailure= alert, so if something else breaks it in some other way, the machine says so. See "Failing loudly".

The version floor that is still 22.18 is the development one — pnpm dev, pnpm seed, pnpm migrate and the api test:* scripts all hand src/*.ts to node. That is a laptop requirement, not a server one.

The API

# /etc/systemd/system/cashumints.service
[Unit]
Description=cashumints.space indexer and API
Wants=network-online.target
After=network-online.target
# Stop after five failures in two minutes rather than restarting forever: a process
# that cannot start will not start on the 4000th attempt either, and `failed` in
# `systemctl status` is a louder signal than a journal scrolling past. Both keys belong
# to [Unit] — under [Service] systemd only warns and ignores them. See "Failing loudly"
# for why the window is 120s and not 60s.
StartLimitIntervalSec=120
StartLimitBurst=5
# And carry that `failed` off the machine. %n is this unit's own name.
OnFailure=cashumints-alert@%n.service

[Service]
Type=simple
User=cashumints
Group=cashumints
WorkingDirectory=/home/cashumints/CashuMints.space/api

# StateDirectory creates /var/lib/cashumints with the service user's ownership.
StateDirectory=cashumints
Environment=NODE_ENV=production
Environment=PORT=8788
Environment=DB_PATH=/var/lib/cashumints/cashumints.db
Environment=ICON_DIR=/var/lib/cashumints/icons
# For Postgres, replace DB_PATH with DATABASE_URL and add After=postgresql.service.
# Keep ICON_DIR either way: cached icons are files, not rows.

# Compiled JavaScript, run by the distribution's own node. See "Node" above for why
# this is not src/index.ts any more.
ExecStart=/usr/bin/node --env-file-if-exists=../.env dist/index.js

Restart=on-failure
RestartSec=5s
# The process finishes its in-flight probe batch and closes the database on SIGTERM.
KillSignal=SIGTERM
TimeoutStopSec=30s
UMask=0027

NoNewPrivileges=true
PrivateTmp=true
PrivateDevices=true
ProtectSystem=strict
ProtectHome=read-only
ReadWritePaths=/var/lib/cashumints
ProtectKernelTunables=true
ProtectKernelModules=true
ProtectControlGroups=true
RestrictAddressFamilies=AF_UNIX AF_INET AF_INET6

[Install]
WantedBy=multi-user.target

Environment= in the unit and --env-file-if-exists=../.env are not interchangeable. Anything set in the unit wins, so a value that must not drift with an edit to .env belongs in the unit; a secret belongs in .env, which is not world-readable.

The site server

web/server.mjs serves the prerendered tree and nothing else — no template, no database handle, no route table. It is a few hundred lines of node:http with no dependencies, and pnpm --filter ./web test covers what it has to get right: a miss answers 404 rather than 200 with a 404-shaped body, a miss under a locale stays in that locale, hashed assets are immutable and markup is not, and nothing outside the root is readable however the path is spelled.

Two behaviours are worth knowing about because they are load-bearing:

  • It refuses to start on a root with no index.html. The failure that prevents is a process that starts cleanly, answers every request with 404 and looks healthy to anything watching the port.
  • It does not serve web/dist. astro build empties dist before it writes, so serving it directly means a rebuild takes the site down for the length of the build. It serves the published copy at WEB_ROOT instead.
# /etc/systemd/system/cashumints-site.service
[Unit]
Description=cashumints.space static site server
Wants=network-online.target
After=network-online.target
StartLimitIntervalSec=120
StartLimitBurst=5
OnFailure=cashumints-alert@%n.service
# Not Requires=cashumints.service: the pages are prerendered, so the site keeps serving
# a correct-as-of-last-build copy while the API is down. Only the islands go quiet.

[Service]
Type=simple
User=cashumints
Group=cashumints
WorkingDirectory=/home/cashumints/CashuMints.space/web

StateDirectory=cashumints
Environment=NODE_ENV=production
Environment=SITE_PORT=8789
Environment=SITE_HOST=127.0.0.1
Environment=WEB_ROOT=/var/lib/cashumints/web
ExecStart=/usr/bin/node server.mjs

Restart=on-failure
RestartSec=5s
KillSignal=SIGTERM
TimeoutStopSec=15s

NoNewPrivileges=true
PrivateTmp=true
PrivateDevices=true
ProtectSystem=strict
# Read-only rather than absent: server.mjs itself lives under /home/cashumints.
ProtectHome=read-only
ReadWritePaths=/var/lib/cashumints
ProtectKernelTunables=true
ProtectKernelModules=true
ProtectControlGroups=true
RestrictAddressFamilies=AF_UNIX AF_INET AF_INET6
RestrictSUIDSGID=true
LockPersonality=true

[Install]
WantedBy=multi-user.target

nginx

The whole public surface. Two upstreams, no root, no try_files, no error_page: a miss is the site server's 404 page, in the right language and with a 404 status, and an nginx error page here would replace it with a blank one and hide which upstream failed.

proxy_cache_path /var/cache/nginx/cashumints levels=1:2 keys_zone=cashumints:10m
                 max_size=256m inactive=10m use_temp_path=off;

# Keep a few connections open to each upstream rather than reconnecting per request.
# Both processes hold idle sockets longer than nginx does, so nginx is always the side
# that closes — the reverse of that race is what produces sporadic 502s under load.
upstream cashumints_site { server 127.0.0.1:8789; keepalive 16; }
upstream cashumints_api  { server 127.0.0.1:8788; keepalive 8; }

server {
  listen 443 ssl http2;
  listen [::]:443 ssl http2;
  server_name cashumints.space;

  ssl_certificate /etc/letsencrypt/live/cashumints.space/fullchain.pem;
  ssl_certificate_key /etc/letsencrypt/live/cashumints.space/privkey.pem;
  include /etc/letsencrypt/options-ssl-nginx.conf;
  ssl_dhparam /etc/letsencrypt/ssl-dhparams.pem;

  # Set here because this is the only part of the stack that knows a request arrived
  # over TLS. Note that add_header replaces rather than merges: any location declaring
  # its own add_header drops this one and has to restate it.
  add_header Strict-Transport-Security "max-age=31536000; includeSubDomains" always;

  # The upstreams send no Content-Encoding, so compression is nginx's to do.
  # gzip_proxied any is required — without it nginx will not compress a proxied response
  # at all. The home page goes out at a fifth of its size.
  gzip on;
  gzip_proxied any;
  gzip_vary on;
  gzip_comp_level 5;
  gzip_min_length 1024;
  gzip_types text/plain text/css text/javascript application/javascript application/json
             application/manifest+json application/xml image/svg+xml;

  # Nothing here accepts an upload; the API's largest body is an 8 KB JSON submission.
  client_max_body_size 16k;

  # Both upstreams are a process on this machine. A slow response is a bug, not a
  # network condition, and failing fast beats holding a worker for a minute.
  proxy_connect_timeout 2s;
  proxy_read_timeout 30s;
  proxy_send_timeout 30s;

  # HTTP/1.1 with an empty Connection header is what makes the keepalive pools work;
  # the default 1.0 opens a new socket per request.
  proxy_http_version 1.1;
  proxy_set_header Connection "";

  # proxy_params (Debian) sets Host, X-Real-IP, X-Forwarded-For and X-Forwarded-Proto.
  # That include is load-bearing: the API's rate limiter reads the last hop of
  # X-Forwarded-For to tell two visitors apart, and without it every request arrives from
  # the loopback peer and shares one bucket. Do not also set X-Forwarded-For beside the
  # include — declaring it twice is what produces nginx's proxy_headers_hash warning.

  # Cache-Control comes from server.mjs: a year and immutable for anything with a
  # content hash in its name, revalidate-every-time for markup. Nothing to restate here.
  location / {
    proxy_pass http://cashumints_site;
    include proxy_params;
  }

  # Health must always reflect the live process. It is the endpoint you page on.
  location = /api/health {
    proxy_pass http://cashumints_api;
    proxy_cache off;
    add_header Strict-Transport-Security "max-age=31536000; includeSubDomains" always;
    add_header Cache-Control "no-store" always;
    include proxy_params;
  }

  # /api/mints and /api/stats change every few minutes but can be asked for by every
  # visitor at once. A micro-cache absorbs that without making the data stale.
  location ~ ^/api/(mints|stats)(?:/|$|\?) {
    proxy_pass http://cashumints_api;
    proxy_cache cashumints;
    proxy_cache_valid 200 30s;
    proxy_cache_valid 404 10s;
    # Serve the previous response while one request refreshes it, so a slow backend
    # never becomes a slow page.
    proxy_cache_use_stale updating error timeout http_500 http_502 http_503;
    proxy_cache_background_update on;
    proxy_cache_lock on;
    add_header Strict-Transport-Security "max-age=31536000; includeSubDomains" always;
    add_header X-Cache-Status $upstream_cache_status always;
    include proxy_params;
  }

  location /api/ {
    proxy_pass http://cashumints_api;
    include proxy_params;
  }

  location /icons/ {
    proxy_pass http://cashumints_api;
    proxy_cache cashumints;
    proxy_cache_valid 200 1d;
    include proxy_params;
  }
}

server {
  listen 80;
  listen [::]:80;
  server_name cashumints.space;

  location /.well-known/acme-challenge/ { root /var/www/html; }
  location / { return 301 https://cashumints.space$request_uri; }
}

/api/ and /icons/ are proxied same-origin because PUBLIC_API_URL is empty, which is the default. Drop those blocks only if you build with it pointing at a separate API host.

nginx 1.25 and later want http2 on; on its own line and warn about the listen … http2 form above; Debian 12 ships 1.22, where the newer form is an unknown directive. The form above is the one that works on both.

Failing loudly

Two separate silences produced today's incident, and they need separate fixes.

The first is a crash loop that never reports a failure. Restart=on-failure with no start limit is an infinite loop by definition: systemd restarts, the process dies, systemd restarts. The unit never reaches failed, so systemctl status stays active (auto-restart), nothing sends anything anywhere, and the only evidence is a journal scrolling past at four lines a second. The API did this 464 times over fifteen hours.

All three units now carry:

StartLimitIntervalSec=120
StartLimitBurst=5
OnFailure=cashumints-alert@%n.service

Why 120 and not 60. RestartSec=5s means five attempts cost a little over twenty seconds of waiting, plus however long each attempt survives before dying. A process that fails slowly — a database connection that times out, a port that takes four seconds to refuse — spreads five failures past a sixty second window, resets the counter, and loops forever anyway. 120s covers the slow case. Both keys belong under [Unit]; systemd takes them under [Service] with only a warning and then ignores them.

Why OnFailure= at all. StartLimitBurst turns the loop into a failed state, which is much better, and is still a state somebody has to go and look at. OnFailure= is what makes the machine speak first. %n expands to the failed unit's own name, which arrives as the template instance in %i.

The second silence is a build that publishes an index it should have refused; that one is the MIN_MINTS_FOR_BUILD gate under "Rebuilds", and the discovery_starved flag under "Discovery starvation" is what feeds it.

The alert unit

# /etc/systemd/system/cashumints-alert@.service
#
# The unit that makes a failure audible.
#
# The other three units each carry `OnFailure=cashumints-alert@%n.service`, so systemd
# starts one of these with the failed unit's name as the instance — `%i` below is
# literally `cashumints.service`, `cashumints-web.service` or `cashumints-site.service`.
#
# Why it exists: the API once crash looped 464 times over fifteen hours and nothing said
# so. `Restart=on-failure` with no start limit is an infinite loop that never reaches a
# `failed` state, so the journal filled with identical lines nobody was reading and
# every signal stayed green. The other half of the fix is StartLimitBurst= in each unit,
# which turns the loop into a failure; this is what carries that failure off the machine.
#
# Install:
#     sudo install -m 0644 deploy/cashumints-alert@.service /etc/systemd/system/
#     sudo install -d -m 0755 /etc/cashumints
#     sudo install -m 0640 -o root -g root deploy/alert.env.example /etc/cashumints/alert.env
#     sudo systemctl daemon-reload
#
# No [Install] section and never enabled: OnFailure= starts it, and a unit that also
# started at boot would page on every reboot.

[Unit]
Description=Notify that %i failed
# No OnFailure= here. An alerter that alerts about its own failure is a loop, and this
# one is written so its worst case is a journal line rather than a retry.

[Service]
Type=oneshot

# The one file an operator edits, and the only reason this unit is configurable at all.
# Absent is a supported state — the leading `-` says so — and then the ExecStart below
# still writes to the journal at ERROR, which is what `systemctl status` and
# `journalctl -p err` read. See alert.env.example.
EnvironmentFile=-/etc/cashumints/alert.env

# So `journalctl -t cashumints-alert` finds every alert, whichever unit triggered it.
SyslogIdentifier=cashumints-alert

# Everything is inside one shell so the "nothing configured" branch is reachable without
# a second unit. The pieces, in order:
#
#   - `printf '<3>…'` on stdout. systemd reads that syslog prefix off a journal stream
#     and files the line at priority 3, ERROR, so `journalctl -p err` is a complete
#     history of failures on a host with no webhook configured at all. `<4>` is warning.
#     A prefix rather than systemd-cat, so the unit needs nothing from the filesystem it
#     has just sandboxed itself away from.
#   - NTFY_URL is a topic URL (https://ntfy.sh/your-topic). It gets a plain-text body
#     naming the failed unit, plus the header names ntfy understands.
#   - WEBHOOK_URL gets a JSON POST instead, for Discord, Slack or anything that speaks
#     `{"content": …}` — every key is sent, so one payload fits all of them.
#   - `--max-time 10` and a `||` fallback on each: an alert that hangs would hold the
#     failed unit's job open, and an alert that fails must not itself become a second
#     failed unit for somebody to notice. The shell ends in `true` for the same reason.
#
# `%i` is the failed unit's name, passed as an argument rather than interpolated into
# the shell text: systemd expands specifiers before /bin/sh ever sees the line, and a
# unit name is not a thing to trust to quoting.
ExecStart=/bin/sh -c '\
  UNIT="$1"; \
  HOST="$(hostname)"; \
  WHEN="$(date -Is)"; \
  TEXT="$UNIT failed on $HOST at $WHEN"; \
  printf "<3>%s\\n" "$TEXT"; \
  if [ -n "$NTFY_URL" ]; then \
    /usr/bin/curl -fsS --max-time 10 \
      -H "Title: cashumints: $UNIT failed" \
      -H "Priority: high" \
      -H "Tags: rotating_light" \
      -d "$TEXT" "$NTFY_URL" >/dev/null \
      || printf "<3>%s\\n" "alert: POST to NTFY_URL failed"; \
  fi; \
  if [ -n "$WEBHOOK_URL" ]; then \
    /usr/bin/curl -fsS --max-time 10 \
      -H "Content-Type: application/json" \
      -d "{\\"unit\\":\\"$UNIT\\",\\"host\\":\\"$HOST\\",\\"at\\":\\"$WHEN\\",\\"text\\":\\"$TEXT\\",\\"content\\":\\"$TEXT\\"}" \
      "$WEBHOOK_URL" >/dev/null \
      || printf "<3>%s\\n" "alert: POST to WEBHOOK_URL failed"; \
  fi; \
  if [ -z "$NTFY_URL" ] && [ -z "$WEBHOOK_URL" ]; then \
    printf "<4>%s\\n" "alert: no NTFY_URL or WEBHOOK_URL in /etc/cashumints/alert.env, journal only"; \
  fi; \
  true' _ %i

# It sends one HTTP request and writes one line. It needs no identity of its own, and
# DynamicUser gives it a throwaway one rather than sharing `nobody` with everything else
# on the host that also could not be bothered to make a user.
DynamicUser=yes
NoNewPrivileges=true
PrivateDevices=true
ProtectSystem=strict
ProtectHome=true
ProtectKernelTunables=true
ProtectKernelModules=true
ProtectControlGroups=true
RestrictAddressFamilies=AF_UNIX AF_INET AF_INET6
RestrictSUIDSGID=true
LockPersonality=true
# An alert that cannot reach the network in ten seconds is not worth a stuck job.
TimeoutStartSec=30

Configuration is one optional file. With it absent, or with both values empty, a failure still lands in the journal at ERROR and is readable with journalctl -p err -t cashumints-alert; the unit is written so that its worst case is a log line rather than a second thing to debug.

# /etc/cashumints/alert.env
#
# Read by cashumints-alert@.service, which systemd starts when any of the three units
# fails. Everything here is optional: with the file absent or both values empty, an
# alert is still written to the journal at ERROR priority and is readable with
#
#     journalctl -p err -t cashumints-alert
#
# Set one or both to have failures leave the machine.
#
# Install it root-owned and not world-readable — a webhook URL is a capability:
#     sudo install -d -m 0755 /etc/cashumints
#     sudo install -m 0640 -o root -g root deploy/alert.env.example /etc/cashumints/alert.env
#     sudo systemctl daemon-reload

# An ntfy topic URL. Free and public at ntfy.sh; pick a topic name nobody will guess,
# because anyone who knows it can read and post to it.
#NTFY_URL=https://ntfy.sh/cashumints-alerts-CHANGE-ME

# Anything that accepts a JSON POST. The body carries `unit`, `host`, `at`, `text` and
# `content` — the last of which is what Discord and most Slack-compatible endpoints read,
# so one payload fits all three.
#WEBHOOK_URL=https://discord.com/api/webhooks/…

Test it without breaking anything:

sudo systemctl start 'cashumints-alert@test.service'
journalctl -t cashumints-alert -n 5 --no-pager

Rebuilds

Mint pages are prerendered, so new mints and new review counts appear at the next build. A nightly rebuild is enough; the site stays correct in between because the islands refresh status and reviews at runtime, and an unbuilt mint still resolves through the client-side fallback on the 404 page.

Publishing is a separate step from building, and the separation is the point: the copy the site server reads is only touched once a build has succeeded, so a failed build leaves the previous site up rather than replacing it with a half-written one.

# /etc/systemd/system/cashumints-web.service
[Unit]
Description=Rebuild the cashumints.space static site
# Every page's data comes from the API over loopback, so the API has to be up.
# Requires= rather than Wants=: a dead API should abort the build, not replace a good
# site with an empty one.
Requires=cashumints.service
After=cashumints.service network-online.target
Wants=network-online.target
OnFailure=cashumints-alert@%n.service

[Service]
Type=oneshot
User=cashumints
Group=cashumints
WorkingDirectory=/home/cashumints/CashuMints.space
StateDirectory=cashumints

Environment=NODE_ENV=production
# Where the build reaches the API. Must match PORT= in cashumints.service.
Environment=API_URL=http://127.0.0.1:8788
Environment=SITE_URL=https://cashumints.space
# Browser-facing origin. Empty means same origin: islands fetch /api/... and nginx
# forwards it. Declared even though it is empty, because systemd's environment wins over
# .env — so what a production build emits cannot drift with an edit to that file.
Environment=PUBLIC_API_URL=

# After= orders the start; it does not wait for the port to accept connections. At boot
# the API is still opening its database and probing, so block until it reports healthy
# rather than letting the first fetch die on ECONNREFUSED. /api/health answers 503 until
# it is genuinely ready, and curl -f treats that as a failure, so the loop keeps waiting.
ExecStartPre=/usr/bin/timeout 90 /bin/sh -c 'until curl -sf -o /dev/null http://127.0.0.1:8788/api/health; do sleep 1; done'
# Check `which pnpm` on the host: a corepack or pnpm-home install sits outside /usr/bin,
# and systemd's PATH does not include it.
ExecStart=/usr/bin/pnpm build
# --delay-updates stages the changed files and renames them in at the end, so the window
# where the tree is a mix of two builds is a rename rather than a whole transfer, and
# --delete-after keeps removals from landing before their replacements. Unchanged files
# — every hashed asset and card, which is nearly all of it — are not touched at all.
ExecStartPost=/usr/bin/rsync -a --delete-after --delay-updates web/dist/ /var/lib/cashumints/web/

# ~500 prerendered pages plus a card per mint. Minutes, not seconds, on a small VPS, and
# TimeoutStartSec is what bounds a Type=oneshot.
TimeoutStartSec=1800
# A nightly rebuild should not starve the API it is reading from.
Nice=10
UMask=0022

NoNewPrivileges=true
PrivateTmp=true
PrivateDevices=true
# ProtectHome is deliberately absent, unlike in the other two units: this one writes
# inside /home/cashumints — web/dist, web/public/og, web/src/generated and the pnpm
# store are all under it.
ProtectSystem=full
ProtectKernelTunables=true
ProtectKernelModules=true
ProtectControlGroups=true
RestrictAddressFamilies=AF_UNIX AF_INET AF_INET6

There is deliberately no [Install] section: a rebuild should be scheduled, not fired on every boot.

# /etc/systemd/system/cashumints-web.timer
[Unit]
Description=Nightly cashumints.space rebuild

[Timer]
OnCalendar=*-*-* 03:30:00
Persistent=true

[Install]
WantedBy=timers.target

First deploy

Order matters once: the site server refuses to start against a root that has no index.html, so the build has to publish before it comes up.

/usr/bin/node --version                        # 20.18 or newer
sudo systemctl enable --now cashumints         # API first: the build reads from it
sudo systemctl start cashumints-web            # build, then publish to /var/lib/cashumints/web
sudo systemctl enable --now cashumints-site    # now it has something to serve
sudo systemctl enable --now cashumints-web.timer
sudo nginx -t && sudo systemctl reload nginx
curl -sI https://cashumints.space | head -1              # 200
curl -sI https://cashumints.space/nope | head -1         # 404, not 200
curl -s  https://cashumints.space/api/health             # status ok

A 502 on / means the site server is down or was never started; a 502 on /api/ means the API is. They fail independently, which is the other thing proxying buys: the prerendered site keeps serving while the API is restarting.

Licence

See LICENSE.

S
Description
No description provided
Readme
6.2 MiB
Languages
TypeScript 55.8%
Astro 28.9%
JavaScript 9.5%
CSS 5.8%