Make a starved discovery cycle say so, in the log and on /api/health.

For about a year the production RELAYS list did not include the relay carrying
the kind 38000/38172 archive. Every backfill read about thirty events, wrote them
faithfully, reported ok=true, and the nightly build republished an index of eight
mints. Nothing measured the difference between "the cycle completed" and "the
cycle read anything", so nothing went red.

Three signals now do:

  - Per-relay attribution. queryRelays() replaces pool.querySync(), which merges
    every relay into one deduplicated array and throws away who sent what. It
    keeps one subscription per relay over the pool's existing sockets and shares
    a single alreadyHaveEvent across them, so an event five relays carry is still
    verified once; receivedEvent fires before that check, which is what makes the
    per-relay count mean "what this relay contributed". The deadline moved out of
    each Subscription's own EOSE timer so `eose` means a frame arrived rather than
    something timed out.

  - A WARN naming any relay that will not connect, on every cycle, and any relay
    that connected and sent nothing, on backfills only. An incremental cycle is
    supposed to come back empty.

  - BACKFILL_MIN_EVENTS, default 200. Under it, ERROR discovery starvation
    suspected and a flag health reports as discovery_starved, forcing 503. Sticky
    across incremental cycles so an hourly cycle finding four events cannot clear
    what a backfill diagnosed; stored in the database so a restart cannot either.

A fresh database is starved until its first backfill lands. That is intended: it
holds the build's health gate rather than publishing a site made from nothing.

Verified against the live relay set — 1528 events, five relays connected, EOSE on
all five, health 200 — and against an unreachable list, which produces the two
WARN lines, the ERROR, and 503.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
michilis
2026-08-25 16:07:01 +02:00
co-authored by Claude Opus 5
parent 65307ba278
commit 9ffa53094d
6 changed files with 455 additions and 18 deletions
+78 -1
View File
@@ -585,6 +585,7 @@ so a systemd `Environment=` line or a one-off `PORT=9000 pnpm dev:api` still ove
| `DB_POOL_MAX` | `10` | Postgres connections held open. Unused by SQLite. |
| `ICON_DIR` | `api/data/icons` | Cached mint icons, served at `/icons/*` |
| `RELAYS` | see `shared/src/nostr.ts` | Comma separated relay list |
| `BACKFILL_MIN_EVENTS` | `200` | Events a backfill has to read before it counts as one. Under it, discovery logs `ERROR discovery starvation suspected` and health goes 503. See [Starvation](#discovery-starvation). |
| `PROBE_INTERVAL_MIN` | `10` | Minutes between probe cycles |
| `DISCOVERY_INTERVAL_MIN` | `60` | Minutes between discovery cycles |
| `PROBE_CONCURRENCY` | `8` | Mints probed in parallel |
@@ -717,7 +718,7 @@ Five endpoints, CORS open, no auth. Four read; the fifth writes.
| Endpoint | Notes |
| -------------------- | ------------------------------------------------------------------ |
| `GET /api/health` | Never cached. 503 when probes are stale or discovery failed. |
| `GET /api/health` | Never cached. 503 when probes are stale, discovery failed, or discovery is starved. Carries the last cycle's per-relay outcome. |
| `GET /api/stats` | Network counters, memoized 60s in process. |
| `GET /api/mints` | Everything listed, online first then score descending. `?limit=`, `?type=`. |
| `GET /api/mints/:host` | One listing plus its ecosystem's own fields, distribution, uptime and probe history. |
@@ -725,6 +726,82 @@ Five endpoints, CORS open, no auth. Four read; the fifth writes.
`/icons/*` serves the cached mint icons.
### Discovery starvation
For about a year, `GET /api/health` said `ok`, every discovery cycle reported
`ok=true`, and the index sat at eight mints. The production `RELAYS` list did not
include the relay carrying the kind 38000/38172 archive, so each backfill read about
thirty events, wrote them faithfully, and the nightly build republished the result.
Nothing was broken in a way anything measured.
What was missing is that "the cycle completed" and "the cycle read anything" are
different claims, and only the first one was being made. Three things now make the
second one:
**Per-relay attribution.** A cycle records, for every relay in `RELAYS`, whether a
socket opened, how many events it sent, and whether it ended in a real EOSE. Counts are
taken before cross-relay deduplication, so they say what each relay contributed rather
than what happened to be new because of it. The one-line cycle log carries the lot:
```
INFO discovery cycle mode=backfill events=1528 … starved=false \
relays=wss://relay.cashumints.space=12 wss://nos.lol=1566 wss://relay.azzamo.net=10 \
wss://relay.snort.social=18 wss://relay.primal.net=1
```
Read that line before changing `RELAYS`. It is also how you find out that most of this
network's archive currently sits behind one relay.
**A WARN per relay, naming it.** A relay that would not connect is warned about on every
cycle. A relay that connected and sent nothing is warned about on backfills only — an
incremental cycle asking for one interval is *supposed* to come back empty, and an
hourly warning about that would train everyone to skip the line that eventually matters.
```
WARN discovery relay unreachable relay=wss://relay.example.invalid mode=backfill
WARN discovery relay returned no events relay=wss://relay.azzamo.net mode=backfill
```
**A floor.** `BACKFILL_MIN_EVENTS`, 200 by default. A backfill asks five relays for the
entire history of four kinds; on a working relay set that is thousands of events. Under
the floor:
```
ERROR discovery starvation suspected events=31 floor=200 relays=5 silent=4 unreachable=0
```
and a flag is set that `GET /api/health` reports as `discovery_starved`, which forces
`status` to `degraded` and the response to **503**. The flag is sticky across
incremental cycles: an hourly cycle finding four events must not clear a starvation a
backfill diagnosed. Only the next backfill clears it.
`GET /api/health` grew four fields for this:
```json
{
"status": "degraded",
"discovery_relays": [
{ "url": "wss://relay.cashumints.space", "connected": true, "events": 12, "eose": true },
{ "url": "wss://relay.example.invalid", "connected": false, "events": 0, "eose": false }
],
"last_discovery_events": 31,
"last_discovery_mode": "backfill",
"discovery_starved": true,
"backfill_min_events": 200
}
```
**A fresh database reports 503 until its first backfill finishes, and that is correct.**
Before any backfill has run, nothing has confirmed that this deployment's relay list
reads anything at all, and answering `ok` would be the original bug in miniature. In
practice it holds `cashumints-web.service` at its health gate — which is the point: a
first deploy should not publish a site built from an empty index. The state is stored in
the database rather than in memory for the same reason, so a restart cannot launder a
starvation into "no cycle yet".
To silence it deliberately on a deployment that genuinely has less history than this —
a private relay, a test rig — set `BACKFILL_MIN_EVENTS=1`.
### Indexing on demand
```bash