Make a starved discovery cycle say so, in the log and on /api/health.
For about a year the production RELAYS list did not include the relay carrying
the kind 38000/38172 archive. Every backfill read about thirty events, wrote them
faithfully, reported ok=true, and the nightly build republished an index of eight
mints. Nothing measured the difference between "the cycle completed" and "the
cycle read anything", so nothing went red.
Three signals now do:
- Per-relay attribution. queryRelays() replaces pool.querySync(), which merges
every relay into one deduplicated array and throws away who sent what. It
keeps one subscription per relay over the pool's existing sockets and shares
a single alreadyHaveEvent across them, so an event five relays carry is still
verified once; receivedEvent fires before that check, which is what makes the
per-relay count mean "what this relay contributed". The deadline moved out of
each Subscription's own EOSE timer so `eose` means a frame arrived rather than
something timed out.
- A WARN naming any relay that will not connect, on every cycle, and any relay
that connected and sent nothing, on backfills only. An incremental cycle is
supposed to come back empty.
- BACKFILL_MIN_EVENTS, default 200. Under it, ERROR discovery starvation
suspected and a flag health reports as discovery_starved, forcing 503. Sticky
across incremental cycles so an hourly cycle finding four events cannot clear
what a backfill diagnosed; stored in the database so a restart cannot either.
A fresh database is starved until its first backfill lands. That is intended: it
holds the build's health gate rather than publishing a site made from nothing.
Verified against the live relay set — 1528 events, five relays connected, EOSE on
all five, health 200 — and against an unreachable list, which produces the two
WARN lines, the ERROR, and 503.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
65307ba278
commit
9ffa53094d
@@ -585,6 +585,7 @@ so a systemd `Environment=` line or a one-off `PORT=9000 pnpm dev:api` still ove
|
||||
| `DB_POOL_MAX` | `10` | Postgres connections held open. Unused by SQLite. |
|
||||
| `ICON_DIR` | `api/data/icons` | Cached mint icons, served at `/icons/*` |
|
||||
| `RELAYS` | see `shared/src/nostr.ts` | Comma separated relay list |
|
||||
| `BACKFILL_MIN_EVENTS` | `200` | Events a backfill has to read before it counts as one. Under it, discovery logs `ERROR discovery starvation suspected` and health goes 503. See [Starvation](#discovery-starvation). |
|
||||
| `PROBE_INTERVAL_MIN` | `10` | Minutes between probe cycles |
|
||||
| `DISCOVERY_INTERVAL_MIN` | `60` | Minutes between discovery cycles |
|
||||
| `PROBE_CONCURRENCY` | `8` | Mints probed in parallel |
|
||||
@@ -717,7 +718,7 @@ Five endpoints, CORS open, no auth. Four read; the fifth writes.
|
||||
|
||||
| Endpoint | Notes |
|
||||
| -------------------- | ------------------------------------------------------------------ |
|
||||
| `GET /api/health` | Never cached. 503 when probes are stale or discovery failed. |
|
||||
| `GET /api/health` | Never cached. 503 when probes are stale, discovery failed, or discovery is starved. Carries the last cycle's per-relay outcome. |
|
||||
| `GET /api/stats` | Network counters, memoized 60s in process. |
|
||||
| `GET /api/mints` | Everything listed, online first then score descending. `?limit=`, `?type=`. |
|
||||
| `GET /api/mints/:host` | One listing plus its ecosystem's own fields, distribution, uptime and probe history. |
|
||||
@@ -725,6 +726,82 @@ Five endpoints, CORS open, no auth. Four read; the fifth writes.
|
||||
|
||||
`/icons/*` serves the cached mint icons.
|
||||
|
||||
### Discovery starvation
|
||||
|
||||
For about a year, `GET /api/health` said `ok`, every discovery cycle reported
|
||||
`ok=true`, and the index sat at eight mints. The production `RELAYS` list did not
|
||||
include the relay carrying the kind 38000/38172 archive, so each backfill read about
|
||||
thirty events, wrote them faithfully, and the nightly build republished the result.
|
||||
Nothing was broken in a way anything measured.
|
||||
|
||||
What was missing is that "the cycle completed" and "the cycle read anything" are
|
||||
different claims, and only the first one was being made. Three things now make the
|
||||
second one:
|
||||
|
||||
**Per-relay attribution.** A cycle records, for every relay in `RELAYS`, whether a
|
||||
socket opened, how many events it sent, and whether it ended in a real EOSE. Counts are
|
||||
taken before cross-relay deduplication, so they say what each relay contributed rather
|
||||
than what happened to be new because of it. The one-line cycle log carries the lot:
|
||||
|
||||
```
|
||||
INFO discovery cycle mode=backfill events=1528 … starved=false \
|
||||
relays=wss://relay.cashumints.space=12 wss://nos.lol=1566 wss://relay.azzamo.net=10 \
|
||||
wss://relay.snort.social=18 wss://relay.primal.net=1
|
||||
```
|
||||
|
||||
Read that line before changing `RELAYS`. It is also how you find out that most of this
|
||||
network's archive currently sits behind one relay.
|
||||
|
||||
**A WARN per relay, naming it.** A relay that would not connect is warned about on every
|
||||
cycle. A relay that connected and sent nothing is warned about on backfills only — an
|
||||
incremental cycle asking for one interval is *supposed* to come back empty, and an
|
||||
hourly warning about that would train everyone to skip the line that eventually matters.
|
||||
|
||||
```
|
||||
WARN discovery relay unreachable relay=wss://relay.example.invalid mode=backfill
|
||||
WARN discovery relay returned no events relay=wss://relay.azzamo.net mode=backfill
|
||||
```
|
||||
|
||||
**A floor.** `BACKFILL_MIN_EVENTS`, 200 by default. A backfill asks five relays for the
|
||||
entire history of four kinds; on a working relay set that is thousands of events. Under
|
||||
the floor:
|
||||
|
||||
```
|
||||
ERROR discovery starvation suspected events=31 floor=200 relays=5 silent=4 unreachable=0
|
||||
```
|
||||
|
||||
and a flag is set that `GET /api/health` reports as `discovery_starved`, which forces
|
||||
`status` to `degraded` and the response to **503**. The flag is sticky across
|
||||
incremental cycles: an hourly cycle finding four events must not clear a starvation a
|
||||
backfill diagnosed. Only the next backfill clears it.
|
||||
|
||||
`GET /api/health` grew four fields for this:
|
||||
|
||||
```json
|
||||
{
|
||||
"status": "degraded",
|
||||
"discovery_relays": [
|
||||
{ "url": "wss://relay.cashumints.space", "connected": true, "events": 12, "eose": true },
|
||||
{ "url": "wss://relay.example.invalid", "connected": false, "events": 0, "eose": false }
|
||||
],
|
||||
"last_discovery_events": 31,
|
||||
"last_discovery_mode": "backfill",
|
||||
"discovery_starved": true,
|
||||
"backfill_min_events": 200
|
||||
}
|
||||
```
|
||||
|
||||
**A fresh database reports 503 until its first backfill finishes, and that is correct.**
|
||||
Before any backfill has run, nothing has confirmed that this deployment's relay list
|
||||
reads anything at all, and answering `ok` would be the original bug in miniature. In
|
||||
practice it holds `cashumints-web.service` at its health gate — which is the point: a
|
||||
first deploy should not publish a site built from an empty index. The state is stored in
|
||||
the database rather than in memory for the same reason, so a restart cannot launder a
|
||||
starvation into "no cycle yet".
|
||||
|
||||
To silence it deliberately on a deployment that genuinely has less history than this —
|
||||
a private relay, a test rig — set `BACKFILL_MIN_EVENTS=1`.
|
||||
|
||||
### Indexing on demand
|
||||
|
||||
```bash
|
||||
|
||||
Reference in New Issue
Block a user