Make a starved discovery cycle say so, in the log and on /api/health.
For about a year the production RELAYS list did not include the relay carrying
the kind 38000/38172 archive. Every backfill read about thirty events, wrote them
faithfully, reported ok=true, and the nightly build republished an index of eight
mints. Nothing measured the difference between "the cycle completed" and "the
cycle read anything", so nothing went red.
Three signals now do:
- Per-relay attribution. queryRelays() replaces pool.querySync(), which merges
every relay into one deduplicated array and throws away who sent what. It
keeps one subscription per relay over the pool's existing sockets and shares
a single alreadyHaveEvent across them, so an event five relays carry is still
verified once; receivedEvent fires before that check, which is what makes the
per-relay count mean "what this relay contributed". The deadline moved out of
each Subscription's own EOSE timer so `eose` means a frame arrived rather than
something timed out.
- A WARN naming any relay that will not connect, on every cycle, and any relay
that connected and sent nothing, on backfills only. An incremental cycle is
supposed to come back empty.
- BACKFILL_MIN_EVENTS, default 200. Under it, ERROR discovery starvation
suspected and a flag health reports as discovery_starved, forcing 503. Sticky
across incremental cycles so an hourly cycle finding four events cannot clear
what a backfill diagnosed; stored in the database so a restart cannot either.
A fresh database is starved until its first backfill lands. That is intended: it
holds the build's health gate rather than publishing a site made from nothing.
Verified against the live relay set — 1528 events, five relays connected, EOSE on
all five, health 200 — and against an unreachable list, which produces the two
WARN lines, the ERROR, and 503.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
65307ba278
commit
9ffa53094d
@@ -112,6 +112,20 @@ export const config = {
|
||||
iconDir: process.env['ICON_DIR'] ?? path.join(apiRoot, 'data', 'icons'),
|
||||
relays: (process.env['RELAYS']?.split(',').map((r) => r.trim()).filter(Boolean) ??
|
||||
[...DEFAULT_RELAYS]) as string[],
|
||||
/**
|
||||
* The floor a backfill cycle has to clear before it counts as a real read.
|
||||
*
|
||||
* For about a year this deployment's RELAYS list did not include the relay carrying
|
||||
* the kind 38000/38172 archive. Every backfill returned about thirty events, wrote
|
||||
* them, reported ok=true, and the index sat at eight mints while every health signal
|
||||
* stayed green. A backfill asks five relays for the whole history of four kinds; on a
|
||||
* working relay set it comes back with thousands. Anything under this is not a quiet
|
||||
* network, it is a misconfigured one, and it says so in the log and on /api/health.
|
||||
*
|
||||
* Raise it on a deployment that genuinely has more history, lower it for a local
|
||||
* test relay. It is deliberately not zero-able: set it to 1 if you mean "off".
|
||||
*/
|
||||
backfillMinEvents: int('BACKFILL_MIN_EVENTS', 200),
|
||||
probeIntervalMin: int('PROBE_INTERVAL_MIN', 10),
|
||||
discoveryIntervalMin: int('DISCOVERY_INTERVAL_MIN', 60),
|
||||
probeConcurrency: int('PROBE_CONCURRENCY', 8),
|
||||
|
||||
Reference in New Issue
Block a user