Make a starved discovery cycle say so, in the log and on /api/health.

For about a year the production RELAYS list did not include the relay carrying
the kind 38000/38172 archive. Every backfill read about thirty events, wrote them
faithfully, reported ok=true, and the nightly build republished an index of eight
mints. Nothing measured the difference between "the cycle completed" and "the
cycle read anything", so nothing went red.

Three signals now do:

  - Per-relay attribution. queryRelays() replaces pool.querySync(), which merges
    every relay into one deduplicated array and throws away who sent what. It
    keeps one subscription per relay over the pool's existing sockets and shares
    a single alreadyHaveEvent across them, so an event five relays carry is still
    verified once; receivedEvent fires before that check, which is what makes the
    per-relay count mean "what this relay contributed". The deadline moved out of
    each Subscription's own EOSE timer so `eose` means a frame arrived rather than
    something timed out.

  - A WARN naming any relay that will not connect, on every cycle, and any relay
    that connected and sent nothing, on backfills only. An incremental cycle is
    supposed to come back empty.

  - BACKFILL_MIN_EVENTS, default 200. Under it, ERROR discovery starvation
    suspected and a flag health reports as discovery_starved, forcing 503. Sticky
    across incremental cycles so an hourly cycle finding four events cannot clear
    what a backfill diagnosed; stored in the database so a restart cannot either.

A fresh database is starved until its first backfill lands. That is intended: it
holds the build's health gate rather than publishing a site made from nothing.

Verified against the live relay set — 1528 events, five relays connected, EOSE on
all five, health 200 — and against an unreachable list, which produces the two
WARN lines, the ERROR, and 503.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
michilis
2026-08-25 16:07:01 +02:00
co-authored by Claude Opus 5
parent 65307ba278
commit 9ffa53094d
6 changed files with 455 additions and 18 deletions
+31 -3
View File
@@ -16,6 +16,7 @@ import {
} from '@cashumints/shared';
import { config, startedAt } from './config.ts';
import { getDb, getStateNumber, getState } from './db.ts';
import { lastDiscoveryReport } from './discovery.ts';
import { mintByHost, parseEcosystem, type MintRow } from './mints.ts';
/**
@@ -379,28 +380,55 @@ export function resetStatsCache(): void {
statsCache = null;
}
/** Health bypasses the stats cache: it is the endpoint you page on. */
/**
* Health bypasses the stats cache: it is the endpoint you page on.
*
* Three things can degrade it, and they are three different failures:
*
* probeStale nothing has checked a mint in three intervals
* !discoveryOk the last discovery cycle threw
* report.starved the last backfill read less than BACKFILL_MIN_EVENTS
*
* The third is the one added after the postmortem, and it is the only one that would
* have caught a year of the index sitting at eight mints: the cycles were completing,
* `ok` was true, the probes were fresh, and the relay list simply did not contain the
* relay holding the archive. "Ran without throwing" is not the same claim as "read
* anything", and only the second one is worth a green light.
*
* A deployment with an empty database reports degraded until its first backfill lands,
* because until then nothing has confirmed the relay set reads anything at all. That is
* intended: it holds `cashumints-web.service` at its health gate rather than letting it
* publish a site built from nothing.
*/
export async function getHealth(): Promise<Health> {
const now = Math.floor(Date.now() / 1000);
const db = await getDb();
const [lastProbe, lastDiscovery, discoveryOkRaw, tracked] = await Promise.all([
const [lastProbe, lastDiscovery, discoveryOkRaw, tracked, report] = await Promise.all([
getStateNumber('last_probe_at'),
getStateNumber('last_discovery_at'),
getState('last_discovery_ok'),
db.get<{ n: number }>('SELECT COUNT(*) AS n FROM mints'),
lastDiscoveryReport(),
]);
const discoveryOk = discoveryOkRaw !== '0';
const staleAfter = config.probeIntervalMin * 60 * 3;
const probeStale = lastProbe === null || now - lastProbe > staleAfter;
// No report at all is starvation by default: see the note above.
const starved = report?.starved ?? true;
return {
status: probeStale || !discoveryOk ? 'degraded' : 'ok',
status: probeStale || !discoveryOk || starved ? 'degraded' : 'ok',
uptime_s: now - startedAt,
last_probe_at: lastProbe,
last_discovery_at: lastDiscovery,
mints_tracked: tracked?.n ?? 0,
updated_at: now,
discovery_relays: report?.relays ?? [],
last_discovery_events: report?.events ?? null,
last_discovery_mode: report?.mode ?? null,
discovery_starved: starved,
backfill_min_events: config.backfillMinEvents,
};
}