Files
MichilisandClaude Opus 5 37370dd079
CI / types, lint and rules coverage (push) Canceled after 0s
CI / tests (postgres) (push) Canceled after 0s
CI / tests (sqlite) (push) Canceled after 0s
CI / playwright (push) Canceled after 0s
phase-8: run it on Postgres, and find out what that was hiding
Six phases claimed the product runs on SQLite and on Postgres. Nothing
had ever run it on Postgres. TEST_DATABASE_URL now points the whole
suite at a real server and CI runs both arms.

The first run found a bug that would have shipped. better-auth's banned
flag is integer 0/1 on SQLite and a real boolean on Postgres, and the
code read it as `banned === 1`, so on Postgres an account someone asked
us to freeze went on receiving email. Both of those flags are now typed
for either dialect and read through isFlagSet.

It also found the migration advisory lock being taken on a pool. An
advisory lock belongs to the session that took it, so a lock on one
pooled connection and an unlock on another leaves it held. Two replicas
migrating at once is the ordinary case in k8s and is precisely what it
was there to protect.

New for the scaled mode: k8s manifests with migrations as an
initContainer, one poller in its own worker Deployment rather than one
per API replica, and an Ingress that exposes the web app only. The
storage drivers finally have tests, S3 included, since scaled mode
requires it and it had never been exercised.

Two acceptance tests, both checked against a deliberately broken build
first: two workers claiming a hundred jobs report 188 claims with SKIP
LOCKED removed, and the in-flight request is cut off with the drain wait
removed.

340 tests on SQLite, 341 on Postgres, 70 Playwright, rules coverage 100%.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 22:03:13 +00:00

39 lines
1.7 KiB
Markdown

# Kubernetes manifests
Scaled mode (SPEC.md section 15): **Postgres and S3 are required.** SQLite is not
supported here and the API refuses the configuration that would need it: a dedicated
worker means `JOBS_INLINE=false`, which env validation rejects on SQLite because a second
poller against one SQLite file is a corruption waiting to happen. Local file storage is
equally unsupported unless every replica mounts the same RWX volume, which S3 exists to
avoid.
Apply in order:
```bash
kubectl apply -f namespace.yaml
kubectl apply -f config.yaml # edit first: hostnames, bucket, replicas
kubectl apply -f secret.example.yaml # do not commit real values, see the note inside
kubectl apply -f api.yaml
kubectl apply -f worker.yaml
kubectl apply -f web.yaml
kubectl apply -f ingress.yaml
kubectl apply -f hpa.yaml
```
What each piece is for:
| File | What it does |
|---|---|
| `namespace.yaml` | One namespace, so everything can be removed in one command |
| `config.yaml` | Non-secret settings, the same keys as `apps/api/.env.example` |
| `secret.example.yaml` | The shape of the Secret. Real values come from your secret store |
| `api.yaml` | The HTTP API, N replicas, migrations as an initContainer, plus its Service |
| `worker.yaml` | The job poller, `ROLE=worker`, safe at N replicas on Postgres |
| `web.yaml` | The Next server, plus its Service. Never talks to the database |
| `ingress.yaml` | Public traffic reaches the web Service only. `/api` is proxied inside it |
| `hpa.yaml` | Scales the API on CPU. The worker is not autoscaled; see the note in it |
Every API and worker pod runs `db:migrate` before it serves. That is safe: the migration
takes a Postgres advisory lock on one pinned connection, so replicas starting together
queue rather than race.