phase-8: run it on Postgres, and find out what that was hiding
CI / types, lint and rules coverage (push) Canceled after 0s
CI / tests (postgres) (push) Canceled after 0s
CI / tests (sqlite) (push) Canceled after 0s
CI / playwright (push) Canceled after 0s

Six phases claimed the product runs on SQLite and on Postgres. Nothing
had ever run it on Postgres. TEST_DATABASE_URL now points the whole
suite at a real server and CI runs both arms.

The first run found a bug that would have shipped. better-auth's banned
flag is integer 0/1 on SQLite and a real boolean on Postgres, and the
code read it as `banned === 1`, so on Postgres an account someone asked
us to freeze went on receiving email. Both of those flags are now typed
for either dialect and read through isFlagSet.

It also found the migration advisory lock being taken on a pool. An
advisory lock belongs to the session that took it, so a lock on one
pooled connection and an unlock on another leaves it held. Two replicas
migrating at once is the ordinary case in k8s and is precisely what it
was there to protect.

New for the scaled mode: k8s manifests with migrations as an
initContainer, one poller in its own worker Deployment rather than one
per API replica, and an Ingress that exposes the web app only. The
storage drivers finally have tests, S3 included, since scaled mode
requires it and it had never been exercised.

Two acceptance tests, both checked against a deliberately broken build
first: two workers claiming a hundred jobs report 188 claims with SKIP
LOCKED removed, and the in-flight request is cut off with the drain wait
removed.

340 tests on SQLite, 341 on Postgres, 70 Playwright, rules coverage 100%.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Michilis
2026-09-06 22:03:13 +00:00
co-authored by Claude Opus 5
parent 8c8463924f
commit 37370dd079
29 changed files with 1349 additions and 43 deletions
+62
View File
@@ -0,0 +1,62 @@
# The job poller, on its own so it can be sized and restarted without touching the API.
#
# Safe at any replica count on Postgres: claiming uses FOR UPDATE SKIP LOCKED, which is
# covered by the hundred-job two-worker test in apps/api/src/modules/jobs/scale.test.ts.
# One is the default because the queue is small; raise it when the bandeja backs up.
apiVersion: apps/v1
kind: Deployment
metadata:
name: impuestos-worker
namespace: impuestos
labels: { app: impuestos, component: worker }
spec:
replicas: 1
selector:
matchLabels: { app: impuestos, component: worker }
template:
metadata:
labels: { app: impuestos, component: worker }
spec:
# A job that is mid-flight when the pod goes is recovered by the stale sweep rather
# than lost, but finishing is cheaper than recovering.
terminationGracePeriodSeconds: 30
containers:
- name: worker
image: ghcr.io/example/impuestos-api:latest
ports:
- name: http
containerPort: 4000
env:
# ROLE=worker is what starts the poller; JOBS_INLINE only decides whether a
# server also runs one.
- name: ROLE
value: 'worker'
envFrom:
- configMapRef: { name: impuestos-config }
- secretRef: { name: impuestos-secrets }
# A worker still serves /healthz and /readyz, which is how the cluster knows the
# poller has a database to poll. It is not behind a Service.
livenessProbe:
httpGet: { path: /healthz, port: http }
initialDelaySeconds: 5
periodSeconds: 10
readinessProbe:
httpGet: { path: /readyz, port: http }
initialDelaySeconds: 2
periodSeconds: 10
resources:
requests: { cpu: 100m, memory: 256Mi }
limits: { memory: 512Mi }
volumeMounts:
- name: tmp
mountPath: /tmp
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities: { drop: ['ALL'] }
volumes:
# The root filesystem is read only, so the one place anything may be written is an
# empty directory that dies with the pod. Nothing durable belongs here: uploads go
# through the storage driver to S3.
- name: tmp
emptyDir: {}