phase-8: run it on Postgres, and find out what that was hiding
Six phases claimed the product runs on SQLite and on Postgres. Nothing had ever run it on Postgres. TEST_DATABASE_URL now points the whole suite at a real server and CI runs both arms. The first run found a bug that would have shipped. better-auth's banned flag is integer 0/1 on SQLite and a real boolean on Postgres, and the code read it as `banned === 1`, so on Postgres an account someone asked us to freeze went on receiving email. Both of those flags are now typed for either dialect and read through isFlagSet. It also found the migration advisory lock being taken on a pool. An advisory lock belongs to the session that took it, so a lock on one pooled connection and an unlock on another leaves it held. Two replicas migrating at once is the ordinary case in k8s and is precisely what it was there to protect. New for the scaled mode: k8s manifests with migrations as an initContainer, one poller in its own worker Deployment rather than one per API replica, and an Ingress that exposes the web app only. The storage drivers finally have tests, S3 included, since scaled mode requires it and it had never been exercised. Two acceptance tests, both checked against a deliberately broken build first: two workers claiming a hundred jobs report 188 claims with SKIP LOCKED removed, and the in-flight request is cut off with the drain wait removed. 340 tests on SQLite, 341 on Postgres, 70 Playwright, rules coverage 100%. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
8c8463924f
commit
37370dd079
@@ -0,0 +1,38 @@
|
||||
# Kubernetes manifests
|
||||
|
||||
Scaled mode (SPEC.md section 15): **Postgres and S3 are required.** SQLite is not
|
||||
supported here and the API refuses the configuration that would need it: a dedicated
|
||||
worker means `JOBS_INLINE=false`, which env validation rejects on SQLite because a second
|
||||
poller against one SQLite file is a corruption waiting to happen. Local file storage is
|
||||
equally unsupported unless every replica mounts the same RWX volume, which S3 exists to
|
||||
avoid.
|
||||
|
||||
Apply in order:
|
||||
|
||||
```bash
|
||||
kubectl apply -f namespace.yaml
|
||||
kubectl apply -f config.yaml # edit first: hostnames, bucket, replicas
|
||||
kubectl apply -f secret.example.yaml # do not commit real values, see the note inside
|
||||
kubectl apply -f api.yaml
|
||||
kubectl apply -f worker.yaml
|
||||
kubectl apply -f web.yaml
|
||||
kubectl apply -f ingress.yaml
|
||||
kubectl apply -f hpa.yaml
|
||||
```
|
||||
|
||||
What each piece is for:
|
||||
|
||||
| File | What it does |
|
||||
|---|---|
|
||||
| `namespace.yaml` | One namespace, so everything can be removed in one command |
|
||||
| `config.yaml` | Non-secret settings, the same keys as `apps/api/.env.example` |
|
||||
| `secret.example.yaml` | The shape of the Secret. Real values come from your secret store |
|
||||
| `api.yaml` | The HTTP API, N replicas, migrations as an initContainer, plus its Service |
|
||||
| `worker.yaml` | The job poller, `ROLE=worker`, safe at N replicas on Postgres |
|
||||
| `web.yaml` | The Next server, plus its Service. Never talks to the database |
|
||||
| `ingress.yaml` | Public traffic reaches the web Service only. `/api` is proxied inside it |
|
||||
| `hpa.yaml` | Scales the API on CPU. The worker is not autoscaled; see the note in it |
|
||||
|
||||
Every API and worker pod runs `db:migrate` before it serves. That is safe: the migration
|
||||
takes a Postgres advisory lock on one pinned connection, so replicas starting together
|
||||
queue rather than race.
|
||||
@@ -0,0 +1,84 @@
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
name: impuestos-api
|
||||
namespace: impuestos
|
||||
labels: { app: impuestos, component: api }
|
||||
spec:
|
||||
replicas: 2
|
||||
selector:
|
||||
matchLabels: { app: impuestos, component: api }
|
||||
template:
|
||||
metadata:
|
||||
labels: { app: impuestos, component: api }
|
||||
spec:
|
||||
# 30s: the process drains in-flight requests for up to 25s before it exits.
|
||||
terminationGracePeriodSeconds: 30
|
||||
initContainers:
|
||||
# Every replica runs this. Concurrent runs are safe: the migration holds a Postgres
|
||||
# advisory lock on one pinned connection, so the second pod waits for the first.
|
||||
- name: migrate
|
||||
image: ghcr.io/example/impuestos-api:latest
|
||||
command: ['node', 'dist/db/migrate.cli.js']
|
||||
envFrom:
|
||||
- configMapRef: { name: impuestos-config }
|
||||
- secretRef: { name: impuestos-secrets }
|
||||
resources:
|
||||
requests: { cpu: 50m, memory: 128Mi }
|
||||
limits: { memory: 256Mi }
|
||||
containers:
|
||||
- name: api
|
||||
image: ghcr.io/example/impuestos-api:latest
|
||||
ports:
|
||||
- name: http
|
||||
containerPort: 4000
|
||||
env:
|
||||
- name: ROLE
|
||||
value: 'server'
|
||||
# The dedicated worker Deployment does the jobs. An API replica that also
|
||||
# polled would multiply the pollers by the replica count.
|
||||
- name: JOBS_INLINE
|
||||
value: 'false'
|
||||
envFrom:
|
||||
- configMapRef: { name: impuestos-config }
|
||||
- secretRef: { name: impuestos-secrets }
|
||||
# Liveness answers as long as the process is alive; readiness also checks the
|
||||
# database and the storage driver, and goes false the moment a drain begins.
|
||||
livenessProbe:
|
||||
httpGet: { path: /healthz, port: http }
|
||||
initialDelaySeconds: 5
|
||||
periodSeconds: 10
|
||||
readinessProbe:
|
||||
httpGet: { path: /readyz, port: http }
|
||||
initialDelaySeconds: 2
|
||||
periodSeconds: 5
|
||||
failureThreshold: 2
|
||||
resources:
|
||||
requests: { cpu: 100m, memory: 256Mi }
|
||||
limits: { memory: 512Mi }
|
||||
volumeMounts:
|
||||
- name: tmp
|
||||
mountPath: /tmp
|
||||
securityContext:
|
||||
allowPrivilegeEscalation: false
|
||||
readOnlyRootFilesystem: true
|
||||
capabilities: { drop: ['ALL'] }
|
||||
volumes:
|
||||
# The root filesystem is read only, so the one place anything may be written is an
|
||||
# empty directory that dies with the pod. Nothing durable belongs here: uploads go
|
||||
# through the storage driver to S3.
|
||||
- name: tmp
|
||||
emptyDir: {}
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Service
|
||||
metadata:
|
||||
name: impuestos-api
|
||||
namespace: impuestos
|
||||
spec:
|
||||
type: ClusterIP
|
||||
selector: { app: impuestos, component: api }
|
||||
ports:
|
||||
- name: http
|
||||
port: 4000
|
||||
targetPort: http
|
||||
@@ -0,0 +1,47 @@
|
||||
# Non-secret configuration. Every key here exists in apps/api/.env.example with the same
|
||||
# meaning; nothing is invented for Kubernetes.
|
||||
apiVersion: v1
|
||||
kind: ConfigMap
|
||||
metadata:
|
||||
name: impuestos-config
|
||||
namespace: impuestos
|
||||
data:
|
||||
NODE_ENV: 'production'
|
||||
PORT: '4000'
|
||||
|
||||
# The origin a person types. Cookies are issued for it and links in email point at it.
|
||||
APP_PUBLIC_URL: 'https://impuestos.example'
|
||||
BETTER_AUTH_URL: 'https://impuestos.example'
|
||||
|
||||
# Postgres connections per pod. Multiply by (api replicas + worker replicas) and keep the
|
||||
# total under the server's max_connections.
|
||||
DATABASE_POOL_MAX: '10'
|
||||
|
||||
JOBS_POLL_INTERVAL_MS: '2000'
|
||||
JOBS_STALE_MINUTES: '10'
|
||||
|
||||
STORAGE_DRIVER: 's3'
|
||||
S3_REGION: 'us-east-1'
|
||||
S3_BUCKET: 'impuestos-comprobantes'
|
||||
S3_FORCE_PATH_STYLE: 'true'
|
||||
# Leave empty for AWS; set it for MinIO or another S3 compatible service.
|
||||
S3_ENDPOINT: ''
|
||||
|
||||
OCR_MODEL: 'claude-sonnet-4-6'
|
||||
DEFAULT_LOCALE: 'es'
|
||||
|
||||
SMTP_PORT: '587'
|
||||
SMTP_FROM: 'avisos@impuestos.example'
|
||||
# RFC 8292 wants an https: or mailto: URL. Without one, push switches itself off.
|
||||
PUSH_VAPID_SUBJECT: 'mailto:soporte@impuestos.example'
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: ConfigMap
|
||||
metadata:
|
||||
name: impuestos-web-config
|
||||
namespace: impuestos
|
||||
data:
|
||||
NODE_ENV: 'production'
|
||||
# Cluster DNS for the api Service. The browser never sees this address.
|
||||
API_INTERNAL_URL: 'http://impuestos-api:4000'
|
||||
NEXT_PUBLIC_DEFAULT_LOCALE: 'es'
|
||||
@@ -0,0 +1,42 @@
|
||||
# The API scales on CPU. Stateless by construction: no local writes outside the storage
|
||||
# driver, no in-memory cache that changes an answer, so a new replica behaves like an old
|
||||
# one from its first request.
|
||||
apiVersion: autoscaling/v2
|
||||
kind: HorizontalPodAutoscaler
|
||||
metadata:
|
||||
name: impuestos-api
|
||||
namespace: impuestos
|
||||
spec:
|
||||
scaleTargetRef:
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
name: impuestos-api
|
||||
minReplicas: 2
|
||||
maxReplicas: 10
|
||||
metrics:
|
||||
- type: Resource
|
||||
resource:
|
||||
name: cpu
|
||||
target:
|
||||
type: Utilization
|
||||
averageUtilization: 70
|
||||
---
|
||||
apiVersion: autoscaling/v2
|
||||
kind: HorizontalPodAutoscaler
|
||||
metadata:
|
||||
name: impuestos-web
|
||||
namespace: impuestos
|
||||
spec:
|
||||
scaleTargetRef:
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
name: impuestos-web
|
||||
minReplicas: 2
|
||||
maxReplicas: 6
|
||||
metrics:
|
||||
- type: Resource
|
||||
resource:
|
||||
name: cpu
|
||||
target:
|
||||
type: Utilization
|
||||
averageUtilization: 70
|
||||
@@ -0,0 +1,27 @@
|
||||
# Public traffic reaches the web Service and nothing else. The API is not exposed: the
|
||||
# browser talks to one origin, the web server forwards /api inside the cluster, and the
|
||||
# session cookie stays first party. SPEC.md is explicit that there is no CORS in v1, and
|
||||
# an Ingress that published the API would be the thing that made CORS necessary.
|
||||
apiVersion: networking.k8s.io/v1
|
||||
kind: Ingress
|
||||
metadata:
|
||||
name: impuestos
|
||||
namespace: impuestos
|
||||
annotations:
|
||||
# A scan is a photograph.
|
||||
nginx.ingress.kubernetes.io/proxy-body-size: '12m'
|
||||
spec:
|
||||
ingressClassName: nginx
|
||||
tls:
|
||||
- hosts: ['impuestos.example']
|
||||
secretName: impuestos-tls
|
||||
rules:
|
||||
- host: impuestos.example
|
||||
http:
|
||||
paths:
|
||||
- path: /
|
||||
pathType: Prefix
|
||||
backend:
|
||||
service:
|
||||
name: impuestos-web
|
||||
port: { name: http }
|
||||
@@ -0,0 +1,4 @@
|
||||
apiVersion: v1
|
||||
kind: Namespace
|
||||
metadata:
|
||||
name: impuestos
|
||||
@@ -0,0 +1,29 @@
|
||||
# The shape only. Do not commit real values: point your secret store at this name instead,
|
||||
# or create it once by hand:
|
||||
#
|
||||
# kubectl -n impuestos create secret generic impuestos-secrets \
|
||||
# --from-literal=DATABASE_URL='postgres://user:pass@host:5432/impuestos' \
|
||||
# --from-literal=BETTER_AUTH_SECRET="$(openssl rand -base64 32)" \
|
||||
# --from-literal=S3_ACCESS_KEY_ID=... --from-literal=S3_SECRET_ACCESS_KEY=...
|
||||
#
|
||||
# Optional keys may be left out entirely. The API degrades honestly without them: no
|
||||
# ANTHROPIC_API_KEY sends a scan with no QR to the manual form, no SMTP_HOST logs
|
||||
# verification codes to stdout, and no VAPID keys hides push everywhere in the UI.
|
||||
apiVersion: v1
|
||||
kind: Secret
|
||||
metadata:
|
||||
name: impuestos-secrets
|
||||
namespace: impuestos
|
||||
type: Opaque
|
||||
stringData:
|
||||
DATABASE_URL: 'postgres://impuestos:change-me@postgres:5432/impuestos'
|
||||
BETTER_AUTH_SECRET: 'change-me-at-least-32-characters-long'
|
||||
S3_ACCESS_KEY_ID: 'change-me'
|
||||
S3_SECRET_ACCESS_KEY: 'change-me'
|
||||
ANTHROPIC_API_KEY: ''
|
||||
PUSH_VAPID_PUBLIC_KEY: ''
|
||||
PUSH_VAPID_PRIVATE_KEY: ''
|
||||
SMTP_HOST: ''
|
||||
SMTP_USER: ''
|
||||
SMTP_PASS: ''
|
||||
TELEGRAM_BOT_TOKEN: ''
|
||||
@@ -0,0 +1,64 @@
|
||||
# The Next server. It holds no secrets and never opens a database connection: it renders
|
||||
# and proxies /api to the api Service.
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
name: impuestos-web
|
||||
namespace: impuestos
|
||||
labels: { app: impuestos, component: web }
|
||||
spec:
|
||||
replicas: 2
|
||||
selector:
|
||||
matchLabels: { app: impuestos, component: web }
|
||||
template:
|
||||
metadata:
|
||||
labels: { app: impuestos, component: web }
|
||||
spec:
|
||||
terminationGracePeriodSeconds: 30
|
||||
containers:
|
||||
- name: web
|
||||
image: ghcr.io/example/impuestos-web:latest
|
||||
ports:
|
||||
- name: http
|
||||
containerPort: 3000
|
||||
envFrom:
|
||||
- configMapRef: { name: impuestos-web-config }
|
||||
livenessProbe:
|
||||
httpGet: { path: /healthz, port: http }
|
||||
initialDelaySeconds: 5
|
||||
periodSeconds: 10
|
||||
readinessProbe:
|
||||
httpGet: { path: /healthz, port: http }
|
||||
initialDelaySeconds: 2
|
||||
periodSeconds: 5
|
||||
resources:
|
||||
requests: { cpu: 100m, memory: 256Mi }
|
||||
limits: { memory: 512Mi }
|
||||
volumeMounts:
|
||||
- name: tmp
|
||||
mountPath: /tmp
|
||||
# Next writes its own cache under the app directory at runtime.
|
||||
- name: next-cache
|
||||
mountPath: /app/apps/web/.next/cache
|
||||
securityContext:
|
||||
allowPrivilegeEscalation: false
|
||||
readOnlyRootFilesystem: true
|
||||
capabilities: { drop: ['ALL'] }
|
||||
volumes:
|
||||
- name: tmp
|
||||
emptyDir: {}
|
||||
- name: next-cache
|
||||
emptyDir: {}
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Service
|
||||
metadata:
|
||||
name: impuestos-web
|
||||
namespace: impuestos
|
||||
spec:
|
||||
type: ClusterIP
|
||||
selector: { app: impuestos, component: web }
|
||||
ports:
|
||||
- name: http
|
||||
port: 3000
|
||||
targetPort: http
|
||||
@@ -0,0 +1,62 @@
|
||||
# The job poller, on its own so it can be sized and restarted without touching the API.
|
||||
#
|
||||
# Safe at any replica count on Postgres: claiming uses FOR UPDATE SKIP LOCKED, which is
|
||||
# covered by the hundred-job two-worker test in apps/api/src/modules/jobs/scale.test.ts.
|
||||
# One is the default because the queue is small; raise it when the bandeja backs up.
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
name: impuestos-worker
|
||||
namespace: impuestos
|
||||
labels: { app: impuestos, component: worker }
|
||||
spec:
|
||||
replicas: 1
|
||||
selector:
|
||||
matchLabels: { app: impuestos, component: worker }
|
||||
template:
|
||||
metadata:
|
||||
labels: { app: impuestos, component: worker }
|
||||
spec:
|
||||
# A job that is mid-flight when the pod goes is recovered by the stale sweep rather
|
||||
# than lost, but finishing is cheaper than recovering.
|
||||
terminationGracePeriodSeconds: 30
|
||||
containers:
|
||||
- name: worker
|
||||
image: ghcr.io/example/impuestos-api:latest
|
||||
ports:
|
||||
- name: http
|
||||
containerPort: 4000
|
||||
env:
|
||||
# ROLE=worker is what starts the poller; JOBS_INLINE only decides whether a
|
||||
# server also runs one.
|
||||
- name: ROLE
|
||||
value: 'worker'
|
||||
envFrom:
|
||||
- configMapRef: { name: impuestos-config }
|
||||
- secretRef: { name: impuestos-secrets }
|
||||
# A worker still serves /healthz and /readyz, which is how the cluster knows the
|
||||
# poller has a database to poll. It is not behind a Service.
|
||||
livenessProbe:
|
||||
httpGet: { path: /healthz, port: http }
|
||||
initialDelaySeconds: 5
|
||||
periodSeconds: 10
|
||||
readinessProbe:
|
||||
httpGet: { path: /readyz, port: http }
|
||||
initialDelaySeconds: 2
|
||||
periodSeconds: 10
|
||||
resources:
|
||||
requests: { cpu: 100m, memory: 256Mi }
|
||||
limits: { memory: 512Mi }
|
||||
volumeMounts:
|
||||
- name: tmp
|
||||
mountPath: /tmp
|
||||
securityContext:
|
||||
allowPrivilegeEscalation: false
|
||||
readOnlyRootFilesystem: true
|
||||
capabilities: { drop: ['ALL'] }
|
||||
volumes:
|
||||
# The root filesystem is read only, so the one place anything may be written is an
|
||||
# empty directory that dies with the pod. Nothing durable belongs here: uploads go
|
||||
# through the storage driver to S3.
|
||||
- name: tmp
|
||||
emptyDir: {}
|
||||
Reference in New Issue
Block a user