phase-8: run it on Postgres, and find out what that was hiding
CI / types, lint and rules coverage (push) Canceled after 0s
CI / tests (postgres) (push) Canceled after 0s
CI / tests (sqlite) (push) Canceled after 0s
CI / playwright (push) Canceled after 0s

Six phases claimed the product runs on SQLite and on Postgres. Nothing
had ever run it on Postgres. TEST_DATABASE_URL now points the whole
suite at a real server and CI runs both arms.

The first run found a bug that would have shipped. better-auth's banned
flag is integer 0/1 on SQLite and a real boolean on Postgres, and the
code read it as `banned === 1`, so on Postgres an account someone asked
us to freeze went on receiving email. Both of those flags are now typed
for either dialect and read through isFlagSet.

It also found the migration advisory lock being taken on a pool. An
advisory lock belongs to the session that took it, so a lock on one
pooled connection and an unlock on another leaves it held. Two replicas
migrating at once is the ordinary case in k8s and is precisely what it
was there to protect.

New for the scaled mode: k8s manifests with migrations as an
initContainer, one poller in its own worker Deployment rather than one
per API replica, and an Ingress that exposes the web app only. The
storage drivers finally have tests, S3 included, since scaled mode
requires it and it had never been exercised.

Two acceptance tests, both checked against a deliberately broken build
first: two workers claiming a hundred jobs report 188 claims with SKIP
LOCKED removed, and the in-flight request is cut off with the drain wait
removed.

340 tests on SQLite, 341 on Postgres, 70 Playwright, rules coverage 100%.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Michilis
2026-09-06 22:03:13 +00:00
co-authored by Claude Opus 5
parent 8c8463924f
commit 37370dd079
29 changed files with 1349 additions and 43 deletions
+131
View File
@@ -0,0 +1,131 @@
name: CI
on:
push:
branches: [main, master]
pull_request:
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: true
jobs:
# Everything that does not need a database, once.
static:
name: types, lint and rules coverage
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: pnpm/action-setup@v4
- uses: actions/setup-node@v4
with:
node-version: 22
cache: pnpm
- run: pnpm install --frozen-lockfile
- run: pnpm typecheck
- run: pnpm lint
# RULES.md section 10: the tax logic is the product, so its coverage is a gate.
- run: pnpm test:coverage
# The same suite against both dialects. SPEC.md section 13: the product claims to run on
# SQLite and on Postgres, so both are run rather than one being assumed from the other.
test:
name: tests (${{ matrix.dialect }})
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
dialect: [sqlite, postgres]
services:
postgres:
image: postgres:16-alpine
env:
POSTGRES_USER: impuestos
POSTGRES_PASSWORD: impuestos
POSTGRES_DB: impuestos
ports: ['5432:5432']
options: >-
--health-cmd "pg_isready -U impuestos"
--health-interval 5s
--health-timeout 5s
--health-retries 10
# Scaled mode requires S3, so the driver is exercised against a real object store
# rather than assumed from the local one.
minio:
image: bitnami/minio:latest
env:
MINIO_ROOT_USER: impuestos
MINIO_ROOT_PASSWORD: impuestos-secret
MINIO_DEFAULT_BUCKETS: comprobantes
ports: ['9000:9000']
options: >-
--health-cmd "curl -f http://localhost:9000/minio/health/live"
--health-interval 5s
--health-timeout 5s
--health-retries 20
steps:
- uses: actions/checkout@v4
- uses: pnpm/action-setup@v4
- uses: actions/setup-node@v4
with:
node-version: 22
cache: pnpm
- run: pnpm install --frozen-lockfile
# TEST_DATABASE_URL unset means SQLite in memory, which is what the harness falls
# back to. Set, every harness builds itself a database of its own inside this server,
# and the two-worker claim test in scale.test.ts stops skipping.
- name: Run the suite
env:
TEST_DATABASE_URL: ${{ matrix.dialect == 'postgres' && 'postgres://impuestos:impuestos@localhost:5432/impuestos' || '' }}
TEST_S3_ENDPOINT: http://localhost:9000
TEST_S3_BUCKET: comprobantes
TEST_S3_REGION: us-east-1
TEST_S3_ACCESS_KEY_ID: impuestos
TEST_S3_SECRET_ACCESS_KEY: impuestos-secret
run: pnpm test
# The golden paths, against a built stack rather than the dev server. SQLite and local
# files: that is the combined mode compose ships, and it is the one the helpers in
# e2e/db.ts can read, since they open the database file to check what a flow wrote.
e2e:
name: playwright
runs-on: ubuntu-latest
env:
DATABASE_URL: sqlite:./apps/api/data/app.db
STORAGE_DRIVER: local
STORAGE_LOCAL_PATH: ./apps/api/data/files
BETTER_AUTH_SECRET: ci-secret-that-is-long-enough-32-chars
BETTER_AUTH_URL: http://localhost:3005
APP_PUBLIC_URL: http://localhost:3005
PORT: '4000'
API_INTERNAL_URL: http://localhost:4000
E2E_BASE_URL: http://localhost:3005
steps:
- uses: actions/checkout@v4
- uses: pnpm/action-setup@v4
- uses: actions/setup-node@v4
with:
node-version: 22
cache: pnpm
- run: pnpm install --frozen-lockfile
- run: pnpm exec playwright install --with-deps chromium
- run: pnpm db:migrate
- run: pnpm db:seed
- run: pnpm build
- name: Start the stack
run: |
node --import tsx apps/api/src/index.ts &
pnpm --filter @impuestos/web start &
npx --yes wait-on http://localhost:4000/readyz http://localhost:3005/healthz -t 120000
- run: pnpm test:e2e
- uses: actions/upload-artifact@v4
if: failure()
with:
name: playwright-report
path: playwright-report/
retention-days: 7
+75
View File
@@ -752,3 +752,78 @@ either has its subject or is a 404, and there is no third case. The crossfade fr
to content is applied where the skeleton is a separate early return; the three list screens
render their skeleton inside the same tree as their heading, where fading the whole tree
would also fade a heading that never changed.
## Phase 8
### The dual-dialect claim is now run, not asserted
Six phases said the product runs on SQLite and on Postgres. Nothing had ever run it on
Postgres. `TEST_DATABASE_URL` now points the whole suite at a real server, and CI runs both
arms of the matrix.
The first run found three things, one of them a bug that would have shipped:
**A frozen account went on being notified on Postgres.** `banned` is integer 0/1 on SQLite
and a real boolean on Postgres, and the code read it as `recipient.banned === 1`. On
Postgres the driver hands back `true`, the comparison is false, and an account someone
asked us to freeze keeps getting email. The two better-auth flags are now typed for both
dialects and read through `isFlagSet`, and a test asserts the behaviour on whichever
dialect the suite is pointed at.
**The migration advisory lock was taken on a pool.** `pg_advisory_lock` belongs to the
session that took it. Issued on one pooled connection and released on another, the unlock
is a no-op with a warning and the lock stays held until that connection happens to close.
Two replicas migrating at once, which is the ordinary case in k8s, is exactly what it was
supposed to protect. It now runs on one pinned connection.
**Writes survived by accident.** `set banned = 1` and `where banned = 0` work on Postgres
only because its boolean input syntax accepts the strings '1' and '0'. That is luck rather
than design, and it is why the read path was the one that broke.
### Each Postgres harness gets a database, not a schema
A schema per harness was the obvious choice and it does not work. Kysely's migrator asks
the introspector whether its bookkeeping tables exist, and with no schema configured a
table of that name in any schema counts, so twenty parallel schemas each holding a
`kysely_migration` convince each other the work is already done. A database each has no
such ambiguity. It is slower, which is what the longer timeout under `TEST_DATABASE_URL`
is for: building a real database is honest work, not a hang.
### DATABASE_POOL_MAX
Connections are the resource that runs out first when replicas multiply. It was hardcoded
at ten. It is now a variable, documented next to the arithmetic that matters: pool size
times (api replicas plus worker replicas) has to stay under the server's max_connections.
### Both new tests were checked against a broken implementation
A green test that cannot fail proves nothing, so both were run against a deliberately
broken version first:
- Remove `FOR UPDATE SKIP LOCKED` from the Postgres claim and the two-worker test reports
188 claims for 100 jobs.
- Remove the drain wait from the shutdown path and the in-flight request is cut off.
The shutdown test spawns a real process and holds a request open by sending its body in two
pieces, so the server is genuinely inside the handler when SIGTERM arrives. Asserting on
the app object would have proved nothing about a socket.
### The manifests are checked for the thing that actually rots
They cannot be applied to a cluster from here, so `deploy.test.ts` checks the part that
drifts: every variable set in a ConfigMap, Secret or `env:` block has to be one the env
schema knows, the grace period has to outlast the 25 second drain, and the API has to be
the one that does not poll. A variable renamed in `src/lib/env.ts` and forgotten in a
ConfigMap is a pod that starts and then refuses to boot, and no generic schema validator
would catch it.
### The e2e job stays on SQLite
SPEC.md section 13 puts the golden paths against the compose stack, which is combined mode:
SQLite and local files. It is also the only shape `e2e/db.ts` can read, since those helpers
open the database file to check what a flow actually wrote. S3 is covered where it belongs,
in the unit suite against MinIO, and Postgres by the dialect matrix.
### The storage drivers had no tests at all
S3 is required in scaled mode and had never been exercised. Both drivers are now held to
the same round trip, with the S3 half gated on `TEST_S3_ENDPOINT`. That half has not run on
this machine: there is no object store here and no Docker to start one. CI runs it.
### Still not run here
The images, the compose stack and the manifests have never been built or applied on this
machine. No Docker daemon it can reach, no cluster. The README says so where someone about
to deploy will read it, rather than only here.
+52 -6
View File
@@ -7,12 +7,13 @@ and ready to file yourself.
Working name. See `docs/` for the specifications, `DECISIONS.md` for choices made along the
way and the gaps that still need answers.
> **Status: phase 7 of 8.** The whole taxpayer path works: scan a comprobante, confirm it,
> watch the position move, and take the resulting Formulario 120 or 515 from review to
> **Status: phase 8 of 8, feature complete.** The whole taxpayer path works: scan a
> comprobante, confirm it, watch the position move, and take the resulting Formulario 120 or 515 from review to
> approved to a PDF you file yourself in Marangatu. Staff have a console: look an account
> up, work the ingestion error queue, read and export the audit log. It installs to a home
> screen, keeps a capture taken with no signal and sends it later, and can send a push when
> a deadline is close. The scale-out work is the last phase.
> a deadline is close. It runs on SQLite on one box or on Postgres and S3 across as many
> replicas as you like, and the whole suite is run against both.
---
@@ -131,6 +132,40 @@ S3. The k8s manifests assume both.
Rate limiting is in memory and therefore per replica. At this scale that is deliberate;
the limiter is behind an interface for when it is not.
### Running it on Kubernetes
`deploy/k8s/` holds the manifests, and its README explains what each one is for. The order
is namespace, config, secret, api, worker, web, ingress, hpa. Three things about the shape
are worth knowing before you read them:
**Migrations run as an initContainer on every API pod.** Two pods starting together is the
normal case, not the exception, and it is safe: the migration takes a Postgres advisory
lock on one pinned connection, so the second waits for the first and then finds nothing to
do.
**Exactly one poller per job, not one per replica.** The API Deployment sets
`JOBS_INLINE=false` and the polling lives in its own worker Deployment. The worker is safe
at any replica count on Postgres: claiming uses `FOR UPDATE SKIP LOCKED`, which
`apps/api/src/modules/jobs/scale.test.ts` holds to a hundred jobs and two workers with no
job claimed twice.
**Only the web app is exposed.** The Ingress routes to the web Service and the API has no
route in from outside. The browser talks to one origin and `/api` is forwarded inside the
cluster, which is why there is no CORS configuration anywhere in this repository.
Sizing: `DATABASE_POOL_MAX` is per pod. Multiply it by (api replicas + worker replicas) and
keep the total under the Postgres `max_connections`, or the tenth pod to start is the one
that cannot connect.
### What has not been run here
The container images, the compose stack and the Kubernetes manifests have never been built
or applied on the machine this was written on: there is no Docker daemon it can reach and
no cluster. The manifests parse, their configuration is checked against the env schema by
`apps/api/src/deploy.test.ts`, and the Dockerfiles are ordinary multi stage Node builds,
but none of that is the same as having watched a pod come up. Treat the first deploy as the
first real test of them.
## Installing it, offline and push
`app/manifest.ts` and the icons in `apps/web/public` make the web app installable; the
@@ -196,9 +231,20 @@ old definition in place: it is what old declarations still render through.
## Testing
Vitest covers the rules package, the API services against SQLite in memory, and catalog
parity. Playwright covers the golden paths against a running stack. CI runs the suite
against both SQLite and Postgres.
Vitest covers the rules package, the API services, and catalog parity. Playwright covers
the golden paths against a running stack.
The same suite runs against both dialects. With no `TEST_DATABASE_URL` it uses SQLite in
memory; point that at a Postgres and every harness builds a database of its own inside it,
which is how CI covers the claim that the product runs on either:
```bash
TEST_DATABASE_URL=postgres://user:pass@localhost:5432/postgres pnpm test
```
Two tests only mean something on Postgres and skip themselves without it: the two-worker
hundred-job claim test, and anything that depends on `FOR UPDATE SKIP LOCKED`. The S3
driver behaves the same way through `TEST_S3_ENDPOINT`.
`pnpm test:e2e` does not start anything: bring the stack up first, then point it at the
right origin. The suite signs in as the demo accounts and confirms, rejects and scans
+4
View File
@@ -45,6 +45,10 @@ S3_FORCE_PATH_STYLE=true
ANTHROPIC_API_KEY=
OCR_MODEL=claude-sonnet-4-6
# Postgres connections per process. Multiply by the number of replicas and keep the total
# under the server's max_connections. Ignored on SQLite.
DATABASE_POOL_MAX=10
# Optional. Web push is hidden in the UI when unset. Generate: npx web-push generate-vapid-keys
PUSH_VAPID_PUBLIC_KEY=
PUSH_VAPID_PRIVATE_KEY=
+3 -2
View File
@@ -20,9 +20,10 @@ export function dialectOf(databaseUrl: string): Dialect {
);
}
export function createDb(databaseUrl: string): DbHandle {
export function createDb(databaseUrl: string, options: { poolMax?: number } = {}): DbHandle {
const dialect = dialectOf(databaseUrl);
const handle = dialect === 'sqlite' ? createSqliteDb(databaseUrl) : createPostgresDb(databaseUrl);
const handle =
dialect === 'sqlite' ? createSqliteDb(databaseUrl) : createPostgresDb(databaseUrl, options);
return { ...handle, dialect };
}
+24 -8
View File
@@ -8,11 +8,14 @@ import type { Database } from './schema';
* Timestamps are stored as ISO-8601 text in our own tables, so the driver is told to
* hand back `numeric` as a number and nothing else needs a type parser.
*/
export function createPostgresDb(databaseUrl: string): {
export function createPostgresDb(
databaseUrl: string,
options: { poolMax?: number } = {},
): {
db: Kysely<Database>;
close: () => Promise<void>;
} {
const pool = new pg.Pool({ connectionString: databaseUrl, max: 10 });
const pool = new pg.Pool({ connectionString: databaseUrl, max: options.poolMax ?? 10 });
const db = new Kysely<Database>({ dialect: new PostgresDialect({ pool }) });
return {
db,
@@ -45,10 +48,23 @@ pg.types.setTypeParser(pg.types.builtins.INT8, (value) => {
* api pod, so two migrators can start at the same moment.
*/
export async function withMigrationLock<T>(db: Kysely<Database>, fn: () => Promise<T>): Promise<T> {
await sql`select pg_advisory_lock(${sql.lit(MIGRATION_ADVISORY_LOCK_KEY)})`.execute(db);
try {
return await fn();
} finally {
await sql`select pg_advisory_unlock(${sql.lit(MIGRATION_ADVISORY_LOCK_KEY)})`.execute(db);
}
/**
* On one pinned connection, not on the pool. An advisory lock belongs to the session that
* took it: issue the lock on one pooled connection and the unlock on another and the
* unlock is a no-op with a warning, leaving the lock held until that connection happens
* to close. The next replica to start then waits on a lock nobody holds any more.
*
* `fn` still uses the pool for its own work, which is what makes this mutual exclusion
* rather than a single threaded migration.
*/
return db.connection().execute(async (connection) => {
await sql`select pg_advisory_lock(${sql.lit(MIGRATION_ADVISORY_LOCK_KEY)})`.execute(connection);
try {
return await fn();
} finally {
await sql`select pg_advisory_unlock(${sql.lit(MIGRATION_ADVISORY_LOCK_KEY)})`.execute(
connection,
);
}
});
}
+17 -2
View File
@@ -37,20 +37,35 @@ export interface Database {
insight_dismissals: InsightDismissalsTable;
}
/**
* better-auth owns this table and emits its own DDL per dialect, which is why the two
* flags are typed for both: on SQLite they come back as integer 0/1 and on Postgres as a
* real boolean. Read them through `isFlagSet` rather than comparing to 1.
*/
export interface UserTable {
id: string;
name: string;
email: string;
emailVerified: number;
emailVerified: number | boolean;
image: string | null;
createdAt: string;
updatedAt: string;
role: string | null;
banned: number | null;
banned: number | boolean | null;
banReason: string | null;
banExpires: string | null;
}
/**
* True for a flag on a better-auth table, whichever dialect wrote it.
*
* `banned === 1` was silently false on Postgres, where the driver hands back `true`, and a
* frozen account went on receiving notifications.
*/
export function isFlagSet(value: number | boolean | null | undefined): boolean {
return value === true || value === 1;
}
export interface SessionTable {
id: string;
expiresAt: string;
+111
View File
@@ -0,0 +1,111 @@
import { readFileSync, readdirSync } from 'node:fs';
import { fileURLToPath } from 'node:url';
import { parseAllDocuments } from 'yaml';
import { describe, expect, it } from 'vitest';
import { ENV_KEYS } from './lib/env';
/**
* The manifests in deploy/k8s cannot be applied to a cluster from here, so what is checked
* is the part that actually rots: the configuration contract between them and the env
* schema. A variable renamed in src/lib/env.ts and forgotten in a ConfigMap is a pod that
* starts and then refuses to boot, and it is exactly the sort of thing nobody notices until
* a deploy.
*/
const K8S = fileURLToPath(new URL('../../../deploy/k8s/', import.meta.url));
interface Manifest {
kind?: string;
metadata?: { name?: string };
data?: Record<string, string>;
stringData?: Record<string, string>;
spec?: {
template?: {
spec?: {
terminationGracePeriodSeconds?: number;
containers?: { name?: string; env?: { name?: string; value?: string }[] }[];
initContainers?: { name?: string }[];
};
};
};
}
function load(): { file: string; doc: Manifest }[] {
return readdirSync(K8S)
.filter((file) => file.endsWith('.yaml'))
.flatMap((file) =>
parseAllDocuments(readFileSync(K8S + file, 'utf8'))
.map((document) => document.toJS() as Manifest)
.filter((doc): doc is Manifest => Boolean(doc))
.map((doc) => ({ file, doc })),
);
}
const manifests = load();
/** Name alone is ambiguous: a Deployment, its Service and its HPA all share one. */
const byName = (name: string, kind = 'Deployment') =>
manifests.find((m) => m.doc.metadata?.name === name && m.doc.kind === kind)?.doc;
describe('the kubernetes manifests', () => {
it('parse, and every one names a kind', () => {
expect(manifests.length).toBeGreaterThan(0);
for (const { file, doc } of manifests) {
expect(doc.kind, file).toBeTruthy();
}
});
it('set only variables the API actually reads', () => {
const known = new Set(ENV_KEYS);
const configured = [
...Object.keys(byName('impuestos-config', 'ConfigMap')?.data ?? {}),
...Object.keys(byName('impuestos-secrets', 'Secret')?.stringData ?? {}),
...(byName('impuestos-api')?.spec?.template?.spec?.containers ?? [])
.flatMap((container) => container.env ?? [])
.map((entry) => entry.name ?? ''),
...(byName('impuestos-worker')?.spec?.template?.spec?.containers ?? [])
.flatMap((container) => container.env ?? [])
.map((entry) => entry.name ?? ''),
];
for (const key of configured) {
expect(known.has(key), `${key} is set in deploy/k8s but no such variable exists`).toBe(true);
}
});
it('carries everything the scaled mode needs', () => {
const config = byName('impuestos-config', 'ConfigMap')?.data ?? {};
const secrets = byName('impuestos-secrets', 'Secret')?.stringData ?? {};
// SPEC.md section 15: scaled mode is Postgres and S3, not a choice.
expect(secrets['DATABASE_URL']).toMatch(/^postgres/);
expect(config['STORAGE_DRIVER']).toBe('s3');
for (const key of ['S3_BUCKET', 'S3_REGION']) expect(config[key]).toBeTruthy();
for (const key of ['BETTER_AUTH_SECRET', 'S3_ACCESS_KEY_ID', 'S3_SECRET_ACCESS_KEY']) {
expect(secrets).toHaveProperty(key);
}
});
it('runs exactly one poller per job, not one per replica', () => {
const api = byName('impuestos-api')?.spec?.template?.spec?.containers?.[0]?.env ?? [];
const worker = byName('impuestos-worker')?.spec?.template?.spec?.containers?.[0]?.env ?? [];
// An API replica that polled would multiply pollers by the replica count.
expect(api.find((entry) => entry.name === 'ROLE')?.value).toBe('server');
expect(api.find((entry) => entry.name === 'JOBS_INLINE')?.value).toBe('false');
expect(worker.find((entry) => entry.name === 'ROLE')?.value).toBe('worker');
});
it('migrates before it serves', () => {
const init = byName('impuestos-api')?.spec?.template?.spec?.initContainers ?? [];
expect(init.map((container) => container.name)).toContain('migrate');
});
it('gives every pod longer to stop than the API spends draining', () => {
// src/index.ts drains for up to 25 seconds. A shorter grace period would have the
// cluster kill the process in the middle of the requests it is trying to finish.
const drainSeconds = 25;
for (const name of ['impuestos-api', 'impuestos-worker', 'impuestos-web']) {
const grace = byName(name)?.spec?.template?.spec?.terminationGracePeriodSeconds;
expect(grace, name).toBeGreaterThan(drainSeconds);
}
});
});
+3 -1
View File
@@ -1,5 +1,6 @@
import { DataExportDto, DependentDto, ProfileDto } from '@impuestos/contracts';
import { afterAll, beforeAll, beforeEach, describe, expect, it } from 'vitest';
import { isFlagSet } from '../../db/schema';
import { createHarness, type Harness } from '../../test/harness';
let h: Harness;
@@ -228,7 +229,8 @@ describe('DELETE /me/account', () => {
.select(['id', 'banned', 'banReason'])
.where('email', '=', 'carlos@demo.local')
.executeTakeFirstOrThrow();
expect(user.banned).toBe(1);
// Integer 0/1 on SQLite, a real boolean on Postgres: the flag is read, not compared.
expect(isFlagSet(user.banned)).toBe(true);
expect(user.banReason).toBe('account_deleted');
const sessions = await h.deps.handle.db
+1 -1
View File
@@ -14,7 +14,7 @@ import { createStorage } from './modules/storage';
const DRAIN_TIMEOUT_MS = 25_000;
const env = loadEnv();
const handle = createDb(env.DATABASE_URL);
const handle = createDb(env.DATABASE_URL, { poolMax: env.DATABASE_POOL_MAX });
const auth = createAuth({
db: handle.db,
dialect: handle.dialect,
+6
View File
@@ -42,6 +42,12 @@ const EnvObject = z.object({
ANTHROPIC_API_KEY: optionalString,
OCR_MODEL: z.string().trim().default('claude-sonnet-4-6'),
/**
* Postgres connections per process. Replicas multiply it, so it has to be sized
* against the server's max_connections rather than left to a library default.
*/
DATABASE_POOL_MAX: z.coerce.number().int().min(1).max(100).default(10),
PUSH_VAPID_PUBLIC_KEY: optionalString,
PUSH_VAPID_PRIVATE_KEY: optionalString,
/** Contact for the push service. Must be an https: or mailto: URL, per RFC 8292. */
+88
View File
@@ -0,0 +1,88 @@
import pg from 'pg';
import { describe, expect, it } from 'vitest';
import { createDb } from '../../db/index';
import { migrateToLatest } from '../../db/migrator';
import { parseEnv } from '../../lib/env';
import { TEST_ENV } from '../../test/harness';
import { claimJob } from './claim';
import { enqueue } from './queue';
/**
* SPEC.md section 13: with Postgres and two workers, enqueue 100 jobs and assert each one
* is claimed exactly once.
*
* This is the whole basis of running more than one replica, so it runs against a real
* Postgres or not at all. `FOR UPDATE SKIP LOCKED` cannot be simulated: SQLite's single
* writer makes any test of it pass for the wrong reason.
*/
const POSTGRES_URL = process.env['TEST_DATABASE_URL'];
const JOBS = 100;
const WORKERS = 2;
describe.skipIf(!POSTGRES_URL)('two workers on one Postgres queue', () => {
it('claims each of a hundred jobs exactly once', async () => {
const database = `scale_${Date.now()}`;
await withAdmin(POSTGRES_URL as string, (db) => db.query(`create database "${database}"`));
const url = new URL(POSTGRES_URL as string);
url.pathname = `/${database}`;
const parsed = parseEnv({ ...TEST_ENV, DATABASE_URL: url.toString(), DATABASE_POOL_MAX: '5' });
const env = parsed.env;
if (!parsed.ok || !env) throw new Error(parsed.message);
// Two handles, because two worker processes would have two pools.
const workers = Array.from({ length: WORKERS }, () =>
createDb(env.DATABASE_URL, { poolMax: 5 }),
);
try {
await migrateToLatest(workers[0] as (typeof workers)[number], env);
for (let index = 0; index < JOBS; index += 1) {
await enqueue(workers[0]!.db, { type: 'classify_document', payload: { index } });
}
const now = new Date().toISOString();
const claimed: string[] = [];
/** One worker draining until the queue gives it nothing. */
async function drain(handle: (typeof workers)[number], instanceId: string): Promise<void> {
for (;;) {
const job = await claimJob(handle.db, handle.dialect, { instanceId, now });
if (!job) return;
claimed.push(job.id);
}
}
// Both at once, which is the only arrangement that can double claim.
await Promise.all(workers.map((handle, index) => drain(handle, `worker-${index}`)));
expect(claimed).toHaveLength(JOBS);
expect(new Set(claimed).size).toBe(JOBS);
// And the queue agrees: everything is running under someone, nothing is pending.
const rows = await workers[0]!.db.selectFrom('jobs').select(['status', 'locked_by']).execute();
expect(rows).toHaveLength(JOBS);
expect(rows.every((row) => row.status === 'running')).toBe(true);
expect(new Set(rows.map((row) => row.locked_by)).size).toBe(WORKERS);
} finally {
await Promise.all(workers.map((handle) => handle.close()));
await withAdmin(POSTGRES_URL as string, (db) =>
db.query(`drop database if exists "${database}" with (force)`),
);
}
});
});
async function withAdmin(
databaseUrl: string,
run: (client: pg.Client) => Promise<unknown>,
): Promise<void> {
const client = new pg.Client({ connectionString: databaseUrl });
await client.connect();
try {
await run(client);
} finally {
await client.end();
}
}
+44 -1
View File
@@ -1,7 +1,8 @@
import { computeRucDv, dueDateFor, formatIsoDate } from '@impuestos/rules';
import { afterEach, describe, expect, it } from 'vitest';
import { isFlagSet } from '../../db/schema';
import { createHarness, type Harness } from '../../test/harness';
import type { Channels, OutgoingMessage } from '../notifications';
import { sendNotification, type Channels, type OutgoingMessage } from '../notifications';
import { runAutoConfirmSweep, runDeadlineSweep, runDigestSweep } from './sweep-handlers';
import { scheduleSweeps, sweepDedupeKey } from './sweeps';
@@ -140,6 +141,48 @@ describe('deadline sweep, ten days out', () => {
});
});
/**
* The flag better-auth writes is integer 0/1 on SQLite and a boolean on Postgres. Reading
* it with `=== 1` was silently false on Postgres, so a frozen account went on being
* notified. This runs on whichever dialect the suite is pointed at.
*/
describe('a frozen account', () => {
it('is not a recipient on either dialect', async () => {
const h = (harness = await createHarness({ channels: recordingChannels() }));
const db = h.deps.handle.db;
const carlos = await db
.selectFrom('user')
.select('id')
.where('email', '=', 'carlos@demo.local')
.executeTakeFirstOrThrow();
await db
.updateTable('notification_prefs')
.set({ email_enabled: 1 })
.where('user_id', '=', carlos.id)
.execute();
await db.updateTable('user').set({ banned: 1 }).where('id', '=', carlos.id).execute();
const stored = await db
.selectFrom('user')
.select('banned')
.where('id', '=', carlos.id)
.executeTakeFirstOrThrow();
// Whatever the driver hands back, the helper agrees it is set.
expect(isFlagSet(stored.banned)).toBe(true);
const result = await sendNotification(
{ db, channels: h.channels },
{
userId: carlos.id,
notification: { kind: 'declaration_ready', form: '120', period: '2026-08', declarationId: 'x' },
},
);
expect(result.delivered).toEqual([]);
});
});
describe('the queued notification actually goes out', () => {
it('reaches the channels the user has enabled, in their language', async () => {
const channels = recordingChannels();
+2 -2
View File
@@ -1,7 +1,7 @@
import { DEFAULT_LOCALE, isLocale, type Locale } from '@impuestos/i18n';
import type { Kysely } from 'kysely';
import { z } from 'zod';
import type { Database } from '../../db/schema';
import { isFlagSet, type Database } from '../../db/schema';
import type { Channels } from './channels';
import { emailSubject, renderNotification, type NotificationKind } from './templates';
@@ -38,7 +38,7 @@ export async function sendNotification(
.executeTakeFirst();
// A frozen account is not a recipient.
if (!recipient || recipient.banned === 1) return { delivered: [] };
if (!recipient || isFlagSet(recipient.banned)) return { delivered: [] };
const locale: Locale = isLocale(recipient.locale) ? recipient.locale : DEFAULT_LOCALE;
const message = renderNotification(locale, args.notification);
@@ -0,0 +1,138 @@
import { mkdtempSync, readdirSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { join } from 'node:path';
import { describe, expect, it } from 'vitest';
import { parseEnv } from '../../lib/env';
import { createStorage, storageKey, type StorageDriver } from './index';
/**
* The storage driver is the one thing scaled mode swaps out (SPEC.md section 15: S3 is
* required there), so both implementations are held to the same round trip.
*
* The S3 half needs a real endpoint and is skipped without one. CI points it at MinIO;
* locally, set TEST_S3_ENDPOINT and the four S3 variables to run it.
*/
const S3_ENDPOINT = process.env['TEST_S3_ENDPOINT'];
const BASE = {
NODE_ENV: 'test',
DATABASE_URL: 'sqlite::memory:',
BETTER_AUTH_SECRET: 'storage-test-secret-long-enough-32ch',
BETTER_AUTH_URL: 'http://localhost:3000',
APP_PUBLIC_URL: 'http://localhost:3000',
};
function driverFor(overrides: Record<string, string>): StorageDriver {
const parsed = parseEnv({ ...BASE, ...overrides });
if (!parsed.ok || !parsed.env) throw new Error(parsed.message);
return createStorage(parsed.env);
}
async function readAll(stream: ReadableStream<Uint8Array>): Promise<Uint8Array> {
const chunks: Uint8Array[] = [];
const reader = stream.getReader();
for (;;) {
const { done, value } = await reader.read();
if (done) break;
if (value) chunks.push(value);
}
return new Uint8Array(chunks.flatMap((chunk) => [...chunk]));
}
describe('storage keys', () => {
it('never carry a name the uploader chose', () => {
// A key is the owner and an id we generated, so "../../etc/passwd" as a filename has
// nowhere to go: it is not part of the key at all.
expect(storageKey('user-1', 'abc')).toBe('user-1/abc');
expect(storageKey('user-1', 'abc')).not.toContain('..');
});
it('keep one account out of the prefix belonging to another', () => {
expect(storageKey('user-1', 'abc').startsWith('user-2/')).toBe(false);
});
});
describe('the driver switch', () => {
it('follows STORAGE_DRIVER and nothing else', () => {
expect(driverFor({ STORAGE_LOCAL_PATH: mkdtempSync(join(tmpdir(), 'st-')) }).kind).toBe('local');
expect(
driverFor({
STORAGE_DRIVER: 's3',
S3_BUCKET: 'whatever',
S3_REGION: 'us-east-1',
S3_ACCESS_KEY_ID: 'a',
S3_SECRET_ACCESS_KEY: 'b',
}).kind,
).toBe('s3');
});
it('refuses the s3 driver with no bucket rather than failing on first upload', () => {
expect(() => driverFor({ STORAGE_DRIVER: 's3' })).toThrow(/S3_BUCKET/);
});
});
describe('the local driver', () => {
const path = mkdtempSync(join(tmpdir(), 'impuestos-storage-'));
const storage = driverFor({ STORAGE_LOCAL_PATH: path });
it('stores, reads back and deletes', async () => {
const key = storageKey('user-1', 'factura');
const body = new TextEncoder().encode('una factura');
await storage.put(key, body, 'text/plain');
expect(new TextDecoder().decode(await readAll(await storage.getStream(key)))).toBe('una factura');
// Local files are served through the authenticated route, never linked directly, and
// the key is escaped into the path rather than pasted into it.
expect(await storage.url(key)).toBe(`/api/files/${encodeURIComponent(key)}`);
await storage.delete(key);
await expect(storage.getStream(key)).rejects.toThrow();
});
it('answers the readiness check once its directory exists', async () => {
await expect(storage.check()).resolves.toBeUndefined();
// Writing created the per-user prefix; nothing else is in there.
expect(readdirSync(path).length).toBeGreaterThanOrEqual(0);
});
});
/** Built inside the tests: a skipped describe still runs its own body. */
function s3Driver(overrides: Record<string, string> = {}): StorageDriver {
return driverFor({
STORAGE_DRIVER: 's3',
S3_ENDPOINT: S3_ENDPOINT ?? '',
S3_REGION: process.env['TEST_S3_REGION'] ?? 'us-east-1',
S3_BUCKET: process.env['TEST_S3_BUCKET'] ?? 'comprobantes',
S3_ACCESS_KEY_ID: process.env['TEST_S3_ACCESS_KEY_ID'] ?? '',
S3_SECRET_ACCESS_KEY: process.env['TEST_S3_SECRET_ACCESS_KEY'] ?? '',
S3_FORCE_PATH_STYLE: 'true',
...overrides,
});
}
describe.skipIf(!S3_ENDPOINT)('the s3 driver', () => {
it('stores, reads back, signs and deletes', async () => {
const storage = s3Driver();
const key = storageKey('user-1', `factura-${Date.now()}`);
const body = new TextEncoder().encode('una factura');
await storage.put(key, body, 'text/plain');
expect(new TextDecoder().decode(await readAll(await storage.getStream(key)))).toBe('una factura');
// A signed URL, which is how a browser gets the file in scaled mode.
const url = await storage.url(key);
expect(url).toContain(key);
expect(url).toMatch(/X-Amz-Signature/);
expect(new TextDecoder().decode(new Uint8Array(await (await fetch(url)).arrayBuffer()))).toBe(
'una factura',
);
await storage.delete(key);
await expect(storage.getStream(key)).rejects.toThrow();
});
it('fails the readiness check when the bucket is not there', async () => {
await expect(s3Driver({ S3_BUCKET: 'no-such-bucket-here' }).check()).rejects.toThrow();
});
});
+148
View File
@@ -0,0 +1,148 @@
import { spawn, type ChildProcess } from 'node:child_process';
import { request as httpRequest } from 'node:http';
import { mkdtempSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { join } from 'node:path';
import { fileURLToPath } from 'node:url';
import { afterAll, describe, expect, it } from 'vitest';
import { createAuth } from './auth/options';
import { createDb } from './db/index';
import { migrateToLatest } from './db/migrator';
import { seed } from './db/seed';
import { parseEnv } from './lib/env';
/**
* SPEC.md section 15: on SIGTERM the API stops accepting, finishes what it is already
* handling, and exits 0. A rolling deploy that cuts a request in half loses a scan someone
* just took, so this runs against a real process and a real socket rather than the app
* object.
*/
const ENTRY = fileURLToPath(new URL('./index.ts', import.meta.url));
const PORT = 4571;
const DIR = mkdtempSync(join(tmpdir(), 'impuestos-shutdown-'));
let child: ChildProcess | null = null;
afterAll(() => {
child?.kill('SIGKILL');
});
describe('graceful shutdown', () => {
it('finishes an in-flight request, then exits 0', async () => {
const env = {
NODE_ENV: 'test',
PORT: String(PORT),
DATABASE_URL: `sqlite:${join(DIR, 'app.db')}`,
STORAGE_LOCAL_PATH: join(DIR, 'files'),
BETTER_AUTH_SECRET: 'shutdown-test-secret-long-enough-32ch',
BETTER_AUTH_URL: `http://127.0.0.1:${PORT}`,
APP_PUBLIC_URL: `http://127.0.0.1:${PORT}`,
};
await prepareDatabase(env);
child = await start(env);
const cookie = await signIn();
/**
* A request whose body arrives in two pieces. The server has the headers and is inside
* the handler waiting for the rest, which is exactly the state a deploy must not cut.
*/
const body = JSON.stringify({ pushEnabled: false });
const [head, tail] = [body.slice(0, 4), body.slice(4)];
const response = new Promise<{ status: number; text: string }>((resolve, reject) => {
const outgoing = httpRequest(
{
host: '127.0.0.1',
port: PORT,
path: '/api/me/notification-prefs',
method: 'PATCH',
headers: {
cookie,
'content-type': 'application/json',
'content-length': Buffer.byteLength(body),
},
},
(incoming) => {
let text = '';
incoming.on('data', (chunk) => (text += chunk));
incoming.on('end', () => resolve({ status: incoming.statusCode ?? 0, text }));
},
);
outgoing.on('error', reject);
outgoing.write(head);
// Held open on purpose.
setTimeout(() => outgoing.end(tail), 1200);
});
// Long enough for the server to be inside the handler, well short of the body.
await new Promise((resolve) => setTimeout(resolve, 400));
const exited = new Promise<number>((resolve) => {
child?.on('exit', (code) => resolve(code ?? -1));
});
child?.kill('SIGTERM');
// Draining, so a load balancer stops sending here before the door closes.
await new Promise((resolve) => setTimeout(resolve, 100));
const health = await fetch(`http://127.0.0.1:${PORT}/healthz`).catch(() => null);
if (health) expect(health.status).toBe(503);
const result = await response;
expect(result.status).toBe(200);
expect(await exited).toBe(0);
// Spawning a process, migrating a file database and holding a request open is slower
// than a unit test has any business being.
}, 60_000);
});
async function prepareDatabase(env: Record<string, string>): Promise<void> {
const parsed = parseEnv(env);
if (!parsed.ok || !parsed.env) throw new Error(parsed.message);
const handle = createDb(parsed.env.DATABASE_URL);
try {
await migrateToLatest(handle, parsed.env);
const auth = createAuth({
db: handle.db,
dialect: handle.dialect,
env: parsed.env,
sendOtp: async () => undefined,
});
await seed(handle, auth);
} finally {
await handle.close();
}
}
async function start(env: Record<string, string>): Promise<ChildProcess> {
const proc = spawn(process.execPath, ['--import', 'tsx', ENTRY], {
// From apps/api, where tsx is a dependency and where `pnpm dev` runs it from.
cwd: fileURLToPath(new URL('..', import.meta.url)),
env: { ...process.env, ...env },
stdio: ['ignore', 'pipe', 'pipe'],
});
proc.stderr?.on('data', (chunk) => console.error(`[api] ${String(chunk).trim()}`));
const deadline = Date.now() + 30_000;
for (;;) {
if (Date.now() > deadline) throw new Error('the api did not come up');
const ok = await fetch(`http://127.0.0.1:${PORT}/healthz`)
.then((response) => response.ok)
.catch(() => false);
if (ok) return proc;
await new Promise((resolve) => setTimeout(resolve, 200));
}
}
async function signIn(): Promise<string> {
const response = await fetch(`http://127.0.0.1:${PORT}/api/auth/sign-in/email`, {
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify({ email: 'maria@demo.local', password: 'demo-maria-1' }),
});
if (!response.ok) throw new Error(`sign in failed: ${response.status}`);
const cookie = response.headers.get('set-cookie');
if (!cookie) throw new Error('sign in returned no cookie');
return cookie.split(';')[0] ?? '';
}
+68 -3
View File
@@ -1,3 +1,5 @@
import pg from 'pg';
import { uuidv7 } from 'uuidv7';
import { mkdtempSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { join } from 'node:path';
@@ -13,9 +15,16 @@ import type { OcrProvider } from '../modules/documents/ocr';
import { createChannels, type Channels } from '../modules/notifications';
import { createStorage, type StorageDriver } from '../modules/storage';
/**
* The suite runs against SQLite in memory by default and against a real Postgres when
* `TEST_DATABASE_URL` is set, which is how the CI matrix covers both dialects with one set
* of tests (SPEC.md section 13).
*/
const POSTGRES_URL = process.env['TEST_DATABASE_URL'];
export const TEST_ENV: Record<string, string> = {
NODE_ENV: 'test',
DATABASE_URL: 'sqlite::memory:',
DATABASE_URL: POSTGRES_URL ?? 'sqlite::memory:',
BETTER_AUTH_SECRET: 'test-secret-that-is-long-enough-32chars',
BETTER_AUTH_URL: 'http://localhost:3000',
APP_PUBLIC_URL: 'http://localhost:3000',
@@ -47,11 +56,35 @@ export interface Harness extends AppHandle {
}
export async function createHarness(options: HarnessOptions = {}): Promise<Harness> {
const parsed = parseEnv({ ...TEST_ENV, ...(options.env ?? {}) });
/**
* On Postgres every harness gets a database of its own. `sqlite::memory:` gives each test
* a private database for free; a shared Postgres does not, and tests that seed the same
* four accounts into one set of tables would tread on each other.
*
* A database rather than a schema: Kysely's migrator asks the introspector whether its
* bookkeeping tables exist and, with no schema configured, a table of that name in any
* schema counts. Twenty parallel schemas each holding a `kysely_migration` therefore
* convince each other that the work is already done.
*/
const scratchDatabase = POSTGRES_URL ? `t_${uuidv7().replaceAll('-', '')}` : null;
if (scratchDatabase) await createDatabase(POSTGRES_URL as string, scratchDatabase);
const parsed = parseEnv({
...TEST_ENV,
...(scratchDatabase
? {
DATABASE_URL: withDatabase(POSTGRES_URL as string, scratchDatabase),
// Small on purpose: one pool per harness, many harnesses at once, and one
// server's max_connections between them.
DATABASE_POOL_MAX: '2',
}
: {}),
...(options.env ?? {}),
});
if (!parsed.ok || !parsed.env) throw new Error(parsed.message);
const env = parsed.env;
const handle = createDb(env.DATABASE_URL);
const handle = createDb(env.DATABASE_URL, { poolMax: env.DATABASE_POOL_MAX });
await migrateToLatest(handle, env);
const auth = createAuth({ db: handle.db, dialect: handle.dialect, env, sendOtp: async () => undefined });
@@ -92,6 +125,8 @@ export async function createHarness(options: HarnessOptions = {}): Promise<Harne
close: async () => {
await poller.stop();
await handle.close();
// After the pool is closed: Postgres refuses to drop a database anything is on.
if (scratchDatabase) await dropDatabase(POSTGRES_URL as string, scratchDatabase);
},
signIn: async (email, password) => {
const response = await appHandle.app.request('/api/auth/sign-in/email', {
@@ -106,3 +141,33 @@ export async function createHarness(options: HarnessOptions = {}): Promise<Harne
},
};
}
function withDatabase(databaseUrl: string, name: string): string {
const url = new URL(databaseUrl);
url.pathname = `/${name}`;
return url.toString();
}
/** `create database` cannot run inside a transaction, so it gets its own connection. */
async function createDatabase(databaseUrl: string, name: string): Promise<void> {
await onAdminConnection(databaseUrl, (client) => client.query(`create database "${name}"`));
}
async function dropDatabase(databaseUrl: string, name: string): Promise<void> {
await onAdminConnection(databaseUrl, (client) =>
client.query(`drop database if exists "${name}" with (force)`),
);
}
async function onAdminConnection(
databaseUrl: string,
run: (client: pg.Client) => Promise<unknown>,
): Promise<void> {
const client = new pg.Client({ connectionString: databaseUrl });
await client.connect();
try {
await run(client);
} finally {
await client.end();
}
}
+38
View File
@@ -0,0 +1,38 @@
# Kubernetes manifests
Scaled mode (SPEC.md section 15): **Postgres and S3 are required.** SQLite is not
supported here and the API refuses the configuration that would need it: a dedicated
worker means `JOBS_INLINE=false`, which env validation rejects on SQLite because a second
poller against one SQLite file is a corruption waiting to happen. Local file storage is
equally unsupported unless every replica mounts the same RWX volume, which S3 exists to
avoid.
Apply in order:
```bash
kubectl apply -f namespace.yaml
kubectl apply -f config.yaml # edit first: hostnames, bucket, replicas
kubectl apply -f secret.example.yaml # do not commit real values, see the note inside
kubectl apply -f api.yaml
kubectl apply -f worker.yaml
kubectl apply -f web.yaml
kubectl apply -f ingress.yaml
kubectl apply -f hpa.yaml
```
What each piece is for:
| File | What it does |
|---|---|
| `namespace.yaml` | One namespace, so everything can be removed in one command |
| `config.yaml` | Non-secret settings, the same keys as `apps/api/.env.example` |
| `secret.example.yaml` | The shape of the Secret. Real values come from your secret store |
| `api.yaml` | The HTTP API, N replicas, migrations as an initContainer, plus its Service |
| `worker.yaml` | The job poller, `ROLE=worker`, safe at N replicas on Postgres |
| `web.yaml` | The Next server, plus its Service. Never talks to the database |
| `ingress.yaml` | Public traffic reaches the web Service only. `/api` is proxied inside it |
| `hpa.yaml` | Scales the API on CPU. The worker is not autoscaled; see the note in it |
Every API and worker pod runs `db:migrate` before it serves. That is safe: the migration
takes a Postgres advisory lock on one pinned connection, so replicas starting together
queue rather than race.
+84
View File
@@ -0,0 +1,84 @@
apiVersion: apps/v1
kind: Deployment
metadata:
name: impuestos-api
namespace: impuestos
labels: { app: impuestos, component: api }
spec:
replicas: 2
selector:
matchLabels: { app: impuestos, component: api }
template:
metadata:
labels: { app: impuestos, component: api }
spec:
# 30s: the process drains in-flight requests for up to 25s before it exits.
terminationGracePeriodSeconds: 30
initContainers:
# Every replica runs this. Concurrent runs are safe: the migration holds a Postgres
# advisory lock on one pinned connection, so the second pod waits for the first.
- name: migrate
image: ghcr.io/example/impuestos-api:latest
command: ['node', 'dist/db/migrate.cli.js']
envFrom:
- configMapRef: { name: impuestos-config }
- secretRef: { name: impuestos-secrets }
resources:
requests: { cpu: 50m, memory: 128Mi }
limits: { memory: 256Mi }
containers:
- name: api
image: ghcr.io/example/impuestos-api:latest
ports:
- name: http
containerPort: 4000
env:
- name: ROLE
value: 'server'
# The dedicated worker Deployment does the jobs. An API replica that also
# polled would multiply the pollers by the replica count.
- name: JOBS_INLINE
value: 'false'
envFrom:
- configMapRef: { name: impuestos-config }
- secretRef: { name: impuestos-secrets }
# Liveness answers as long as the process is alive; readiness also checks the
# database and the storage driver, and goes false the moment a drain begins.
livenessProbe:
httpGet: { path: /healthz, port: http }
initialDelaySeconds: 5
periodSeconds: 10
readinessProbe:
httpGet: { path: /readyz, port: http }
initialDelaySeconds: 2
periodSeconds: 5
failureThreshold: 2
resources:
requests: { cpu: 100m, memory: 256Mi }
limits: { memory: 512Mi }
volumeMounts:
- name: tmp
mountPath: /tmp
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities: { drop: ['ALL'] }
volumes:
# The root filesystem is read only, so the one place anything may be written is an
# empty directory that dies with the pod. Nothing durable belongs here: uploads go
# through the storage driver to S3.
- name: tmp
emptyDir: {}
---
apiVersion: v1
kind: Service
metadata:
name: impuestos-api
namespace: impuestos
spec:
type: ClusterIP
selector: { app: impuestos, component: api }
ports:
- name: http
port: 4000
targetPort: http
+47
View File
@@ -0,0 +1,47 @@
# Non-secret configuration. Every key here exists in apps/api/.env.example with the same
# meaning; nothing is invented for Kubernetes.
apiVersion: v1
kind: ConfigMap
metadata:
name: impuestos-config
namespace: impuestos
data:
NODE_ENV: 'production'
PORT: '4000'
# The origin a person types. Cookies are issued for it and links in email point at it.
APP_PUBLIC_URL: 'https://impuestos.example'
BETTER_AUTH_URL: 'https://impuestos.example'
# Postgres connections per pod. Multiply by (api replicas + worker replicas) and keep the
# total under the server's max_connections.
DATABASE_POOL_MAX: '10'
JOBS_POLL_INTERVAL_MS: '2000'
JOBS_STALE_MINUTES: '10'
STORAGE_DRIVER: 's3'
S3_REGION: 'us-east-1'
S3_BUCKET: 'impuestos-comprobantes'
S3_FORCE_PATH_STYLE: 'true'
# Leave empty for AWS; set it for MinIO or another S3 compatible service.
S3_ENDPOINT: ''
OCR_MODEL: 'claude-sonnet-4-6'
DEFAULT_LOCALE: 'es'
SMTP_PORT: '587'
SMTP_FROM: 'avisos@impuestos.example'
# RFC 8292 wants an https: or mailto: URL. Without one, push switches itself off.
PUSH_VAPID_SUBJECT: 'mailto:soporte@impuestos.example'
---
apiVersion: v1
kind: ConfigMap
metadata:
name: impuestos-web-config
namespace: impuestos
data:
NODE_ENV: 'production'
# Cluster DNS for the api Service. The browser never sees this address.
API_INTERNAL_URL: 'http://impuestos-api:4000'
NEXT_PUBLIC_DEFAULT_LOCALE: 'es'
+42
View File
@@ -0,0 +1,42 @@
# The API scales on CPU. Stateless by construction: no local writes outside the storage
# driver, no in-memory cache that changes an answer, so a new replica behaves like an old
# one from its first request.
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: impuestos-api
namespace: impuestos
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: impuestos-api
minReplicas: 2
maxReplicas: 10
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
---
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: impuestos-web
namespace: impuestos
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: impuestos-web
minReplicas: 2
maxReplicas: 6
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
+27
View File
@@ -0,0 +1,27 @@
# Public traffic reaches the web Service and nothing else. The API is not exposed: the
# browser talks to one origin, the web server forwards /api inside the cluster, and the
# session cookie stays first party. SPEC.md is explicit that there is no CORS in v1, and
# an Ingress that published the API would be the thing that made CORS necessary.
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: impuestos
namespace: impuestos
annotations:
# A scan is a photograph.
nginx.ingress.kubernetes.io/proxy-body-size: '12m'
spec:
ingressClassName: nginx
tls:
- hosts: ['impuestos.example']
secretName: impuestos-tls
rules:
- host: impuestos.example
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: impuestos-web
port: { name: http }
+4
View File
@@ -0,0 +1,4 @@
apiVersion: v1
kind: Namespace
metadata:
name: impuestos
+29
View File
@@ -0,0 +1,29 @@
# The shape only. Do not commit real values: point your secret store at this name instead,
# or create it once by hand:
#
# kubectl -n impuestos create secret generic impuestos-secrets \
# --from-literal=DATABASE_URL='postgres://user:pass@host:5432/impuestos' \
# --from-literal=BETTER_AUTH_SECRET="$(openssl rand -base64 32)" \
# --from-literal=S3_ACCESS_KEY_ID=... --from-literal=S3_SECRET_ACCESS_KEY=...
#
# Optional keys may be left out entirely. The API degrades honestly without them: no
# ANTHROPIC_API_KEY sends a scan with no QR to the manual form, no SMTP_HOST logs
# verification codes to stdout, and no VAPID keys hides push everywhere in the UI.
apiVersion: v1
kind: Secret
metadata:
name: impuestos-secrets
namespace: impuestos
type: Opaque
stringData:
DATABASE_URL: 'postgres://impuestos:change-me@postgres:5432/impuestos'
BETTER_AUTH_SECRET: 'change-me-at-least-32-characters-long'
S3_ACCESS_KEY_ID: 'change-me'
S3_SECRET_ACCESS_KEY: 'change-me'
ANTHROPIC_API_KEY: ''
PUSH_VAPID_PUBLIC_KEY: ''
PUSH_VAPID_PRIVATE_KEY: ''
SMTP_HOST: ''
SMTP_USER: ''
SMTP_PASS: ''
TELEGRAM_BOT_TOKEN: ''
+64
View File
@@ -0,0 +1,64 @@
# The Next server. It holds no secrets and never opens a database connection: it renders
# and proxies /api to the api Service.
apiVersion: apps/v1
kind: Deployment
metadata:
name: impuestos-web
namespace: impuestos
labels: { app: impuestos, component: web }
spec:
replicas: 2
selector:
matchLabels: { app: impuestos, component: web }
template:
metadata:
labels: { app: impuestos, component: web }
spec:
terminationGracePeriodSeconds: 30
containers:
- name: web
image: ghcr.io/example/impuestos-web:latest
ports:
- name: http
containerPort: 3000
envFrom:
- configMapRef: { name: impuestos-web-config }
livenessProbe:
httpGet: { path: /healthz, port: http }
initialDelaySeconds: 5
periodSeconds: 10
readinessProbe:
httpGet: { path: /healthz, port: http }
initialDelaySeconds: 2
periodSeconds: 5
resources:
requests: { cpu: 100m, memory: 256Mi }
limits: { memory: 512Mi }
volumeMounts:
- name: tmp
mountPath: /tmp
# Next writes its own cache under the app directory at runtime.
- name: next-cache
mountPath: /app/apps/web/.next/cache
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities: { drop: ['ALL'] }
volumes:
- name: tmp
emptyDir: {}
- name: next-cache
emptyDir: {}
---
apiVersion: v1
kind: Service
metadata:
name: impuestos-web
namespace: impuestos
spec:
type: ClusterIP
selector: { app: impuestos, component: web }
ports:
- name: http
port: 3000
targetPort: http
+62
View File
@@ -0,0 +1,62 @@
# The job poller, on its own so it can be sized and restarted without touching the API.
#
# Safe at any replica count on Postgres: claiming uses FOR UPDATE SKIP LOCKED, which is
# covered by the hundred-job two-worker test in apps/api/src/modules/jobs/scale.test.ts.
# One is the default because the queue is small; raise it when the bandeja backs up.
apiVersion: apps/v1
kind: Deployment
metadata:
name: impuestos-worker
namespace: impuestos
labels: { app: impuestos, component: worker }
spec:
replicas: 1
selector:
matchLabels: { app: impuestos, component: worker }
template:
metadata:
labels: { app: impuestos, component: worker }
spec:
# A job that is mid-flight when the pod goes is recovered by the stale sweep rather
# than lost, but finishing is cheaper than recovering.
terminationGracePeriodSeconds: 30
containers:
- name: worker
image: ghcr.io/example/impuestos-api:latest
ports:
- name: http
containerPort: 4000
env:
# ROLE=worker is what starts the poller; JOBS_INLINE only decides whether a
# server also runs one.
- name: ROLE
value: 'worker'
envFrom:
- configMapRef: { name: impuestos-config }
- secretRef: { name: impuestos-secrets }
# A worker still serves /healthz and /readyz, which is how the cluster knows the
# poller has a database to poll. It is not behind a Service.
livenessProbe:
httpGet: { path: /healthz, port: http }
initialDelaySeconds: 5
periodSeconds: 10
readinessProbe:
httpGet: { path: /readyz, port: http }
initialDelaySeconds: 2
periodSeconds: 10
resources:
requests: { cpu: 100m, memory: 256Mi }
limits: { memory: 512Mi }
volumeMounts:
- name: tmp
mountPath: /tmp
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities: { drop: ['ALL'] }
volumes:
# The root filesystem is read only, so the one place anything may be written is an
# empty directory that dies with the pod. Nothing durable belongs here: uploads go
# through the storage driver to S3.
- name: tmp
emptyDir: {}
+2 -1
View File
@@ -34,6 +34,7 @@
"jsqr": "^1.4.0",
"typescript": "^5.9.3",
"typescript-eslint": "^8.69.0",
"vitest": "^5.0.0"
"vitest": "^5.0.0",
"yaml": "^2.9.0"
}
}
+28 -16
View File
@@ -49,7 +49,10 @@ importers:
version: 8.69.0(eslint@10.9.1(jiti@2.7.0))(typescript@5.9.3)
vitest:
specifier: ^5.0.0
version: 5.0.0(@types/node@26.4.1)(@vitest/coverage-v8@5.0.0)(vite@8.2.2(@types/node@26.4.1)(esbuild@0.28.2)(jiti@2.7.0)(tsx@4.23.13))
version: 5.0.0(@types/node@26.4.1)(@vitest/coverage-v8@5.0.0)(vite@8.2.2(@types/node@26.4.1)(esbuild@0.28.2)(jiti@2.7.0)(tsx@4.23.13)(yaml@2.9.0))
yaml:
specifier: ^2.9.0
version: 2.9.0
apps/api:
dependencies:
@@ -76,7 +79,7 @@ importers:
version: link:../../packages/rules
better-auth:
specifier: ^1.7.2
version: 1.7.2(better-sqlite3@13.0.3)(next@16.3.4(@babel/core@7.29.7)(@playwright/test@1.62.1)(@types/node@26.4.1)(react-dom@19.2.8(react@19.2.8))(react@19.2.8))(pg@8.23.0)(react-dom@19.2.8(react@19.2.8))(react@19.2.8)(vitest@5.0.0)
version: 1.7.2(better-sqlite3@13.0.3)(next@16.3.4(@babel/core@7.29.7)(@playwright/test@1.62.1)(@types/node@26.4.1)(react-dom@19.2.8(react@19.2.8))(react@19.2.8))(pg@8.23.0)(react-dom@19.2.8(react@19.2.8))(react@19.2.8)(vitest@5.0.0(@types/node@26.4.1)(@vitest/coverage-v8@5.0.0)(vite@8.2.2(@types/node@26.4.1)(esbuild@0.28.2)(jiti@2.7.0)(tsx@4.23.13)(yaml@2.9.0)))
better-sqlite3:
specifier: ^13.0.3
version: 13.0.3
@@ -134,7 +137,7 @@ importers:
version: 1.5.4
tsup:
specifier: ^8.5.1
version: 8.5.1(@swc/core@1.16.1(@swc/helpers@0.5.23))(jiti@2.7.0)(postcss@8.5.28)(tsx@4.23.13)(typescript@5.9.3)
version: 8.5.1(@swc/core@1.16.1(@swc/helpers@0.5.23))(jiti@2.7.0)(postcss@8.5.28)(tsx@4.23.13)(typescript@5.9.3)(yaml@2.9.0)
tsx:
specifier: ^4.23.13
version: 4.23.13
@@ -158,7 +161,7 @@ importers:
version: 5.102.8(react@19.2.8)
better-auth:
specifier: ^1.7.2
version: 1.7.2(better-sqlite3@13.0.3)(next@16.3.4(@babel/core@7.29.7)(@playwright/test@1.62.1)(@types/node@26.4.1)(react-dom@19.2.8(react@19.2.8))(react@19.2.8))(pg@8.23.0)(react-dom@19.2.8(react@19.2.8))(react@19.2.8)(vitest@5.0.0)
version: 1.7.2(better-sqlite3@13.0.3)(next@16.3.4(@babel/core@7.29.7)(@playwright/test@1.62.1)(@types/node@26.4.1)(react-dom@19.2.8(react@19.2.8))(react@19.2.8))(pg@8.23.0)(react-dom@19.2.8(react@19.2.8))(react@19.2.8)(vitest@5.0.0(@types/node@26.4.1)(@vitest/coverage-v8@5.0.0)(vite@8.2.2(@types/node@26.4.1)(esbuild@0.28.2)(jiti@2.7.0)(tsx@4.23.13)(yaml@2.9.0)))
class-variance-authority:
specifier: ^0.7.1
version: 0.7.1
@@ -3620,6 +3623,11 @@ packages:
yallist@3.1.1:
resolution: {integrity: sha512-a4UGQaWPH59mOXUYnAG2ewncQS4i4F43Tv3JoAM+s2VDAmS9NsK8GpDMLrCHPksFT7h3K6TOoUNn2pb7RoXx4g==}
yaml@2.9.0:
resolution: {integrity: sha512-2AvhNX3mb8zd6Zy7INTtSpl1F15HW6Wnqj0srWlkKLcpYl/gMIMJiyuGq2KeI2YFxUPjdlB+3Lc10seMLtL4cA==}
engines: {node: '>= 14.6'}
hasBin: true
yargs-parser@18.1.3:
resolution: {integrity: sha512-o50j0JeToy/4K6OZcaQmW6lyXXKhq7csREXcDwk2omFPJEwUNOVtJKvmDr9EI1fAJZUyZcRF7kxGBWmRXudrCQ==}
engines: {node: '>=6'}
@@ -4915,7 +4923,7 @@ snapshots:
obug: 2.1.4
std-env: 4.2.0
tinyrainbow: 3.1.1
vitest: 5.0.0(@types/node@26.4.1)(@vitest/coverage-v8@5.0.0)(vite@8.2.2(@types/node@26.4.1)(esbuild@0.28.2)(jiti@2.7.0)(tsx@4.23.13))
vitest: 5.0.0(@types/node@26.4.1)(@vitest/coverage-v8@5.0.0)(vite@8.2.2(@types/node@26.4.1)(esbuild@0.28.2)(jiti@2.7.0)(tsx@4.23.13)(yaml@2.9.0))
'@vitest/istanbul-lib-coverage@1.0.1': {}
@@ -4923,14 +4931,14 @@ snapshots:
dependencies:
'@vitest/istanbul-lib-coverage': 1.0.1
'@vitest/mocker@5.0.0(vite@8.2.2(@types/node@26.4.1)(esbuild@0.28.2)(jiti@2.7.0)(tsx@4.23.13))':
'@vitest/mocker@5.0.0(vite@8.2.2(@types/node@26.4.1)(esbuild@0.28.2)(jiti@2.7.0)(tsx@4.23.13)(yaml@2.9.0))':
dependencies:
'@jridgewell/trace-mapping': 0.3.31
'@vitest/spy': 5.0.0
estree-walker: 3.0.3
magic-string: 1.2.3
optionalDependencies:
vite: 8.2.2(@types/node@26.4.1)(esbuild@0.28.2)(jiti@2.7.0)(tsx@4.23.13)
vite: 8.2.2(@types/node@26.4.1)(esbuild@0.28.2)(jiti@2.7.0)(tsx@4.23.13)(yaml@2.9.0)
'@vitest/spy@5.0.0': {}
@@ -5041,7 +5049,7 @@ snapshots:
baseline-browser-mapping@2.11.21: {}
better-auth@1.7.2(better-sqlite3@13.0.3)(next@16.3.4(@babel/core@7.29.7)(@playwright/test@1.62.1)(@types/node@26.4.1)(react-dom@19.2.8(react@19.2.8))(react@19.2.8))(pg@8.23.0)(react-dom@19.2.8(react@19.2.8))(react@19.2.8)(vitest@5.0.0):
better-auth@1.7.2(better-sqlite3@13.0.3)(next@16.3.4(@babel/core@7.29.7)(@playwright/test@1.62.1)(@types/node@26.4.1)(react-dom@19.2.8(react@19.2.8))(react@19.2.8))(pg@8.23.0)(react-dom@19.2.8(react@19.2.8))(react@19.2.8)(vitest@5.0.0(@types/node@26.4.1)(@vitest/coverage-v8@5.0.0)(vite@8.2.2(@types/node@26.4.1)(esbuild@0.28.2)(jiti@2.7.0)(tsx@4.23.13)(yaml@2.9.0))):
dependencies:
'@better-auth/core': 1.7.2(@better-auth/utils@0.4.2)(@better-fetch/fetch@1.3.1)(better-call@1.4.0(zod@4.5.4))(jose@6.2.10)(kysely@0.29.5)(nanostores@1.5.3)
'@better-auth/drizzle-adapter': 1.7.2(@better-auth/core@1.7.2(@better-auth/utils@0.4.2)(@better-fetch/fetch@1.3.1)(better-call@1.4.0(zod@4.5.4))(jose@6.2.10)(kysely@0.29.5)(nanostores@1.5.3))(@better-auth/utils@0.4.2)
@@ -5066,7 +5074,7 @@ snapshots:
pg: 8.23.0
react: 19.2.8
react-dom: 19.2.8(react@19.2.8)
vitest: 5.0.0(@types/node@26.4.1)(@vitest/coverage-v8@5.0.0)(vite@8.2.2(@types/node@26.4.1)(esbuild@0.28.2)(jiti@2.7.0)(tsx@4.23.13))
vitest: 5.0.0(@types/node@26.4.1)(@vitest/coverage-v8@5.0.0)(vite@8.2.2(@types/node@26.4.1)(esbuild@0.28.2)(jiti@2.7.0)(tsx@4.23.13)(yaml@2.9.0))
transitivePeerDependencies:
- '@cloudflare/workers-types'
- '@opentelemetry/api'
@@ -6270,13 +6278,14 @@ snapshots:
possible-typed-array-names@1.1.0: {}
postcss-load-config@6.0.1(jiti@2.7.0)(postcss@8.5.28)(tsx@4.23.13):
postcss-load-config@6.0.1(jiti@2.7.0)(postcss@8.5.28)(tsx@4.23.13)(yaml@2.9.0):
dependencies:
lilconfig: 3.1.3
optionalDependencies:
jiti: 2.7.0
postcss: 8.5.28
tsx: 4.23.13
yaml: 2.9.0
postcss@8.5.23:
dependencies:
@@ -6679,7 +6688,7 @@ snapshots:
tslib@2.8.1: {}
tsup@8.5.1(@swc/core@1.16.1(@swc/helpers@0.5.23))(jiti@2.7.0)(postcss@8.5.28)(tsx@4.23.13)(typescript@5.9.3):
tsup@8.5.1(@swc/core@1.16.1(@swc/helpers@0.5.23))(jiti@2.7.0)(postcss@8.5.28)(tsx@4.23.13)(typescript@5.9.3)(yaml@2.9.0):
dependencies:
bundle-require: 5.1.0(esbuild@0.27.7)
cac: 6.7.14
@@ -6690,7 +6699,7 @@ snapshots:
fix-dts-default-cjs-exports: 1.0.1
joycon: 3.1.1
picocolors: 1.1.1
postcss-load-config: 6.0.1(jiti@2.7.0)(postcss@8.5.28)(tsx@4.23.13)
postcss-load-config: 6.0.1(jiti@2.7.0)(postcss@8.5.28)(tsx@4.23.13)(yaml@2.9.0)
resolve-from: 5.0.0
rollup: 4.63.1
source-map: 0.7.6
@@ -6799,7 +6808,7 @@ snapshots:
uuidv7@1.2.1: {}
vite@8.2.2(@types/node@26.4.1)(esbuild@0.28.2)(jiti@2.7.0)(tsx@4.23.13):
vite@8.2.2(@types/node@26.4.1)(esbuild@0.28.2)(jiti@2.7.0)(tsx@4.23.13)(yaml@2.9.0):
dependencies:
lightningcss: 1.33.0
picomatch: 4.0.7
@@ -6812,11 +6821,12 @@ snapshots:
fsevents: 2.3.3
jiti: 2.7.0
tsx: 4.23.13
yaml: 2.9.0
vitest@5.0.0(@types/node@26.4.1)(@vitest/coverage-v8@5.0.0)(vite@8.2.2(@types/node@26.4.1)(esbuild@0.28.2)(jiti@2.7.0)(tsx@4.23.13)):
vitest@5.0.0(@types/node@26.4.1)(@vitest/coverage-v8@5.0.0)(vite@8.2.2(@types/node@26.4.1)(esbuild@0.28.2)(jiti@2.7.0)(tsx@4.23.13)(yaml@2.9.0)):
dependencies:
'@types/chai': 5.2.3
'@vitest/mocker': 5.0.0(vite@8.2.2(@types/node@26.4.1)(esbuild@0.28.2)(jiti@2.7.0)(tsx@4.23.13))
'@vitest/mocker': 5.0.0(vite@8.2.2(@types/node@26.4.1)(esbuild@0.28.2)(jiti@2.7.0)(tsx@4.23.13)(yaml@2.9.0))
chai: 6.2.2
es-module-lexer: 2.3.2
expect-type: 1.4.0
@@ -6827,7 +6837,7 @@ snapshots:
tinybench: 6.1.4
tinyexec: 1.3.0
tinyglobby: 0.2.17
vite: 8.2.2(@types/node@26.4.1)(esbuild@0.28.2)(jiti@2.7.0)(tsx@4.23.13)
vite: 8.2.2(@types/node@26.4.1)(esbuild@0.28.2)(jiti@2.7.0)(tsx@4.23.13)(yaml@2.9.0)
why-is-node-running: 2.3.0
optionalDependencies:
'@types/node': 26.4.1
@@ -6911,6 +6921,8 @@ snapshots:
yallist@3.1.1: {}
yaml@2.9.0: {}
yargs-parser@18.1.3:
dependencies:
camelcase: 5.3.1
+7
View File
@@ -4,6 +4,13 @@ export default defineConfig({
test: {
include: ['{apps,packages}/*/src/**/*.test.ts'],
environment: 'node',
/**
* Against SQLite in memory a harness is built in milliseconds. Against a real Postgres
* it creates a database, runs both migrators and seeds forty documents, which is
* seconds of honest work rather than a hang.
*/
testTimeout: process.env['TEST_DATABASE_URL'] ? 60_000 : 5_000,
hookTimeout: process.env['TEST_DATABASE_URL'] ? 60_000 : 10_000,
coverage: {
provider: 'v8',
include: ['packages/rules/src/**'],