Backups & Upgrades
Where state lives
| Store | Holds | If it is lost |
|---|---|---|
| Database (PostgreSQL or SQLite) | Organizations and their encrypted signing keys, users, clients, branding, roles, sign-in sessions, consent approvals, MFA factors and recovery codes, the audit log | Everything. Restore from backup. |
Redis (cache.*) | Authorization codes, refresh tokens and their revocation records, pushed authorization requests, device codes, DPoP nonces and replay records, brute-force throttles, rate-limit windows (with rate_limit.backend: redis), cached discovery and branding | Every refresh token and every half-finished sign-in. Browser sessions survive, so users are usually signed back in to their applications without a password prompt. |
| Secret store | The KEK, the MFA seed key, signing keys, credentials | See Keys & Secrets. |
Redis is not a disposable cache here: refresh tokens exist nowhere else. Run it so that it keeps what it is given:
- Turn on persistence (AOF) if refresh tokens should survive a Redis restart.
- Do not let Redis evict keys under memory pressure (
maxmemory-policy noeviction). An evicted refresh token is indistinguishable from a revoked one. - Size it for your traffic and watch
nauthera_cache_operations_totalandnauthera_cache_operation_seconds.
With rate_limit.backend: redis, a Redis outage makes every protocol and API endpoint
answer 429 until Redis returns, because those endpoints fail closed. See
Rate Limits & Lockout.
Backups
There is no built-in backup. Back up the database with your usual tooling (for example
CloudNativePG's backups or pg_dump), and back up the secrets that make it readable at
the same time:
- the key-encryption key and every retired one (
keys.encryption_key,keys.retired_keys), - the MFA seed key and every retired one (
mfa.encryption_key,mfa.retired_keys), - the main organization's signing keys.
A restore needs all of them. The database alone restores the main organization; it cannot restore other organizations' signing keys without the KEK, or anyone's second factor without the MFA key.
On SQLite, back up the database file while the server is stopped, or use SQLite's online backup.
Schema migrations
The migrations are compiled into the binary. With database.migrate: true, each replica
applies any pending migrations when it starts, before it opens its listeners.
- Forward only. There are no down migrations. Rolling back to an older binary after a migration has run is not supported; roll forward, or restore a backup.
- Safe with several replicas. On PostgreSQL the runner holds an advisory lock on one dedicated connection, so replicas migrate one after another instead of racing.
- Atomic. Each migration and its record in
schema_migrationscommit in the same transaction. - Tamper-evident. A checksum of every applied migration is stored and checked on every start. If an applied migration was changed, startup fails.
- PgBouncer. The advisory lock needs a single session. Behind PgBouncer in
transaction-pooling mode (
database.pgbouncer: true), migrations should go over a direct or session-pooled connection; the server logs a warning when they would not.
The binary has no migrate-and-exit mode, so a Kubernetes Job or init container that
runs the image with migrate: true never completes. The migration-job.yaml in the
repository's deploy/k8s/ directory does not work for this reason. Let the Deployment
migrate instead.
With database.migrate: false, the server does not check the schema version at all. A
new binary on an old schema then fails at the first query that needs a new table or
column.
Upgrading
No release has been published yet (#307), so upgrades today mean building a newer commit. When releases exist, the same procedure will apply:
- Back up the database and the secrets listed above.
- Roll out the new image with
database.migrate: true. The chart's Deployment and the raw manifests usemaxUnavailable: 0, so old pods keep serving until new ones are ready. - During the rollout, old and new pods share one schema. The repository's rule for migrations is expand, then contract: add columns and tables before code uses them, and drop them only once no running version needs them.
On SIGTERM a replica reports not ready, waits server.drain_delay (5s) for load
balancers to stop sending it traffic, and then drains for up to
server.shutdown_timeout (25s). A second signal forces it to exit.
What is cleaned up automatically
| Data | Cleanup |
|---|---|
| Sign-in sessions | Sessions past their absolute lifetime are deleted every hour by each replica. |
| Audit log | Partitions or rows older than audit.retention (400 days) are dropped every audit.maintenance_interval (6h). See Audit Log. |
| Codes, refresh tokens, device codes, nonces | Expire in Redis with their own lifetime. |
| Consent approvals | Never: they do not expire (#321). |
| Expired pending MFA sign-ins | Not swept yet. The table grows slowly. |