feat: daily encrypted database backup to a bucket of its own
All checks were successful
Build and deploy / Validate source (push) Successful in 8s
Build and deploy / Integration suite on a real stack (push) Successful in 3m38s
Build and deploy / Secret scan and release gate (push) Successful in 7s
Build and deploy / Publish images (push) Successful in 1m2s

A backup service runs pg_dump every day at 03:00 Brasília, checks the archive,
encrypts it with age to a public key and uploads it with a token for that
bucket only. The server cannot read or delete backups: the private key stays
with the owner, the bucket's lifecycle rule expires copies and its lock stops
early deletion. Each run is recorded and shown on the Kanban's Integrations
tab. tests/backup_test.py backs up, restores into a scratch database and
compares the rows in CI. Setup and restore: docs/BACKUP.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
Cauê Faleiros
2026-09-29 18:49:21 -03:00
parent 1b57313496
commit 20403c5132
14 changed files with 565 additions and 2 deletions

87
docs/BACKUP.md Normal file
View File

@@ -0,0 +1,87 @@
# Database backup
The `backup` service copies the database off the server once a day
(`ops/db_backup.py`). It runs `pg_dump`, checks the archive, encrypts it with
[age](https://age-encryption.org) and uploads it to an R2 bucket used for
nothing else. Each run is recorded, and the Kanban shows the latest one under
Integrações → "Backup do banco" ("Em dia", "Verificar" or "Não configurado").
The artwork is not backed up. Originals and print files are temporary (7 to 30
days) and already stored on R2. The database holds what cannot be recreated:
orders, customers, quotes, payments, the Tiny connection and the operators.
## Why the backup cannot be read or deleted from the server
- **The server only encrypts.** It holds the public key (`age1…`). The private
key, which is the only way to open a backup, stays with the owner.
- **Its own bucket and token.** The backup token reaches the backup bucket
only, and the artwork token cannot reach it.
- **Nothing is deleted by the job.** The bucket's lifecycle rule removes old
copies. The bucket lock keeps anyone, including someone holding the token,
from deleting or overwriting a copy before its time.
## Setup (once)
1. **Key pair.** Generate it on your own computer, not on the server:
```bash
docker run --rm gitea.blyzer.com.br/blyzer/dtf-api:latest age-keygen
```
Or, with age installed (`sudo pacman -S age`, `apt install age`), just
run `age-keygen`.
Save the whole output (the `AGE-SECRET-KEY-1…` line is the private key) in
the password manager. Only the `public key: age1…` value goes to Portainer.
Without the private key, no backup can ever be restored.
2. **Cloudflare R2 → Create bucket.** Use a name such as `dtf-backups`.
3. **Settings on that bucket:**
- Object lifecycle rules: delete objects after 35 days.
- Bucket lock rules: retain for 30 days, prefix `db/`.
4. **R2 → Manage API tokens → Create API token.**
- Permission: Object Read & Write.
- Scope: the `dtf-backups` bucket only.
5. **Portainer.** Add these to the stack's environment, then redeploy:
| Variable | Value |
|---|---|
| `BACKUP_S3_ENDPOINT` | same as `R2_ENDPOINT` |
| `BACKUP_BUCKET` | `dtf-backups` |
| `BACKUP_ACCESS_KEY_ID` | the new token's Access Key ID |
| `BACKUP_SECRET_ACCESS_KEY` | the new token's Secret Access Key |
| `BACKUP_AGE_RECIPIENT` | the public key, `age1…` |
| `BACKUP_HOUR` | optional, default `3` (03:00 Brasília) |
The first backup runs as soon as the service starts. After that it runs daily,
and a failed run is retried an hour later.
## Restore test
Do this after setup, and then once a month. Open the console of the `backup`
container in Portainer (Containers → `…_backup…` → Console) and run:
```bash
python -m ops.db_backup list
python -m ops.db_backup verify --identity -
```
`--identity -` asks for the private key to be pasted. It is kept only in the
container's memory while the command runs. `verify` restores the latest backup
into a scratch database, prints its row counts and drops it. The live database
is not touched.
## Restoring for real
From the `backup` container's console:
```bash
python -m ops.db_backup restore db/2026/10/dtf-20261001T060000Z.dump.age --into dtf_restored --identity -
```
This restores into a new database (or an empty one), never over the running
one. Then do one of these:
- **Check the restored data**, then point the stack at it.
- **On a new server**, deploy only the `db` service first. Restore into `dtf`
while it is still empty, then deploy the rest. `db-init` then applies the
schema and grants on top.

View File

@@ -864,6 +864,12 @@ print-file evidence still need correction before this item can close.
- `[ ]` 5.13 — Define production recovery: scheduled encrypted offsite database
and object backups, a consistent snapshot boundary, Swarm data placement and
a restore rehearsal that opens every required live order file.
**2026-09-29:** database part built. The `backup` service dumps daily,
encrypts with age (the server holds only the public key) and uploads to a
bucket of its own, whose lifecycle rule expires copies and whose lock keeps
them from being deleted early. The Kanban shows the latest run, and
`tests/backup_test.py` restores a copy in CI. Setup and restore:
docs/BACKUP.md. Artwork is not copied: it is temporary and already on R2.
- `[ ]` 5.14 — Promote and verify one immutable release. **2026-09-24:** green
pushes to `main` now publish images and Portainer's pull-and-redeploy is the
release gate; the source preflight is advisory unless enforced by variable,