chore: remove the unused second stack definition

deploy/stack.yaml arrived in the first commit and was never deployed. Portainer
runs the repository's docker-compose.yml. Keeping both meant two definitions
drifting apart, with the documentation naming the one nobody used, which is how
the credential question came up at all.

The hardening it offered is narrower than it looks: Docker secrets keep values
out of docker inspect and the Portainer console, but local/secrets.py loads them
into the process environment regardless, and anyone able to read docker inspect
can already read the secret files. With a single Portainer user, the benefit that
remains does not outweigh maintaining a divergent copy.

local/secrets.py stays: inert against the deployed file, and it lets a stack
switch to Docker secrets later without touching code. The preflight and its tests
degrade cleanly when no such stack is present.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Cauê Faleiros
2026-09-21 13:48:22 -03:00
parent e95a42dbed
commit 9926d3ab55
5 changed files with 25 additions and 344 deletions

View File

@@ -35,7 +35,6 @@ jobs:
# taken on the machine itself. Known occupants of that host:
# 8000, 9443 Portainer (the Edge tunnel and its UI)
# 18080/18081 the production dtf-cloud stack (docker-compose.yml defaults)
# 8080/8081 deploy/stack.yaml defaults
# 9000/9001 MinIO defaults elsewhere
# This block avoids all of them. Ephemeral ports are not an option: the
# published port is baked into PUBLIC_ORIGIN, ALLOWED_ORIGINS and the CSP

View File

@@ -222,38 +222,30 @@ refactor: `'This runtime only supports APP_ENV=local'`,
- Replace marker matching with behavioural assertions (import the module, assert
the adapter classes in use).
### `[ ]` 2.12 — Production credentials are environment variables, and the docs name the wrong file
### `[x]` 2.12 — Two divergent stack definitions; the docs named the wrong one
Found 2026-09-21, by asking which file Portainer deploys.
Found 2026-09-21 by asking which file Portainer deploys. `PORTAINER.md` called
`deploy/stack.yaml` "the production stack"; the deployed file is the repository's
`docker-compose.yml`. `deploy/stack.yaml` came from the first commit and was never
deployed — it supplied credentials as Docker secrets where the deployed file uses
plain environment variables.
`PORTAINER.md` called `deploy/stack.yaml` "the production stack". The deployed
file is the repository's `docker-compose.yml`. They are not equivalent:
**Decision (2026-09-21): keep `docker-compose.yml`, delete `deploy/stack.yaml`.**
| File | Credentials | Deployed |
|---|---|---|
| `deploy/stack.yaml` | 11 Docker secret files | no |
| `docker-compose.yml` | 8 plain environment variables | yes |
The gain from Docker secrets here is narrower than it sounds. It keeps values out
of `docker inspect` and the Portainer UI, but `local/secrets.py` loads them into
the process environment anyway, and anyone who can read `docker inspect` is
already root or in the docker group and could read the secret files directly. The
operator is the only Portainer user, so the main benefit — limiting what a
lower-privileged console user can see — does not apply. Maintaining two
definitions that drift was the larger real cost.
So `DATABASE_PASSWORD`, `AWS_SECRET_ACCESS_KEY` and `OPERATOR_PASSWORD` sit in the
container environment, readable through `docker inspect`, `docker service inspect`,
the Portainer stack editor, and `/proc/<pid>/environ` for anything in that
container. The R2 secret key is the worst of them: it grants read and write over
every customer's artwork.
`local/secrets.py` stays. It is inert against the deployed file and costs nothing,
and it means a stack can switch to Docker secrets later without a code change.
The `*_FILE` loading from 2.3 makes the hardened file bootable, so the work is
done — what remains is a decision, because changing how production receives its
credentials is not a change to make quietly:
- **Migrate to `deploy/stack.yaml`:** create the 11 secrets in Portainer, point the
stack at that file. Best posture, most operational steps, and it also switches the
published ports (8080/8081 vs 18080/18081) so the reverse proxy needs updating.
- **Add Docker secrets to `docker-compose.yml`:** smaller change, keeps ports and
the current stack definition, still removes the values from the environment.
- **Accept it explicitly** and delete `deploy/stack.yaml` so two divergent
definitions stop drifting.
Rotate the R2 key and operator password whichever is chosen, since the current
values have been readable from the stack environment.
Still open: **rotate the R2 secret key.** Not because of Portainer, but because it
grants read and write over every customer's artwork and has been readable from the
stack environment for some time. The operator password is worth rotating with it.
### `[x]` 2.6 — Base images are not pinned `(F10)`

View File

@@ -4,7 +4,7 @@ The DTF application is one Portainer-owned Docker Swarm stack. Gitea builds,
tests, scans, and publishes the two application images, then calls the stack's
Portainer webhook. Start with the short operator guide in `../PORTAINER.md`.
- `stack.yaml` — the single Portainer stack.
- The deployed stack is the repository's `docker-compose.yml`, not a file here.
- `Dockerfile.api` and `Dockerfile.web` — prebuilt registry images.
- `portainer.env.example` — non-secret Portainer variables.
- `production_preflight.py` — fail-closed application/configuration validator.

View File

@@ -1,310 +0,0 @@
version: "3.8"
x-app-environment: &app-environment
APP_ENV: production
DATABASE_URL_FILE: /run/secrets/database_url
S3_ENDPOINT: ${R2_ENDPOINT:?set R2_ENDPOINT}
S3_PUBLIC_ENDPOINT: ${R2_PUBLIC_ENDPOINT:?set R2_PUBLIC_ENDPOINT}
S3_BUCKET: ${R2_BUCKET:?set R2_BUCKET}
AWS_ACCESS_KEY_ID_FILE: /run/secrets/r2_access_key_id
AWS_SECRET_ACCESS_KEY_FILE: /run/secrets/r2_secret_access_key
AWS_DEFAULT_REGION: auto
# The Kanban authenticates by email; the runtime reads OPERATOR_EMAIL.
OPERATOR_EMAIL: ${OPERATOR_EMAIL:?set OPERATOR_EMAIL}
OPERATOR_PASSWORD_FILE: /run/secrets/operator_password
PAYMENT_ADAPTER: ${PAYMENT_ADAPTER:?set PAYMENT_ADAPTER}
FREIGHT_ADAPTER: ${FREIGHT_ADAPTER:?set FREIGHT_ADAPTER}
TINY_ADAPTER: ${TINY_ADAPTER:?set TINY_ADAPTER}
WHATSAPP_ADAPTER: ${WHATSAPP_ADAPTER:?set WHATSAPP_ADAPTER}
STORAGE_ADAPTER: s3-r2
PAYMENT_TOKEN_FILE: /run/secrets/payment_token
PAYMENT_WEBHOOK_SECRET_FILE: /run/secrets/payment_webhook_secret
TINY_TOKEN_FILE: /run/secrets/tiny_token
WHATSAPP_TOKEN_FILE: /run/secrets/whatsapp_token
PUBLIC_ORIGIN: ${PUBLIC_ORIGIN:?set PUBLIC_ORIGIN}
PUBLIC_HOST: ${PUBLIC_HOST:?set PUBLIC_HOST}
ALLOWED_HOSTS: ${PUBLIC_HOST:?set PUBLIC_HOST},${KANBAN_HOST:?set KANBAN_HOST}
ALLOWED_ORIGINS: ${PUBLIC_ORIGIN:?set PUBLIC_ORIGIN},https://${KANBAN_HOST:?set KANBAN_HOST}
COOKIE_SECURE: "true"
MAX_UPLOAD_BYTES: ${MAX_UPLOAD_BYTES:-5368709120}
UPLOAD_PART_BYTES: ${UPLOAD_PART_BYTES:-8388608}
STORAGE_QUOTA_BYTES: ${STORAGE_QUOTA_BYTES:?set STORAGE_QUOTA_BYTES}
OWNER_UPLOAD_QUOTA_BYTES: ${OWNER_UPLOAD_QUOTA_BYTES:?set OWNER_UPLOAD_QUOTA_BYTES}
MAX_PENDING_UPLOADS: ${MAX_PENDING_UPLOADS:-10}
SCAN_MAX_BYTES: ${SCAN_MAX_BYTES:-134217728}
x-app-secrets: &app-secrets
- database_url
- r2_access_key_id
- r2_secret_access_key
- operator_password
- payment_token
- payment_webhook_secret
- tiny_token
- whatsapp_token
x-rolling: &rolling
update_config:
parallelism: 1
delay: 10s
order: start-first
failure_action: rollback
monitor: 45s
rollback_config:
parallelism: 1
delay: 5s
order: start-first
failure_action: pause
monitor: 45s
restart_policy:
condition: on-failure
delay: 5s
max_attempts: 5
window: 60s
services:
db:
image: ${POSTGRES_IMAGE:?set POSTGRES_IMAGE}
environment:
POSTGRES_DB: ${POSTGRES_DB:?set POSTGRES_DB}
POSTGRES_USER: ${POSTGRES_USER:?set POSTGRES_USER}
POSTGRES_PASSWORD_FILE: /run/secrets/db_admin_password
secrets: [db_admin_password]
volumes:
- postgres-data:/var/lib/postgresql/data
networks: [backend]
healthcheck:
test: [CMD-SHELL, 'pg_isready -U "$$POSTGRES_USER" -d "$$POSTGRES_DB"']
interval: 10s
timeout: 5s
retries: 12
start_period: 20s
stop_grace_period: 60s
deploy:
replicas: 1
placement:
constraints: [node.labels.dtf_database == true]
update_config:
parallelism: 1
order: stop-first
failure_action: rollback
monitor: 60s
rollback_config:
parallelism: 1
order: stop-first
failure_action: pause
monitor: 60s
restart_policy:
condition: on-failure
delay: 10s
max_attempts: 5
window: 120s
resources:
limits: {cpus: "2.0", memory: 4G}
reservations: {cpus: "0.5", memory: 1G}
db-init:
image: ${API_IMAGE:?set API_IMAGE}:${IMAGE_TAG:-latest}
command: python -m local.bootstrap
environment:
APP_ENV: production
DATABASE_ADMIN_URL_FILE: /run/secrets/database_admin_url
APP_DB_USER: ${APP_DB_USER:?set APP_DB_USER}
APP_DB_PASSWORD_FILE: /run/secrets/app_db_password
secrets: [database_admin_url, app_db_password]
networks: [backend]
deploy:
replicas: 1
restart_policy: {condition: none}
placement:
constraints: [node.platform.os == linux]
resources:
limits: {cpus: "0.5", memory: 512M}
scanner:
image: ${CLAMAV_IMAGE:?set CLAMAV_IMAGE}
user: "100:101"
entrypoint: [clamd, --foreground=true, --config-file=/etc/clamav/clamd.conf]
configs:
- source: clamd_config
target: /etc/clamav/clamd.conf
mode: 0444
networks: [backend]
read_only: true
cap_drop: [ALL]
security_opt: [no-new-privileges:true]
tmpfs:
- /tmp:uid=100,gid=101,mode=0750
- /run/clamav:uid=100,gid=101,mode=0750
- /var/log/clamav:uid=100,gid=101,mode=0750
healthcheck:
test: [CMD, clamdscan, --config-file=/etc/clamav/clamd.conf, --ping, "3"]
interval: 15s
timeout: 5s
retries: 20
start_period: 90s
deploy:
replicas: 1
restart_policy: {condition: on-failure, delay: 10s}
resources:
limits: {cpus: "2.0", memory: 3G}
reservations: {cpus: "0.5", memory: 1G}
api:
image: ${API_IMAGE:?set API_IMAGE}:${IMAGE_TAG:-latest}
environment: *app-environment
secrets: *app-secrets
networks: [backend, egress]
read_only: true
tmpfs: [/tmp]
init: true
cap_drop: [ALL]
security_opt: [no-new-privileges:true]
healthcheck:
test:
- CMD-SHELL
- >-
python -c "import os,urllib.request; r=urllib.request.Request('http://localhost:8000/health',headers={'Host':os.environ['PUBLIC_HOST']}); urllib.request.urlopen(r,timeout=3)"
interval: 10s
timeout: 5s
retries: 12
start_period: 30s
stop_grace_period: 30s
deploy:
<<: *rolling
replicas: 2
resources:
limits: {cpus: "1.0", memory: 1G}
reservations: {cpus: "0.25", memory: 256M}
worker:
image: ${API_IMAGE:?set API_IMAGE}:${IMAGE_TAG:-latest}
command: python -m local.worker
environment:
<<: *app-environment
CLAMD_HOST: scanner
secrets: *app-secrets
networks: [backend, egress]
read_only: true
tmpfs: [/tmp]
init: true
cap_drop: [ALL]
security_opt: [no-new-privileges:true]
healthcheck:
test: [CMD, python, -c, "import urllib.request; urllib.request.urlopen('http://localhost:8002/health',timeout=3)"]
interval: 15s
timeout: 5s
retries: 12
start_period: 90s
stop_grace_period: 60s
deploy:
<<: *rolling
replicas: 1
update_config:
parallelism: 1
order: stop-first
failure_action: rollback
monitor: 60s
rollback_config:
parallelism: 1
order: stop-first
failure_action: pause
monitor: 60s
resources:
limits: {cpus: "1.5", memory: 2G}
reservations: {cpus: "0.25", memory: 512M}
site:
image: ${WEB_IMAGE:?set WEB_IMAGE}:${IMAGE_TAG:-latest}
environment:
WEB_INDEX: index.html
PUBLIC_HOST: ${PUBLIC_HOST:?set PUBLIC_HOST}
S3_PUBLIC_ENDPOINT: ${R2_PUBLIC_ENDPOINT:?set R2_PUBLIC_ENDPOINT}
networks: [backend]
ports:
- target: 8080
published: ${SITE_PORT:-8080}
protocol: tcp
mode: ingress
read_only: true
tmpfs:
- /tmp:uid=101,gid=101,mode=0750
- /var/cache/nginx:uid=101,gid=101,mode=0750
- /var/run:uid=101,gid=101,mode=0750
- /etc/nginx/conf.d:uid=101,gid=101,mode=0750
cap_drop: [ALL]
security_opt: [no-new-privileges:true]
healthcheck:
test: [CMD-SHELL, 'wget -q --header="Host: $$PUBLIC_HOST" -O /dev/null http://127.0.0.1:8080/health']
interval: 10s
timeout: 5s
retries: 12
start_period: 15s
deploy:
<<: *rolling
replicas: 2
resources:
limits: {cpus: "0.5", memory: 256M}
reservations: {cpus: "0.1", memory: 64M}
kanban:
image: ${WEB_IMAGE:?set WEB_IMAGE}:${IMAGE_TAG:-latest}
environment:
WEB_INDEX: kanban.html
PUBLIC_HOST: ${KANBAN_HOST:?set KANBAN_HOST}
S3_PUBLIC_ENDPOINT: ${R2_PUBLIC_ENDPOINT:?set R2_PUBLIC_ENDPOINT}
networks: [backend]
ports:
- target: 8080
published: ${KANBAN_PORT:-8081}
protocol: tcp
mode: ingress
read_only: true
tmpfs:
- /tmp:uid=101,gid=101,mode=0750
- /var/cache/nginx:uid=101,gid=101,mode=0750
- /var/run:uid=101,gid=101,mode=0750
- /etc/nginx/conf.d:uid=101,gid=101,mode=0750
cap_drop: [ALL]
security_opt: [no-new-privileges:true]
healthcheck:
test: [CMD-SHELL, 'wget -q --header="Host: $$PUBLIC_HOST" -O /dev/null http://127.0.0.1:8080/health']
interval: 10s
timeout: 5s
retries: 12
start_period: 15s
deploy:
<<: *rolling
replicas: 1
resources:
limits: {cpus: "0.5", memory: 256M}
reservations: {cpus: "0.1", memory: 64M}
configs:
clamd_config:
file: ../local/clamd.conf
secrets:
database_url: {external: true, name: "${DATABASE_URL_SECRET:?set DATABASE_URL_SECRET}"}
database_admin_url: {external: true, name: "${DATABASE_ADMIN_URL_SECRET:?set DATABASE_ADMIN_URL_SECRET}"}
db_admin_password: {external: true, name: "${DB_ADMIN_PASSWORD_SECRET:?set DB_ADMIN_PASSWORD_SECRET}"}
app_db_password: {external: true, name: "${APP_DB_PASSWORD_SECRET:?set APP_DB_PASSWORD_SECRET}"}
r2_access_key_id: {external: true, name: "${R2_ACCESS_KEY_ID_SECRET:?set R2_ACCESS_KEY_ID_SECRET}"}
r2_secret_access_key: {external: true, name: "${R2_SECRET_ACCESS_KEY_SECRET:?set R2_SECRET_ACCESS_KEY_SECRET}"}
operator_password: {external: true, name: "${OPERATOR_PASSWORD_SECRET:?set OPERATOR_PASSWORD_SECRET}"}
payment_token: {external: true, name: "${PAYMENT_TOKEN_SECRET:?set PAYMENT_TOKEN_SECRET}"}
payment_webhook_secret: {external: true, name: "${PAYMENT_WEBHOOK_SECRET:?set PAYMENT_WEBHOOK_SECRET}"}
tiny_token: {external: true, name: "${TINY_TOKEN_SECRET:?set TINY_TOKEN_SECRET}"}
whatsapp_token: {external: true, name: "${WHATSAPP_TOKEN_SECRET:?set WHATSAPP_TOKEN_SECRET}"}
volumes:
postgres-data:
external: true
name: ${POSTGRES_VOLUME:?set POSTGRES_VOLUME}
networks:
backend:
driver: overlay
internal: true
egress:
driver: overlay

View File

@@ -1,9 +1,9 @@
"""Resolve Docker secret files into the environment before configuration is read.
Swarm mounts each secret as a file and the stack passes its path as `<NAME>_FILE`.
Nothing read `_FILE` settings, so `deploy/stack.yaml` could not boot: the runtime
looked for `DATABASE_URL`, `AWS_ACCESS_KEY_ID` and `OPERATOR_PASSWORD` while the
stack supplied only the `_FILE` form.
The deployed `docker-compose.yml` passes credentials as plain environment
variables, so this module is inert there. It exists so a stack can supply them as
Docker secrets instead without any code change; see `ROADMAP.md` 2.12.
Call `load()` in every entrypoint before any configuration is read.
@@ -12,9 +12,9 @@ secrets` elsewhere in the package to the standard library, not to this file.
"""
import os
# The settings production supplies as secret files. Any other `*_FILE` variable is
# The settings a stack may supply as secret files. Any other `*_FILE` variable is
# resolved the same way; this list documents the contract and is what the release
# gate checks against, so keep it in step with `deploy/stack.yaml`.
# gate checks against.
SECRET_FILE_SETTINGS = (
'DATABASE_URL',
'DATABASE_ADMIN_URL',