Commit Graph

89 Commits

Author SHA1 Message Date
Cauê Faleiros
7cab210674 ci: move the integration stack clear of the ports that host already uses
8000 was Portainer's Edge tunnel, not a stray process. The first attempt at a
fix picked 18080/18081, which are the production dtf-cloud stack's own defaults
in docker-compose.yml: it would have passed only while that stack was down and
collided again the moment it came back.

Use 28080/28081/28000/29000/29001, clear of Portainer (8000, 9443), both
production stack definitions (18080/18081 and 8080/8081) and the usual MinIO
ports. The occupants are listed in the workflow so the next person choosing a
port can see what is taken.

Ephemeral ports would remove the guesswork but do not work here: the published
port is baked into PUBLIC_ORIGIN, ALLOWED_ORIGINS and the CSP when the
containers start, so it has to be known before they run.

Full suite verified on the new block, including the browser end-to-end.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 13:04:56 -03:00
Cauê Faleiros
24cb52d992 fix: stop the CI stack colliding with ports already used on the runner
The integration job failed with "Bind for 0.0.0.0:8000 failed: port is already
allocated". The runner shares the host's Docker daemon, so every published port
is claimed on the machine itself, where other services already listen. Port 8000
was the first collision; 8080, 8081, 9000 and 9001 were equally exposed.

MinIO's ports were hardcoded, and S3_PUBLIC_ENDPOINT was pinned to
localhost:9000 independently, so moving storage would have broken the presigned
URLs the browser fetches. Both now derive from STORAGE_PORT and move together.

CI runs on 18080/18081/18000/19000/19001. Local defaults are unchanged.

Verified by running the whole stack and the full suite on exactly those ports,
including the browser end-to-end, which downloads through a presigned URL and so
proves the storage endpoint followed the port.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 13:00:17 -03:00
Cauê Faleiros
6a50e6db4d fix: bake scanner and storage config into images instead of bind-mounting
Some checks failed
Build and deploy / Validate source (push) Successful in 6s
Build and deploy / Integration suite on a real stack (push) Failing after 54s
Build and deploy / Secret scan and release gate (push) Successful in 5s
Build and deploy / Publish images and notify Portainer (push) Has been skipped
The integration job failed starting the scanner:

  error mounting ".../local/clamd.conf" to rootfs at "/etc/clamav/clamd.conf":
  not a directory

The files are in the repository, so this was not a missing checkout. A
containerised CI runner shares the host's Docker daemon, so "./local/clamd.conf"
resolves to a workspace path that exists inside the runner but not on the host
where the daemon creates the mount. The daemon makes an empty directory there
and the container cannot start. Only the bind-mounting services were affected,
which is why PostgreSQL and MinIO came up first.

Build the scanner and storage-init images with their configuration copied in, so
compose.local.yaml no longer bind-mounts anything from the host and works
regardless of how the runner reaches the daemon. Both bases stay overridable
through CLAMAV_IMAGE and MINIO_IMAGE.

The production stack is unaffected: it ships clamd.conf as a Swarm config, which
the manager reads at deploy time.

Verified from a clean slate: the stack starts, the scanner runs the baked
configuration, storage provisioning runs from the baked script, and the full
suite passes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 12:51:08 -03:00
Cauê Faleiros
dbc9ba4dee docs: record an intermittent browser test failure
Some checks failed
Build and deploy / Validate source (push) Successful in 5s
Build and deploy / Integration suite on a real stack (push) Failing after 1m10s
Build and deploy / Secret scan and release gate (push) Successful in 5s
Build and deploy / Publish images and notify Portainer (push) Has been skipped
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 12:33:58 -03:00
Cauê Faleiros
6c52ad655e fix: pull MinIO from quay.io so CI can start the stack
The integration job failed on the runner with "pull access denied for
minio/minio ... may require 'docker login'". Docker Hub now refuses anonymous
pulls of minio/minio: an unauthenticated manifest request returns 401
UNAUTHORIZED, while library/postgres returns 200, which is why only MinIO
failed. It worked locally only because this machine is logged in to Docker Hub.

quay.io serves the same release anonymously, and it is the same image: both
registries resolve to image ID sha256:a1ea29fa2835. MINIO_IMAGE overrides it for
anyone mirroring into their own registry.

Verified by deleting the Docker Hub copy locally and starting the stack from
quay alone, then running the full suite against it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 12:33:46 -03:00
Cauê Faleiros
010c2a162f docs: record base image pinning and the CRITICAL image gate
Some checks failed
Build and deploy / Validate source (push) Successful in 6s
Build and deploy / Integration suite on a real stack (push) Failing after 6s
Build and deploy / Secret scan and release gate (push) Successful in 11s
Build and deploy / Publish images and notify Portainer (push) Has been skipped
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 11:58:46 -03:00
Cauê Faleiros
4c9fa2436e build: pin base image digests and clear every fixable image finding
The production images built on mutable tags with --pull, so the same commit
could produce different bases, and neither Dockerfile upgraded its OS packages
even though the local ones did. The published API image carried 56 HIGH and 3
CRITICAL findings, 15 of them with an upstream fix available.

Pin both bases by digest and upgrade OS packages in the production images. That
removes all 3 CRITICAL and 13 of the 15 fixable findings. The remaining two,
msgpack and setuptools, come from a third-party SBOM; neither package is
importable or listed by pip in the built image, which I confirmed rather than
taking the previous report's word for it.

The web image could not be fixed this way: the official 1.28 line pins
nginx=1.28.3-r1 in /etc/apk/world, so apk upgrade leaves five HIGH findings in
place even though Alpine ships 1.28.3-r7. Moving to nginx:alpine (1.31.6)
clears them completely; 1.29-alpine scans worse, at 37 HIGH. Same uid 101 and
the same template entrypoint, and the local images now use the same pinned
bases so the integration suite exercises what ships. Full suite passes on
nginx 1.31.6, including the browser end-to-end.

With both images at zero CRITICAL, the image scan now blocks on CRITICAL and
reports HIGH, instead of reporting everything. PYTHON_BASE_IMAGE and
NGINX_BASE_IMAGE are wired through to the builds so a base can move forward
without editing the repository, which is what PORTAINER.md already promised.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 11:58:32 -03:00
Cauê Faleiros
bbcc8ab100 docs: record the release gates and what they do not cover
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 11:45:23 -03:00
Cauê Faleiros
9da2a7dcad ci: add the release gates, and document the ones that do not gate
PORTAINER.md and SECURITY_REPORT.md described a pipeline that required
regressions, HIGH/CRITICAL secret, misconfiguration and image gates, and stated
that the source preflight stopped this application from publishing. None of it
ran: the workflow built and called the webhook unconditionally.

Add a blocking Trivy secret scan. Verified both ways: a planted AWS key pair,
GitHub token and private key block the job, and the repository passes clean.
Note that Trivy allowlists documented example credentials, so this gate is a
backstop, not permission to commit secrets.

The source preflight now runs on every push and always prints its verdict, but
enforces only when ENFORCE_PRODUCTION_PREFLIGHT is true. Enforcing it today
would block every deployment, because it refuses a release while the payment
and messaging adapters are fake, which is the deliberate state the stack runs
in. Set the variable when real adapters land.

Image vulnerabilities are reported after each build rather than enforced. The
current bases carry 56 HIGH and 3 CRITICAL findings, only 15 of them with an
upstream fix, so failing on them would stop releases without making anything
safer. Pinning digests and triaging the fixable ones is ROADMAP 2.6.

Both documents now carry a table of what gates and what does not, instead of
describing checks that did not exist.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 11:45:11 -03:00
Cauê Faleiros
d2f7b2c03b docs: record secret-file loading and the behavioural release gate
Some checks failed
Build and deploy / Validate source (push) Successful in 1m25s
Build and deploy / Integration suite on a real stack (push) Failing after 7s
Build and deploy / Publish images and notify Portainer (push) Has been skipped
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 11:37:16 -03:00
Cauê Faleiros
341f154c36 feat: load Docker secret files so the production stack can boot
deploy/stack.yaml passes DATABASE_URL_FILE, AWS_ACCESS_KEY_ID_FILE,
OPERATOR_PASSWORD_FILE and the provider tokens as Swarm secret paths, but the
runtime only ever read the plain names. That stack could not start: the database
URL and R2 credentials were absent, and operator login raised KeyError, so it
returned 500 instead of the intended 503.

local/secrets.py resolves every <NAME>_FILE into <NAME> before configuration is
read, from the API, worker and bootstrap entrypoints. It fails closed on an
unreadable or empty secret and on a name supplied both directly and as a file,
because starting with a credential nobody intended is worse than not starting.
Only one trailing newline is stripped, so a generated password keeps any
whitespace that belongs to it, and no value reaches an error message.

The stack also passed OPERATOR_USER while the Kanban authenticates by email;
it now passes OPERATOR_EMAIL, matching the runtime.

The release gate checked this by searching local/secrets.py for the literal
"DATABASE_URL_FILE", which would pass for any file containing that string. It
now loads the module and makes it resolve every secret the stack declares, and
asserts it fails closed on a missing one. Four marker strings that stopped
matching when R2 support landed are removed rather than left to rot; the two
that still describe real blockers stay, so the gate continues to refuse a
release while payment and messaging adapters are fake.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 11:36:59 -03:00
Cauê Faleiros
9b38c9fe3d docs: record Block 2 rate-limiting work and CI coverage
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 18:55:33 -03:00
Cauê Faleiros
5d669cbcce ci: run the real suites before publishing images
The pipeline ran py_compile plus four unit tests, then built and called the
Portainer webhook. None of that starts the application, so a missing import in
local/auth.py passed every check and reached production, where it returned 500
on every session, login and registration.

Add an integration job that builds the localhost stack and runs the suites that
already existed but were never executed automatically: smoke, workflow,
security, scanning, retention, runtime security, and the two browser tests.
publish-and-deploy now depends on it, so a failure blocks the deploy instead of
shipping.

Verified by reintroducing the original defect: py_compile and the unit tests
still passed, and smoke_test failed on /session, which would have stopped the
release.

The browser tests need a real Chrome and are skipped with a warning when the
runner has none; installing google-chrome-stable or setting CHROME_BIN makes
them gate too. Every other suite gates unconditionally.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 18:55:07 -03:00
Cauê Faleiros
52ae13c9a1 fix: rate-limit and audit by real client address
uvicorn does not trust forwarded headers from a peer outside
forwarded_allow_ips, so request.client.host was the web gateway for every
request. The auth-source bucket therefore counted all customers together:
60 failed logins from one attacker locked out everyone. Security events
recorded the gateway address, which made the audit trail useless for
attribution.

The gateway now overwrites X-Forwarded-For with the peer address it observed
instead of appending to whatever the client sent, so the header carries one
value the client cannot choose, and client_ip() resolves it with a fallback to
the connection peer.

The guest-session limiter was keyed on the environment name, making it one
global bucket of 120 per 15 minutes: roughly eight new visitors a minute for
the whole site before legitimate traffic started receiving 429. It is now per
source, and the ceiling is deliberately generous because offices and mobile
carriers put many real customers behind a single address.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 18:48:31 -03:00
Cauê Faleiros
aa0eba3457 docs: track outstanding work in ROADMAP.md
All checks were successful
Build and deploy / Validate source (push) Successful in 14s
Build and deploy / Publish images and notify Portainer (push) Successful in 1m7s
Findings from the 2026-09-18 audit, ordered by block, each with its acceptance
criterion and audit id. Block 0 is closed; the remaining blocks record security,
architecture, scale and maintenance work, including the decisions that need a
product answer before any code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 18:33:28 -03:00
Cauê Faleiros
fc6f708e95 chore: restore a working localhost stack
docker-compose.yml became the production/R2 stack, but LOCAL_SETUP.md still
documented "docker compose up --build" against .env.example, which fails on
missing R2_ENDPOINT, SITE_DOMAIN and KANBAN_DOMAIN, and MinIO was gone.

Add compose.local.yaml: builds from source, MinIO storage, fake providers,
disposable credentials, an app database password distinct from the
administrator one, and published origins in ALLOWED_ORIGINS so browser writes
are not rejected. Correct the documented command and the Kanban login, which
listed a username the email-validated model rejects.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 18:33:27 -03:00
Cauê Faleiros
0f16868f2d fix: stop the application database role reusing the admin password
docker-compose.yml passed POSTGRES_PASSWORD as APP_DB_PASSWORD, so the DML-only
dtf_app role and the owning administrator shared one credential and the
privilege separation bootstrap.py sets up was decorative.

APP_DB_PASSWORD is now its own required variable, and bootstrap refuses to run
when it matches the administrator password, in both the URL and discrete-field
configuration forms.

Deploying this requires APP_DB_PASSWORD to be set in the stack environment
first; db-init rotates the role to it on the same deploy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 18:32:55 -03:00
Cauê Faleiros
67aa6c970c fix: make the by-metre product reachable and its declaration verified
Navigation links carried data-modo-cta and opened loose artwork directly, so
every entry point except the price card bypassed "Arquivo por metro". Dropping
a PNG or JPG from the by-metre editor then switched the order to loose artwork
and repriced it, while the same screen advertised PNG/JPG in its drop zone and
marked that path as the cheaper one.

Replace the guessing with a declaration: the product is the choice. #tipoEnvio
shows ready sheet and loose artwork side by side with both prices, in all four
modes, reversible until a file is attached and locked afterwards. sel() no
longer reassigns modo, so a file can never change product or price on its own.

Verify the declaration instead of trusting it. medirFolha returns dpiFolha, and
an image without the pixels to span the film width at DPI_RECUSA is refused as
a sheet, with the numbers shown and one click to send it as loose artwork.
Previously a small image declared as a sheet was billed by the metre.

Scope pintaCaminhos to #caminhos .cam: its global selector was clearing the new
control. Move the checkout status out of #carr, which the success path hides,
so a paid order still confirms itself to the customer.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 18:32:38 -03:00
Cauê Faleiros
9046ec6db0 fix: restore session creation broken by missing import
local/auth.py used os.environ without importing os, so new_session raised
NameError. Every first visit to /api/session, every registration and every
login returned 500, which left Site checkout, cart recovery and the customer
portal unusable since the R2 stack change.

Read COOKIE_SECURE once as a module constant and share it with local/app.py
instead of resolving the same variable in two places.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 18:32:17 -03:00
Cauê Faleiros
99fd92bdb8 fix: use packed artwork model for standard image uploads
All checks were successful
Build and deploy / Validate source (push) Successful in 10s
Build and deploy / Publish images and notify Portainer (push) Successful in 53s
2026-09-18 16:35:24 -03:00
Cauê Faleiros
483a083a3d fix: open artwork packing from DTF links 2026-09-18 16:00:47 -03:00
Cauê Faleiros
56275f3aba fix: route artwork uploads to packing flow 2026-09-18 15:41:39 -03:00
Cauê Faleiros
4ce28aa0d0 fix: restore ready-sheet preview behavior
All checks were successful
Build and deploy / Validate source (push) Successful in 8s
Build and deploy / Publish images and notify Portainer (push) Successful in 51s
2026-09-18 15:29:50 -03:00
Cauê Faleiros
56d8fde0db fix: show repeated artwork in preview
All checks were successful
Build and deploy / Validate source (push) Successful in 12s
Build and deploy / Publish images and notify Portainer (push) Successful in 1m2s
2026-09-18 15:14:29 -03:00
Cauê Faleiros
2225b70401 fix: preview uploaded artwork by metre
All checks were successful
Build and deploy / Validate source (push) Successful in 8s
Build and deploy / Publish images and notify Portainer (push) Successful in 53s
2026-09-18 14:54:10 -03:00
Cauê Faleiros
88a2a6f060 fix: prevent stale Kanban script after deployment
All checks were successful
Build and deploy / Validate source (push) Successful in 7s
Build and deploy / Publish images and notify Portainer (push) Successful in 53s
2026-09-18 14:27:33 -03:00
Cauê Faleiros
4d707009ba feat: configure Kanban login with optional email
All checks were successful
Build and deploy / Validate source (push) Successful in 10s
Build and deploy / Publish images and notify Portainer (push) Successful in 56s
2026-09-18 14:11:55 -03:00
Cauê Faleiros
d4190ebfeb Revert "feat: configure Kanban login by operator email"
All checks were successful
Build and deploy / Validate source (push) Successful in 12s
Build and deploy / Publish images and notify Portainer (push) Successful in 54s
This reverts commit 508fa03664.
2026-09-18 13:45:08 -03:00
Cauê Faleiros
508fa03664 feat: configure Kanban login by operator email
All checks were successful
Build and deploy / Validate source (push) Successful in 13s
Build and deploy / Publish images and notify Portainer (push) Successful in 1m2s
2026-09-18 12:54:04 -03:00
Cauê Faleiros
9d348ca893 fix: support arbitrary production database passwords
All checks were successful
Build and deploy / Validate source (push) Successful in 7s
Build and deploy / Publish images and notify Portainer (push) Successful in 49s
2026-09-18 12:18:43 -03:00
Cauê Faleiros
fbb620bd85 fix: keep web services up during API rollout 2026-09-18 12:17:34 -03:00
Cauê Faleiros
7605ca9918 fix: start web services in Portainer Swarm
All checks were successful
Build and deploy / Validate source (push) Successful in 7s
Build and deploy / Publish images and notify Portainer (push) Successful in 43s
2026-09-18 12:02:31 -03:00
Cauê Faleiros
e3e37f674d feat: run DTF stack with Cloudflare R2
All checks were successful
Build and deploy / Validate source (push) Successful in 7s
Build and deploy / Publish images and notify Portainer (push) Successful in 44s
2026-09-18 11:51:54 -03:00
Cauê Faleiros
dabb4db2e3 fix: make DTF compose stack Swarm compatible
All checks were successful
Build and deploy / Validate source (push) Successful in 7s
Build and deploy / Publish images and notify Portainer (push) Successful in 39s
2026-09-18 11:27:36 -03:00
Cauê Faleiros
6db107f0fc fix: use Compose deploy resource limits
All checks were successful
Build and deploy / Validate source (push) Successful in 6s
Build and deploy / Publish images and notify Portainer (push) Successful in 42s
2026-09-18 11:20:21 -03:00
Cauê Faleiros
38812884e3 chore: rename root compose file
All checks were successful
Build and deploy / Validate source (push) Successful in 1m40s
Build and deploy / Publish images and notify Portainer (push) Successful in 41s
2026-09-18 11:03:20 -03:00
Cauê Faleiros
6c0c94373a ci: publish DTF images and notify Portainer
All checks were successful
Build and deploy / Validate source (push) Successful in 1m29s
Build and deploy / Publish images and notify Portainer (push) Successful in 1m17s
2026-09-17 17:21:21 -03:00
Cauê Faleiros
26f7d2eb04 chore: remove generated artifacts from repository
Some checks failed
Validate, publish and deploy / validate (push) Successful in 7s
Validate, publish and deploy / publish-and-deploy (push) Failing after 6s
2026-09-15 16:58:26 -03:00
Cauê Faleiros
98c951d374 first commit
Some checks failed
Validate, publish and deploy / validate (push) Successful in 2m2s
Validate, publish and deploy / publish-and-deploy (push) Failing after 8s
2026-09-15 16:42:34 -03:00