deploy/stack.yaml arrived in the first commit and was never deployed. Portainer
runs the repository's docker-compose.yml. Keeping both meant two definitions
drifting apart, with the documentation naming the one nobody used, which is how
the credential question came up at all.
The hardening it offered is narrower than it looks: Docker secrets keep values
out of docker inspect and the Portainer console, but local/secrets.py loads them
into the process environment regardless, and anyone able to read docker inspect
can already read the secret files. With a single Portainer user, the benefit that
remains does not outweigh maintaining a divergent copy.
local/secrets.py stays: inert against the deployed file, and it lets a stack
switch to Docker secrets later without touching code. The preflight and its tests
degrade cleanly when no such stack is present.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The suites connected to localhost:<published port>, which works for a developer
but not on a containerised runner: published ports live in the host's network
namespace, so the runner container gets connection refused.
Run them from inside the stack instead, against the gateway by service name.
SITE_BASE_URL and SITE_HOST_HEADER make that possible without weakening what is
under test: the Host stays "localhost", so the gateway's host check and
TrustedHostMiddleware see exactly what a localhost run produces, and the tests
that deliberately send their own Host still override it.
S3_PUBLIC_ENDPOINT has to agree, because presigned URLs are signed against it
and the signature covers the host, so it cannot be rewritten afterwards. CI
points the whole stack at http://storage:9000 so the URLs it hands out are
reachable by whoever follows them.
The browser suites still need Chrome to reach the stack from the runner, which
the same namespace split prevents. They now check reachability and skip with a
warning instead of failing with a bare connection error; recorded as ROADMAP
5.10, since they are the only coverage for the artwork editor.
Verified both ways: the six suites pass inside the network, and an unchanged
developer localhost run still passes, as do both browser suites locally.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
8000 was Portainer's Edge tunnel, not a stray process. The first attempt at a
fix picked 18080/18081, which are the production dtf-cloud stack's own defaults
in docker-compose.yml: it would have passed only while that stack was down and
collided again the moment it came back.
Use 28080/28081/28000/29000/29001, clear of Portainer (8000, 9443), both
production stack definitions (18080/18081 and 8080/8081) and the usual MinIO
ports. The occupants are listed in the workflow so the next person choosing a
port can see what is taken.
Ephemeral ports would remove the guesswork but do not work here: the published
port is baked into PUBLIC_ORIGIN, ALLOWED_ORIGINS and the CSP when the
containers start, so it has to be known before they run.
Full suite verified on the new block, including the browser end-to-end.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The integration job failed with "Bind for 0.0.0.0:8000 failed: port is already
allocated". The runner shares the host's Docker daemon, so every published port
is claimed on the machine itself, where other services already listen. Port 8000
was the first collision; 8080, 8081, 9000 and 9001 were equally exposed.
MinIO's ports were hardcoded, and S3_PUBLIC_ENDPOINT was pinned to
localhost:9000 independently, so moving storage would have broken the presigned
URLs the browser fetches. Both now derive from STORAGE_PORT and move together.
CI runs on 18080/18081/18000/19000/19001. Local defaults are unchanged.
Verified by running the whole stack and the full suite on exactly those ports,
including the browser end-to-end, which downloads through a presigned URL and so
proves the storage endpoint followed the port.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The production images built on mutable tags with --pull, so the same commit
could produce different bases, and neither Dockerfile upgraded its OS packages
even though the local ones did. The published API image carried 56 HIGH and 3
CRITICAL findings, 15 of them with an upstream fix available.
Pin both bases by digest and upgrade OS packages in the production images. That
removes all 3 CRITICAL and 13 of the 15 fixable findings. The remaining two,
msgpack and setuptools, come from a third-party SBOM; neither package is
importable or listed by pip in the built image, which I confirmed rather than
taking the previous report's word for it.
The web image could not be fixed this way: the official 1.28 line pins
nginx=1.28.3-r1 in /etc/apk/world, so apk upgrade leaves five HIGH findings in
place even though Alpine ships 1.28.3-r7. Moving to nginx:alpine (1.31.6)
clears them completely; 1.29-alpine scans worse, at 37 HIGH. Same uid 101 and
the same template entrypoint, and the local images now use the same pinned
bases so the integration suite exercises what ships. Full suite passes on
nginx 1.31.6, including the browser end-to-end.
With both images at zero CRITICAL, the image scan now blocks on CRITICAL and
reports HIGH, instead of reporting everything. PYTHON_BASE_IMAGE and
NGINX_BASE_IMAGE are wired through to the builds so a base can move forward
without editing the repository, which is what PORTAINER.md already promised.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
PORTAINER.md and SECURITY_REPORT.md described a pipeline that required
regressions, HIGH/CRITICAL secret, misconfiguration and image gates, and stated
that the source preflight stopped this application from publishing. None of it
ran: the workflow built and called the webhook unconditionally.
Add a blocking Trivy secret scan. Verified both ways: a planted AWS key pair,
GitHub token and private key block the job, and the repository passes clean.
Note that Trivy allowlists documented example credentials, so this gate is a
backstop, not permission to commit secrets.
The source preflight now runs on every push and always prints its verdict, but
enforces only when ENFORCE_PRODUCTION_PREFLIGHT is true. Enforcing it today
would block every deployment, because it refuses a release while the payment
and messaging adapters are fake, which is the deliberate state the stack runs
in. Set the variable when real adapters land.
Image vulnerabilities are reported after each build rather than enforced. The
current bases carry 56 HIGH and 3 CRITICAL findings, only 15 of them with an
upstream fix, so failing on them would stop releases without making anything
safer. Pinning digests and triaging the fixable ones is ROADMAP 2.6.
Both documents now carry a table of what gates and what does not, instead of
describing checks that did not exist.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
deploy/stack.yaml passes DATABASE_URL_FILE, AWS_ACCESS_KEY_ID_FILE,
OPERATOR_PASSWORD_FILE and the provider tokens as Swarm secret paths, but the
runtime only ever read the plain names. That stack could not start: the database
URL and R2 credentials were absent, and operator login raised KeyError, so it
returned 500 instead of the intended 503.
local/secrets.py resolves every <NAME>_FILE into <NAME> before configuration is
read, from the API, worker and bootstrap entrypoints. It fails closed on an
unreadable or empty secret and on a name supplied both directly and as a file,
because starting with a credential nobody intended is worse than not starting.
Only one trailing newline is stripped, so a generated password keeps any
whitespace that belongs to it, and no value reaches an error message.
The stack also passed OPERATOR_USER while the Kanban authenticates by email;
it now passes OPERATOR_EMAIL, matching the runtime.
The release gate checked this by searching local/secrets.py for the literal
"DATABASE_URL_FILE", which would pass for any file containing that string. It
now loads the module and makes it resolve every secret the stack declares, and
asserts it fails closed on a missing one. Four marker strings that stopped
matching when R2 support landed are removed rather than left to rot; the two
that still describe real blockers stay, so the gate continues to refuse a
release while payment and messaging adapters are fake.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The pipeline ran py_compile plus four unit tests, then built and called the
Portainer webhook. None of that starts the application, so a missing import in
local/auth.py passed every check and reached production, where it returned 500
on every session, login and registration.
Add an integration job that builds the localhost stack and runs the suites that
already existed but were never executed automatically: smoke, workflow,
security, scanning, retention, runtime security, and the two browser tests.
publish-and-deploy now depends on it, so a failure blocks the deploy instead of
shipping.
Verified by reintroducing the original defect: py_compile and the unit tests
still passed, and smoke_test failed on /session, which would have stopped the
release.
The browser tests need a real Chrome and are skipped with a warning when the
runner has none; installing google-chrome-stable or setting CHROME_BIN makes
them gate too. Every other suite gates unconditionally.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>