Commit Graph

12 Commits

Author SHA1 Message Date
Cauê Faleiros
a87403338d feat: give each Kanban operator their own account
All checks were successful
Build and deploy / Validate source (push) Successful in 5s
Build and deploy / Integration suite on a real stack (push) Successful in 1m17s
Build and deploy / Secret scan and release gate (push) Successful in 5s
Build and deploy / Publish images and notify Portainer (push) Successful in 1m37s
One OPERATOR_EMAIL and OPERATOR_PASSWORD served the whole factory, so every card
movement recorded the same name and the movement history could not answer who
did what. Traceability was one of the things the project set out to provide.

Accounts live in dtf_local.operators, authenticated with the same scrypt hashing
as customer accounts and with comparable work whether or not the account exists,
so absence is not observable by timing. Administration is a CLI in the API
container, like the schema migration: list, add, password, disable, enable.
Passwords are read from the terminal rather than an argument so they stay out of
shell history and the process list, and disabling deletes that operator's open
sessions instead of leaving them valid for the rest of the eight-hour window.

Migration is the part that could hurt: an empty table means 503 and a factory
locked out of its Kanban. OPERATOR_EMAIL and OPERATOR_PASSWORD seed the first
account, and only when that email is absent, so a password changed through the
CLI survives a redeploy carrying a stale environment variable. The first attempt
at this silently did nothing, because db-init receives its own small environment
and had neither variable; both compose files now pass them to it.

Verified against a running stack: bootstrap seeds the existing credential, that
credential still logs in unchanged, a second operator authenticates separately,
wrong passwords and unknown accounts are rejected alike, and disabling revokes
an open session immediately.

Roles are left out on purpose. The separation of duties the meeting described
governs rework authorisation, which this system does not implement, so a role
model would have no consumer to serve.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 14:13:47 -03:00
Cauê Faleiros
da903db32a feat: serve pdf.js from this origin instead of a CDN
All checks were successful
Build and deploy / Validate source (push) Successful in 5s
Build and deploy / Integration suite on a real stack (push) Successful in 1m10s
Build and deploy / Secret scan and release gate (push) Successful in 5s
Build and deploy / Publish images and notify Portainer (push) Successful in 1m43s
The Site pulled pdf.js 3.11.174 from cdnjs with no integrity attribute, and the
policy trusted the whole of cdnjs.cloudflare.com for both script-src and
worker-src. Anything that host served would have executed, and a customer
measuring a PDF sheet depended on it being reachable.

Vendor both files instead of pinning a hash: it removes the dependency rather
than constraining it, and lets the policy name only 'self'. Provenance and
SHA-256 digests are recorded in local/static/vendor/README.md, verified on
download against the SRI digests cdnjs publishes for that release.

cdnjs is now absent from script-src, worker-src and connect-src in both gateway
templates. Workers are 'self' plus blob:, which the Site needs for the worker it
constructs itself.

Verified in a browser against the running stack: pdf.js loads from /vendor/, the
blob worker starts, and a real seven-page PDF parses with no CSP violation. Both
browser suites and the full integration suite pass.

The version is deliberately unchanged. 3.11.174 is old, but its known eval path
is already closed by isEvalSupported:false, and upgrading is an API change that
needs its own testing rather than riding along with this.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 13:56:02 -03:00
Cauê Faleiros
9926d3ab55 chore: remove the unused second stack definition
deploy/stack.yaml arrived in the first commit and was never deployed. Portainer
runs the repository's docker-compose.yml. Keeping both meant two definitions
drifting apart, with the documentation naming the one nobody used, which is how
the credential question came up at all.

The hardening it offered is narrower than it looks: Docker secrets keep values
out of docker inspect and the Portainer console, but local/secrets.py loads them
into the process environment regardless, and anyone able to read docker inspect
can already read the secret files. With a single Portainer user, the benefit that
remains does not outweigh maintaining a divergent copy.

local/secrets.py stays: inert against the deployed file, and it lets a stack
switch to Docker secrets later without touching code. The preflight and its tests
degrade cleanly when no such stack is present.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 13:48:22 -03:00
Cauê Faleiros
e95a42dbed docs: name the stack Portainer actually deploys, and its credential exposure
PORTAINER.md called deploy/stack.yaml the production stack. The deployed file is
docker-compose.yml, which supplies eight credentials as plain environment
variables where stack.yaml uses Docker secrets. That puts the database password,
operator password and the R2 secret key in the container environment, readable
through docker inspect and the Portainer stack editor.

Recorded as ROADMAP 2.12 with the three options rather than changed here:
altering how production receives credentials is not a quiet change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 13:42:12 -03:00
Cauê Faleiros
c1a07a75fa ci: run the integration suites inside the stack's own network
All checks were successful
Build and deploy / Validate source (push) Successful in 4s
Build and deploy / Integration suite on a real stack (push) Successful in 1m12s
Build and deploy / Secret scan and release gate (push) Successful in 5s
Build and deploy / Publish images and notify Portainer (push) Successful in 1m50s
The suites connected to localhost:<published port>, which works for a developer
but not on a containerised runner: published ports live in the host's network
namespace, so the runner container gets connection refused.

Run them from inside the stack instead, against the gateway by service name.
SITE_BASE_URL and SITE_HOST_HEADER make that possible without weakening what is
under test: the Host stays "localhost", so the gateway's host check and
TrustedHostMiddleware see exactly what a localhost run produces, and the tests
that deliberately send their own Host still override it.

S3_PUBLIC_ENDPOINT has to agree, because presigned URLs are signed against it
and the signature covers the host, so it cannot be rewritten afterwards. CI
points the whole stack at http://storage:9000 so the URLs it hands out are
reachable by whoever follows them.

The browser suites still need Chrome to reach the stack from the runner, which
the same namespace split prevents. They now check reachability and skip with a
warning instead of failing with a bare connection error; recorded as ROADMAP
5.10, since they are the only coverage for the artwork editor.

Verified both ways: the six suites pass inside the network, and an unchanged
developer localhost run still passes, as do both browser suites locally.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 13:29:15 -03:00
Cauê Faleiros
3b92813491 fix: recover the customer address behind the host reverse proxy
Some checks failed
Build and deploy / Validate source (push) Successful in 7s
Build and deploy / Integration suite on a real stack (push) Failing after 51s
Build and deploy / Secret scan and release gate (push) Successful in 9s
Build and deploy / Publish images and notify Portainer (push) Has been skipped
The production gateway does not face the internet: nginx-proxy-manager owns
80/443 on the host and proxies to it. So $remote_addr inside the gateway is that
proxy, and overwriting X-Forwarded-For with it discarded the customer address
the proxy had already recorded. Every request would have been attributed to one
internal address, which is exactly the fault 2.1 set out to fix, reintroduced in
production only.

Use real_ip to take the customer address from the proxy's header, trusting only
private networks. A request that reaches the published port directly from the
internet is not trusted, so its header is ignored and $remote_addr stays the
real peer: the anti-spoofing property is kept.

Also downgrade 2.9. TLS is not missing, it is terminated by that proxy. The gap
is that the repository never says so, which would break every session cookie if
the stack moved to a host without one.

Validated with nginx -t against the rendered production configuration.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 13:07:56 -03:00
Cauê Faleiros
dbc9ba4dee docs: record an intermittent browser test failure
Some checks failed
Build and deploy / Validate source (push) Successful in 5s
Build and deploy / Integration suite on a real stack (push) Failing after 1m10s
Build and deploy / Secret scan and release gate (push) Successful in 5s
Build and deploy / Publish images and notify Portainer (push) Has been skipped
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 12:33:58 -03:00
Cauê Faleiros
010c2a162f docs: record base image pinning and the CRITICAL image gate
Some checks failed
Build and deploy / Validate source (push) Successful in 6s
Build and deploy / Integration suite on a real stack (push) Failing after 6s
Build and deploy / Secret scan and release gate (push) Successful in 11s
Build and deploy / Publish images and notify Portainer (push) Has been skipped
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 11:58:46 -03:00
Cauê Faleiros
bbcc8ab100 docs: record the release gates and what they do not cover
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 11:45:23 -03:00
Cauê Faleiros
d2f7b2c03b docs: record secret-file loading and the behavioural release gate
Some checks failed
Build and deploy / Validate source (push) Successful in 1m25s
Build and deploy / Integration suite on a real stack (push) Failing after 7s
Build and deploy / Publish images and notify Portainer (push) Has been skipped
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 11:37:16 -03:00
Cauê Faleiros
9b38c9fe3d docs: record Block 2 rate-limiting work and CI coverage
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 18:55:33 -03:00
Cauê Faleiros
aa0eba3457 docs: track outstanding work in ROADMAP.md
All checks were successful
Build and deploy / Validate source (push) Successful in 14s
Build and deploy / Publish images and notify Portainer (push) Successful in 1m7s
Findings from the 2026-09-18 audit, ordered by block, each with its acceptance
criterion and audit id. Block 0 is closed; the remaining blocks record security,
architecture, scale and maintenance work, including the decisions that need a
product answer before any code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 18:33:28 -03:00