Commit Graph

21 Commits

Author SHA1 Message Date
Cauê Faleiros
86b8199612 fix: create indexes after the tables they name
All checks were successful
Build and deploy / Validate source (push) Successful in 5s
Build and deploy / Integration suite on a real stack (push) Successful in 1m20s
Build and deploy / Secret scan and release gate (push) Successful in 6s
Build and deploy / Publish images and notify Portainer (push) Successful in 1m37s
The 4.1 indexes were added beside the existing uploads_owner, which sits partway
through schema.sql, so CREATE INDEX ... ON dtf_local.order_files ran before that
table was created and bootstrap aborted with UndefinedTable. db-init then
restarted on failure without ever completing, and everything waiting on it timed
out.

Every local run passed because those volumes already had the tables. Only a
clean database exposes it, which is what CI has and my checks did not.

All eleven indexes now sit at the end of the file, after every table, with an
assertion in the change that each indexed table is created before its index.
Verified from docker compose down -v: the stack starts, bootstrap completes,
eleven indexes exist, and the full suite passes.

Recorded as ROADMAP 5.11: nothing exercises the schema against an empty
database, which is the only way this class of fault appears.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 15:00:10 -03:00
Cauê Faleiros
543a9a9fb4 perf: index the real queries, bound the board, and correct what the Site promises
Some checks failed
Build and deploy / Validate source (push) Successful in 5s
Build and deploy / Integration suite on a real stack (push) Failing after 10m30s
Build and deploy / Secret scan and release gate (push) Successful in 5s
Build and deploy / Publish images and notify Portainer (push) Has been skipped
Five items that needed no decisions.

Indexes: the schema indexed only uploads(owner), so the worker's once-a-second
outbox poll scanned a table that only grows, and every per-customer and
per-order lookup did the same. Ten indexes now follow queries the application
actually issues, and no more, since each one is paid for on every write. The
outbox and live uploads use partial indexes so they stay the size of the backlog
rather than of all history. Confirmed against the database that the planner
chooses them.

Board: /api/operator/board returned every order ever created. Finished orders
are terminal, so they were pure growth. It now returns everything still in
progress however old, plus a window of recent finished ones and the true
finished total, and the Kanban column says "50 de 213" rather than letting the
count read as an all-time figure. An operator cannot lose a card they could act
on.

Dependencies: the root requirements.txt was the prototype's, pinned by wildcard,
listing packages this system does not use, next to the hash-locked lock file.
Deleted. pip was pinned as a runtime dependency, which installed a package
manager into the read-only production image; nothing depended on it, so it is
gone from both the direct list and the lock, and the base image's pip performs
the hash-enforced install.

Retention copy: the Site told customers their artwork was kept 90 days with 12
months of history, and invited them to reorder without uploading again. Files
are kept 30 days. The copy now matches the policy and drops the promise the
system cannot keep.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 14:29:07 -03:00
Cauê Faleiros
a87403338d feat: give each Kanban operator their own account
All checks were successful
Build and deploy / Validate source (push) Successful in 5s
Build and deploy / Integration suite on a real stack (push) Successful in 1m17s
Build and deploy / Secret scan and release gate (push) Successful in 5s
Build and deploy / Publish images and notify Portainer (push) Successful in 1m37s
One OPERATOR_EMAIL and OPERATOR_PASSWORD served the whole factory, so every card
movement recorded the same name and the movement history could not answer who
did what. Traceability was one of the things the project set out to provide.

Accounts live in dtf_local.operators, authenticated with the same scrypt hashing
as customer accounts and with comparable work whether or not the account exists,
so absence is not observable by timing. Administration is a CLI in the API
container, like the schema migration: list, add, password, disable, enable.
Passwords are read from the terminal rather than an argument so they stay out of
shell history and the process list, and disabling deletes that operator's open
sessions instead of leaving them valid for the rest of the eight-hour window.

Migration is the part that could hurt: an empty table means 503 and a factory
locked out of its Kanban. OPERATOR_EMAIL and OPERATOR_PASSWORD seed the first
account, and only when that email is absent, so a password changed through the
CLI survives a redeploy carrying a stale environment variable. The first attempt
at this silently did nothing, because db-init receives its own small environment
and had neither variable; both compose files now pass them to it.

Verified against a running stack: bootstrap seeds the existing credential, that
credential still logs in unchanged, a second operator authenticates separately,
wrong passwords and unknown accounts are rejected alike, and disabling revokes
an open session immediately.

Roles are left out on purpose. The separation of duties the meeting described
governs rework authorisation, which this system does not implement, so a role
model would have no consumer to serve.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 14:13:47 -03:00
Cauê Faleiros
da903db32a feat: serve pdf.js from this origin instead of a CDN
All checks were successful
Build and deploy / Validate source (push) Successful in 5s
Build and deploy / Integration suite on a real stack (push) Successful in 1m10s
Build and deploy / Secret scan and release gate (push) Successful in 5s
Build and deploy / Publish images and notify Portainer (push) Successful in 1m43s
The Site pulled pdf.js 3.11.174 from cdnjs with no integrity attribute, and the
policy trusted the whole of cdnjs.cloudflare.com for both script-src and
worker-src. Anything that host served would have executed, and a customer
measuring a PDF sheet depended on it being reachable.

Vendor both files instead of pinning a hash: it removes the dependency rather
than constraining it, and lets the policy name only 'self'. Provenance and
SHA-256 digests are recorded in local/static/vendor/README.md, verified on
download against the SRI digests cdnjs publishes for that release.

cdnjs is now absent from script-src, worker-src and connect-src in both gateway
templates. Workers are 'self' plus blob:, which the Site needs for the worker it
constructs itself.

Verified in a browser against the running stack: pdf.js loads from /vendor/, the
blob worker starts, and a real seven-page PDF parses with no CSP violation. Both
browser suites and the full integration suite pass.

The version is deliberately unchanged. 3.11.174 is old, but its known eval path
is already closed by isEvalSupported:false, and upgrading is an API change that
needs its own testing rather than riding along with this.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 13:56:02 -03:00
Cauê Faleiros
9926d3ab55 chore: remove the unused second stack definition
deploy/stack.yaml arrived in the first commit and was never deployed. Portainer
runs the repository's docker-compose.yml. Keeping both meant two definitions
drifting apart, with the documentation naming the one nobody used, which is how
the credential question came up at all.

The hardening it offered is narrower than it looks: Docker secrets keep values
out of docker inspect and the Portainer console, but local/secrets.py loads them
into the process environment regardless, and anyone able to read docker inspect
can already read the secret files. With a single Portainer user, the benefit that
remains does not outweigh maintaining a divergent copy.

local/secrets.py stays: inert against the deployed file, and it lets a stack
switch to Docker secrets later without touching code. The preflight and its tests
degrade cleanly when no such stack is present.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 13:48:22 -03:00
Cauê Faleiros
c1a07a75fa ci: run the integration suites inside the stack's own network
All checks were successful
Build and deploy / Validate source (push) Successful in 4s
Build and deploy / Integration suite on a real stack (push) Successful in 1m12s
Build and deploy / Secret scan and release gate (push) Successful in 5s
Build and deploy / Publish images and notify Portainer (push) Successful in 1m50s
The suites connected to localhost:<published port>, which works for a developer
but not on a containerised runner: published ports live in the host's network
namespace, so the runner container gets connection refused.

Run them from inside the stack instead, against the gateway by service name.
SITE_BASE_URL and SITE_HOST_HEADER make that possible without weakening what is
under test: the Host stays "localhost", so the gateway's host check and
TrustedHostMiddleware see exactly what a localhost run produces, and the tests
that deliberately send their own Host still override it.

S3_PUBLIC_ENDPOINT has to agree, because presigned URLs are signed against it
and the signature covers the host, so it cannot be rewritten afterwards. CI
points the whole stack at http://storage:9000 so the URLs it hands out are
reachable by whoever follows them.

The browser suites still need Chrome to reach the stack from the runner, which
the same namespace split prevents. They now check reachability and skip with a
warning instead of failing with a bare connection error; recorded as ROADMAP
5.10, since they are the only coverage for the artwork editor.

Verified both ways: the six suites pass inside the network, and an unchanged
developer localhost run still passes, as do both browser suites locally.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 13:29:15 -03:00
Cauê Faleiros
6a50e6db4d fix: bake scanner and storage config into images instead of bind-mounting
Some checks failed
Build and deploy / Validate source (push) Successful in 6s
Build and deploy / Integration suite on a real stack (push) Failing after 54s
Build and deploy / Secret scan and release gate (push) Successful in 5s
Build and deploy / Publish images and notify Portainer (push) Has been skipped
The integration job failed starting the scanner:

  error mounting ".../local/clamd.conf" to rootfs at "/etc/clamav/clamd.conf":
  not a directory

The files are in the repository, so this was not a missing checkout. A
containerised CI runner shares the host's Docker daemon, so "./local/clamd.conf"
resolves to a workspace path that exists inside the runner but not on the host
where the daemon creates the mount. The daemon makes an empty directory there
and the container cannot start. Only the bind-mounting services were affected,
which is why PostgreSQL and MinIO came up first.

Build the scanner and storage-init images with their configuration copied in, so
compose.local.yaml no longer bind-mounts anything from the host and works
regardless of how the runner reaches the daemon. Both bases stay overridable
through CLAMAV_IMAGE and MINIO_IMAGE.

The production stack is unaffected: it ships clamd.conf as a Swarm config, which
the manager reads at deploy time.

Verified from a clean slate: the stack starts, the scanner runs the baked
configuration, storage provisioning runs from the baked script, and the full
suite passes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 12:51:08 -03:00
Cauê Faleiros
4c9fa2436e build: pin base image digests and clear every fixable image finding
The production images built on mutable tags with --pull, so the same commit
could produce different bases, and neither Dockerfile upgraded its OS packages
even though the local ones did. The published API image carried 56 HIGH and 3
CRITICAL findings, 15 of them with an upstream fix available.

Pin both bases by digest and upgrade OS packages in the production images. That
removes all 3 CRITICAL and 13 of the 15 fixable findings. The remaining two,
msgpack and setuptools, come from a third-party SBOM; neither package is
importable or listed by pip in the built image, which I confirmed rather than
taking the previous report's word for it.

The web image could not be fixed this way: the official 1.28 line pins
nginx=1.28.3-r1 in /etc/apk/world, so apk upgrade leaves five HIGH findings in
place even though Alpine ships 1.28.3-r7. Moving to nginx:alpine (1.31.6)
clears them completely; 1.29-alpine scans worse, at 37 HIGH. Same uid 101 and
the same template entrypoint, and the local images now use the same pinned
bases so the integration suite exercises what ships. Full suite passes on
nginx 1.31.6, including the browser end-to-end.

With both images at zero CRITICAL, the image scan now blocks on CRITICAL and
reports HIGH, instead of reporting everything. PYTHON_BASE_IMAGE and
NGINX_BASE_IMAGE are wired through to the builds so a base can move forward
without editing the repository, which is what PORTAINER.md already promised.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 11:58:32 -03:00
Cauê Faleiros
341f154c36 feat: load Docker secret files so the production stack can boot
deploy/stack.yaml passes DATABASE_URL_FILE, AWS_ACCESS_KEY_ID_FILE,
OPERATOR_PASSWORD_FILE and the provider tokens as Swarm secret paths, but the
runtime only ever read the plain names. That stack could not start: the database
URL and R2 credentials were absent, and operator login raised KeyError, so it
returned 500 instead of the intended 503.

local/secrets.py resolves every <NAME>_FILE into <NAME> before configuration is
read, from the API, worker and bootstrap entrypoints. It fails closed on an
unreadable or empty secret and on a name supplied both directly and as a file,
because starting with a credential nobody intended is worse than not starting.
Only one trailing newline is stripped, so a generated password keeps any
whitespace that belongs to it, and no value reaches an error message.

The stack also passed OPERATOR_USER while the Kanban authenticates by email;
it now passes OPERATOR_EMAIL, matching the runtime.

The release gate checked this by searching local/secrets.py for the literal
"DATABASE_URL_FILE", which would pass for any file containing that string. It
now loads the module and makes it resolve every secret the stack declares, and
asserts it fails closed on a missing one. Four marker strings that stopped
matching when R2 support landed are removed rather than left to rot; the two
that still describe real blockers stay, so the gate continues to refuse a
release while payment and messaging adapters are fake.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 11:36:59 -03:00
Cauê Faleiros
52ae13c9a1 fix: rate-limit and audit by real client address
uvicorn does not trust forwarded headers from a peer outside
forwarded_allow_ips, so request.client.host was the web gateway for every
request. The auth-source bucket therefore counted all customers together:
60 failed logins from one attacker locked out everyone. Security events
recorded the gateway address, which made the audit trail useless for
attribution.

The gateway now overwrites X-Forwarded-For with the peer address it observed
instead of appending to whatever the client sent, so the header carries one
value the client cannot choose, and client_ip() resolves it with a fallback to
the connection peer.

The guest-session limiter was keyed on the environment name, making it one
global bucket of 120 per 15 minutes: roughly eight new visitors a minute for
the whole site before legitimate traffic started receiving 429. It is now per
source, and the ceiling is deliberately generous because offices and mobile
carriers put many real customers behind a single address.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 18:48:31 -03:00
Cauê Faleiros
0f16868f2d fix: stop the application database role reusing the admin password
docker-compose.yml passed POSTGRES_PASSWORD as APP_DB_PASSWORD, so the DML-only
dtf_app role and the owning administrator shared one credential and the
privilege separation bootstrap.py sets up was decorative.

APP_DB_PASSWORD is now its own required variable, and bootstrap refuses to run
when it matches the administrator password, in both the URL and discrete-field
configuration forms.

Deploying this requires APP_DB_PASSWORD to be set in the stack environment
first; db-init rotates the role to it on the same deploy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 18:32:55 -03:00
Cauê Faleiros
67aa6c970c fix: make the by-metre product reachable and its declaration verified
Navigation links carried data-modo-cta and opened loose artwork directly, so
every entry point except the price card bypassed "Arquivo por metro". Dropping
a PNG or JPG from the by-metre editor then switched the order to loose artwork
and repriced it, while the same screen advertised PNG/JPG in its drop zone and
marked that path as the cheaper one.

Replace the guessing with a declaration: the product is the choice. #tipoEnvio
shows ready sheet and loose artwork side by side with both prices, in all four
modes, reversible until a file is attached and locked afterwards. sel() no
longer reassigns modo, so a file can never change product or price on its own.

Verify the declaration instead of trusting it. medirFolha returns dpiFolha, and
an image without the pixels to span the film width at DPI_RECUSA is refused as
a sheet, with the numbers shown and one click to send it as loose artwork.
Previously a small image declared as a sheet was billed by the metre.

Scope pintaCaminhos to #caminhos .cam: its global selector was clearing the new
control. Move the checkout status out of #carr, which the success path hides,
so a paid order still confirms itself to the customer.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 18:32:38 -03:00
Cauê Faleiros
9046ec6db0 fix: restore session creation broken by missing import
local/auth.py used os.environ without importing os, so new_session raised
NameError. Every first visit to /api/session, every registration and every
login returned 500, which left Site checkout, cart recovery and the customer
portal unusable since the R2 stack change.

Read COOKIE_SECURE once as a module constant and share it with local/app.py
instead of resolving the same variable in two places.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 18:32:17 -03:00
Cauê Faleiros
99fd92bdb8 fix: use packed artwork model for standard image uploads
All checks were successful
Build and deploy / Validate source (push) Successful in 10s
Build and deploy / Publish images and notify Portainer (push) Successful in 53s
2026-09-18 16:35:24 -03:00
Cauê Faleiros
88a2a6f060 fix: prevent stale Kanban script after deployment
All checks were successful
Build and deploy / Validate source (push) Successful in 7s
Build and deploy / Publish images and notify Portainer (push) Successful in 53s
2026-09-18 14:27:33 -03:00
Cauê Faleiros
4d707009ba feat: configure Kanban login with optional email
All checks were successful
Build and deploy / Validate source (push) Successful in 10s
Build and deploy / Publish images and notify Portainer (push) Successful in 56s
2026-09-18 14:11:55 -03:00
Cauê Faleiros
d4190ebfeb Revert "feat: configure Kanban login by operator email"
All checks were successful
Build and deploy / Validate source (push) Successful in 12s
Build and deploy / Publish images and notify Portainer (push) Successful in 54s
This reverts commit 508fa03664.
2026-09-18 13:45:08 -03:00
Cauê Faleiros
508fa03664 feat: configure Kanban login by operator email
All checks were successful
Build and deploy / Validate source (push) Successful in 13s
Build and deploy / Publish images and notify Portainer (push) Successful in 1m2s
2026-09-18 12:54:04 -03:00
Cauê Faleiros
9d348ca893 fix: support arbitrary production database passwords
All checks were successful
Build and deploy / Validate source (push) Successful in 7s
Build and deploy / Publish images and notify Portainer (push) Successful in 49s
2026-09-18 12:18:43 -03:00
Cauê Faleiros
e3e37f674d feat: run DTF stack with Cloudflare R2
All checks were successful
Build and deploy / Validate source (push) Successful in 7s
Build and deploy / Publish images and notify Portainer (push) Successful in 44s
2026-09-18 11:51:54 -03:00
Cauê Faleiros
98c951d374 first commit
Some checks failed
Validate, publish and deploy / validate (push) Successful in 2m2s
Validate, publish and deploy / publish-and-deploy (push) Failing after 8s
2026-09-15 16:42:34 -03:00