The Site's page sat at the repository root while its scripts lived in
local/static, a split with no reason behind it. They are together in web/ now,
with the page as index.html, which is also what the image serves.
app.py held the adapters, the configuration, the shared query helpers and
nineteen routes; customer.py held fourteen more but could not import from it
without a cycle, so it was wired by passing nine callables into install_routes.
Configuration and shared helpers move to local/runtime.py, the rules for
attaching artwork to an order move to local/artwork.py where a customer
correction and an operator final-file set can share them, and the routes become
seven routers under local/api. app.py is 48 lines that create the application,
apply the middleware and include them. Routers import downwards only.
Three faults came out of the extraction and are worth recording, because each
passed a check that looked sufficient. ast reports a function's line at the def,
so every decorator on the line above fell outside the extracted range: twelve
routes and the security middleware were defined but never registered, and the
files still imported and parsed cleanly. Names the old closure renamed on the
way in, and a Jsonb import, were missing in three modules. A name-resolution
pass over every new module found those; the route count matching the original
exactly, 32, is what confirmed the first.
The release gate's marker for the fake payment adapter pointed at app.py and the
adapter moved to runtime.py, so the gate passed while the condition it guards was
unchanged. That is the same silent decay 2.5 set out to fix. A test now asserts
every marker still matches something in its file, so the next move fails loudly.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The previous commit used git add -A and swept in four binaries that were
deliberately untracked: the week-1 client report as .docx and .pdf, a duplicate
of it under output/documents, and imagem-teste.jpg, an input dropped in to test
with. None of them are the repository's to version. They are untracked here,
left on disk, and covered by .gitignore so the mistake cannot repeat.
Three files had no sensible home. The meeting notes sat at the repository root
under a 78-character name with spaces and accents, the roadmap generator lived
in tmp/ — a directory otherwise ignored as scratch — and its output in output/,
which is otherwise generated evidence. They are now docs/reuniao-2026-09-09-
anotacoes.pdf, tools/generate_dtf_report.py and docs/roadmap-cliente.pdf, with
CONTEXT.md and ROADMAP.md updated to match and a docs/README.md saying what each
document is for.
.gitignore no longer needs four rules to keep one generator out of an ignored
directory; tmp/ is scratch again.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
dtf-site.html held commercial rules, the nesting engine, PDF analysis, the cart
and every handler in a single inline script, 42% of the runtime code in one
file, and the money logic lived in the middle of it.
It is now nine files under local/static, cut at the section markers the original
author left, so no function was split across a boundary: config, product modes,
upload, sheet analysis, PDF, quality, packing, cart, flow. They load as classic
scripts in the original order and share one global scope, so evaluation is
exactly what it was; the extraction was checked byte-identical against the
original before the tags replaced it. dtf-site.html is 1,394 lines of markup and
style.
With no inline script left anywhere, the policy no longer needs a hash
allowlist: script-src is now 'self' alone, which is stronger than what it
replaced and cannot drift as the page changes.
Three things depended on the old shape and were updated rather than worked
around. The pricing parity test read the ladder out of the HTML and now reads it
from site-config.js, still proving the server agrees with what the customer is
shown. The isolated artwork test served four hardcoded script paths and now
serves any script that resolves inside local/static, so the next file added does
not silently 404. The CSP assertion checked the whole policy for 'unsafe-inline'
and now checks the script-src directive alone, since style-src legitimately
carries it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
portal/, kanban/ and agente/ were 2,034 lines implementing the original
Tiny-first model: token upload links, a second SQLite Kanban, a factory agent.
Nothing imported or started any of it, and several endpoints took the acting
user from the request body with no authentication at all. Their real cost was
that a reader arriving at this repository found two Kanbans and two portals and
had to work out which one was real. The root schema.sql and .env.exemplo went
with them: both code paths load local/schema.sql, and having .env.exemplo beside
.env.example differing by one letter was a trap rather than a convenience.
The documents describing that model are archived rather than deleted. They
record decisions and reasoning the current documents do not repeat, so they are
worth keeping as background, with a header saying plainly that they are not
instructions.
README.md keeps its business case — the capacity figures and the cost argument
are still the reason this project exists — but now states where the prototype
documentation begins and that the code it describes is gone.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The 4.1 indexes were added beside the existing uploads_owner, which sits partway
through schema.sql, so CREATE INDEX ... ON dtf_local.order_files ran before that
table was created and bootstrap aborted with UndefinedTable. db-init then
restarted on failure without ever completing, and everything waiting on it timed
out.
Every local run passed because those volumes already had the tables. Only a
clean database exposes it, which is what CI has and my checks did not.
All eleven indexes now sit at the end of the file, after every table, with an
assertion in the change that each indexed table is created before its index.
Verified from docker compose down -v: the stack starts, bootstrap completes,
eleven indexes exist, and the full suite passes.
Recorded as ROADMAP 5.11: nothing exercises the schema against an empty
database, which is the only way this class of fault appears.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Five items that needed no decisions.
Indexes: the schema indexed only uploads(owner), so the worker's once-a-second
outbox poll scanned a table that only grows, and every per-customer and
per-order lookup did the same. Ten indexes now follow queries the application
actually issues, and no more, since each one is paid for on every write. The
outbox and live uploads use partial indexes so they stay the size of the backlog
rather than of all history. Confirmed against the database that the planner
chooses them.
Board: /api/operator/board returned every order ever created. Finished orders
are terminal, so they were pure growth. It now returns everything still in
progress however old, plus a window of recent finished ones and the true
finished total, and the Kanban column says "50 de 213" rather than letting the
count read as an all-time figure. An operator cannot lose a card they could act
on.
Dependencies: the root requirements.txt was the prototype's, pinned by wildcard,
listing packages this system does not use, next to the hash-locked lock file.
Deleted. pip was pinned as a runtime dependency, which installed a package
manager into the read-only production image; nothing depended on it, so it is
gone from both the direct list and the lock, and the base image's pip performs
the hash-enforced install.
Retention copy: the Site told customers their artwork was kept 90 days with 12
months of history, and invited them to reorder without uploading again. Files
are kept 30 days. The copy now matches the policy and drops the promise the
system cannot keep.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
One OPERATOR_EMAIL and OPERATOR_PASSWORD served the whole factory, so every card
movement recorded the same name and the movement history could not answer who
did what. Traceability was one of the things the project set out to provide.
Accounts live in dtf_local.operators, authenticated with the same scrypt hashing
as customer accounts and with comparable work whether or not the account exists,
so absence is not observable by timing. Administration is a CLI in the API
container, like the schema migration: list, add, password, disable, enable.
Passwords are read from the terminal rather than an argument so they stay out of
shell history and the process list, and disabling deletes that operator's open
sessions instead of leaving them valid for the rest of the eight-hour window.
Migration is the part that could hurt: an empty table means 503 and a factory
locked out of its Kanban. OPERATOR_EMAIL and OPERATOR_PASSWORD seed the first
account, and only when that email is absent, so a password changed through the
CLI survives a redeploy carrying a stale environment variable. The first attempt
at this silently did nothing, because db-init receives its own small environment
and had neither variable; both compose files now pass them to it.
Verified against a running stack: bootstrap seeds the existing credential, that
credential still logs in unchanged, a second operator authenticates separately,
wrong passwords and unknown accounts are rejected alike, and disabling revokes
an open session immediately.
Roles are left out on purpose. The separation of duties the meeting described
governs rework authorisation, which this system does not implement, so a role
model would have no consumer to serve.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The Site pulled pdf.js 3.11.174 from cdnjs with no integrity attribute, and the
policy trusted the whole of cdnjs.cloudflare.com for both script-src and
worker-src. Anything that host served would have executed, and a customer
measuring a PDF sheet depended on it being reachable.
Vendor both files instead of pinning a hash: it removes the dependency rather
than constraining it, and lets the policy name only 'self'. Provenance and
SHA-256 digests are recorded in local/static/vendor/README.md, verified on
download against the SRI digests cdnjs publishes for that release.
cdnjs is now absent from script-src, worker-src and connect-src in both gateway
templates. Workers are 'self' plus blob:, which the Site needs for the worker it
constructs itself.
Verified in a browser against the running stack: pdf.js loads from /vendor/, the
blob worker starts, and a real seven-page PDF parses with no CSP violation. Both
browser suites and the full integration suite pass.
The version is deliberately unchanged. 3.11.174 is old, but its known eval path
is already closed by isEvalSupported:false, and upgrading is an API change that
needs its own testing rather than riding along with this.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
deploy/stack.yaml arrived in the first commit and was never deployed. Portainer
runs the repository's docker-compose.yml. Keeping both meant two definitions
drifting apart, with the documentation naming the one nobody used, which is how
the credential question came up at all.
The hardening it offered is narrower than it looks: Docker secrets keep values
out of docker inspect and the Portainer console, but local/secrets.py loads them
into the process environment regardless, and anyone able to read docker inspect
can already read the secret files. With a single Portainer user, the benefit that
remains does not outweigh maintaining a divergent copy.
local/secrets.py stays: inert against the deployed file, and it lets a stack
switch to Docker secrets later without touching code. The preflight and its tests
degrade cleanly when no such stack is present.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
PORTAINER.md called deploy/stack.yaml the production stack. The deployed file is
docker-compose.yml, which supplies eight credentials as plain environment
variables where stack.yaml uses Docker secrets. That puts the database password,
operator password and the R2 secret key in the container environment, readable
through docker inspect and the Portainer stack editor.
Recorded as ROADMAP 2.12 with the three options rather than changed here:
altering how production receives credentials is not a quiet change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The suites connected to localhost:<published port>, which works for a developer
but not on a containerised runner: published ports live in the host's network
namespace, so the runner container gets connection refused.
Run them from inside the stack instead, against the gateway by service name.
SITE_BASE_URL and SITE_HOST_HEADER make that possible without weakening what is
under test: the Host stays "localhost", so the gateway's host check and
TrustedHostMiddleware see exactly what a localhost run produces, and the tests
that deliberately send their own Host still override it.
S3_PUBLIC_ENDPOINT has to agree, because presigned URLs are signed against it
and the signature covers the host, so it cannot be rewritten afterwards. CI
points the whole stack at http://storage:9000 so the URLs it hands out are
reachable by whoever follows them.
The browser suites still need Chrome to reach the stack from the runner, which
the same namespace split prevents. They now check reachability and skip with a
warning instead of failing with a bare connection error; recorded as ROADMAP
5.10, since they are the only coverage for the artwork editor.
Verified both ways: the six suites pass inside the network, and an unchanged
developer localhost run still passes, as do both browser suites locally.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The production gateway does not face the internet: nginx-proxy-manager owns
80/443 on the host and proxies to it. So $remote_addr inside the gateway is that
proxy, and overwriting X-Forwarded-For with it discarded the customer address
the proxy had already recorded. Every request would have been attributed to one
internal address, which is exactly the fault 2.1 set out to fix, reintroduced in
production only.
Use real_ip to take the customer address from the proxy's header, trusting only
private networks. A request that reaches the published port directly from the
internet is not trusted, so its header is ignored and $remote_addr stays the
real peer: the anti-spoofing property is kept.
Also downgrade 2.9. TLS is not missing, it is terminated by that proxy. The gap
is that the repository never says so, which would break every session cookie if
the stack moved to a host without one.
Validated with nginx -t against the rendered production configuration.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
8000 was Portainer's Edge tunnel, not a stray process. The first attempt at a
fix picked 18080/18081, which are the production dtf-cloud stack's own defaults
in docker-compose.yml: it would have passed only while that stack was down and
collided again the moment it came back.
Use 28080/28081/28000/29000/29001, clear of Portainer (8000, 9443), both
production stack definitions (18080/18081 and 8080/8081) and the usual MinIO
ports. The occupants are listed in the workflow so the next person choosing a
port can see what is taken.
Ephemeral ports would remove the guesswork but do not work here: the published
port is baked into PUBLIC_ORIGIN, ALLOWED_ORIGINS and the CSP when the
containers start, so it has to be known before they run.
Full suite verified on the new block, including the browser end-to-end.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The integration job failed with "Bind for 0.0.0.0:8000 failed: port is already
allocated". The runner shares the host's Docker daemon, so every published port
is claimed on the machine itself, where other services already listen. Port 8000
was the first collision; 8080, 8081, 9000 and 9001 were equally exposed.
MinIO's ports were hardcoded, and S3_PUBLIC_ENDPOINT was pinned to
localhost:9000 independently, so moving storage would have broken the presigned
URLs the browser fetches. Both now derive from STORAGE_PORT and move together.
CI runs on 18080/18081/18000/19000/19001. Local defaults are unchanged.
Verified by running the whole stack and the full suite on exactly those ports,
including the browser end-to-end, which downloads through a presigned URL and so
proves the storage endpoint followed the port.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The integration job failed starting the scanner:
error mounting ".../local/clamd.conf" to rootfs at "/etc/clamav/clamd.conf":
not a directory
The files are in the repository, so this was not a missing checkout. A
containerised CI runner shares the host's Docker daemon, so "./local/clamd.conf"
resolves to a workspace path that exists inside the runner but not on the host
where the daemon creates the mount. The daemon makes an empty directory there
and the container cannot start. Only the bind-mounting services were affected,
which is why PostgreSQL and MinIO came up first.
Build the scanner and storage-init images with their configuration copied in, so
compose.local.yaml no longer bind-mounts anything from the host and works
regardless of how the runner reaches the daemon. Both bases stay overridable
through CLAMAV_IMAGE and MINIO_IMAGE.
The production stack is unaffected: it ships clamd.conf as a Swarm config, which
the manager reads at deploy time.
Verified from a clean slate: the stack starts, the scanner runs the baked
configuration, storage provisioning runs from the baked script, and the full
suite passes.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The integration job failed on the runner with "pull access denied for
minio/minio ... may require 'docker login'". Docker Hub now refuses anonymous
pulls of minio/minio: an unauthenticated manifest request returns 401
UNAUTHORIZED, while library/postgres returns 200, which is why only MinIO
failed. It worked locally only because this machine is logged in to Docker Hub.
quay.io serves the same release anonymously, and it is the same image: both
registries resolve to image ID sha256:a1ea29fa2835. MINIO_IMAGE overrides it for
anyone mirroring into their own registry.
Verified by deleting the Docker Hub copy locally and starting the stack from
quay alone, then running the full suite against it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The production images built on mutable tags with --pull, so the same commit
could produce different bases, and neither Dockerfile upgraded its OS packages
even though the local ones did. The published API image carried 56 HIGH and 3
CRITICAL findings, 15 of them with an upstream fix available.
Pin both bases by digest and upgrade OS packages in the production images. That
removes all 3 CRITICAL and 13 of the 15 fixable findings. The remaining two,
msgpack and setuptools, come from a third-party SBOM; neither package is
importable or listed by pip in the built image, which I confirmed rather than
taking the previous report's word for it.
The web image could not be fixed this way: the official 1.28 line pins
nginx=1.28.3-r1 in /etc/apk/world, so apk upgrade leaves five HIGH findings in
place even though Alpine ships 1.28.3-r7. Moving to nginx:alpine (1.31.6)
clears them completely; 1.29-alpine scans worse, at 37 HIGH. Same uid 101 and
the same template entrypoint, and the local images now use the same pinned
bases so the integration suite exercises what ships. Full suite passes on
nginx 1.31.6, including the browser end-to-end.
With both images at zero CRITICAL, the image scan now blocks on CRITICAL and
reports HIGH, instead of reporting everything. PYTHON_BASE_IMAGE and
NGINX_BASE_IMAGE are wired through to the builds so a base can move forward
without editing the repository, which is what PORTAINER.md already promised.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
PORTAINER.md and SECURITY_REPORT.md described a pipeline that required
regressions, HIGH/CRITICAL secret, misconfiguration and image gates, and stated
that the source preflight stopped this application from publishing. None of it
ran: the workflow built and called the webhook unconditionally.
Add a blocking Trivy secret scan. Verified both ways: a planted AWS key pair,
GitHub token and private key block the job, and the repository passes clean.
Note that Trivy allowlists documented example credentials, so this gate is a
backstop, not permission to commit secrets.
The source preflight now runs on every push and always prints its verdict, but
enforces only when ENFORCE_PRODUCTION_PREFLIGHT is true. Enforcing it today
would block every deployment, because it refuses a release while the payment
and messaging adapters are fake, which is the deliberate state the stack runs
in. Set the variable when real adapters land.
Image vulnerabilities are reported after each build rather than enforced. The
current bases carry 56 HIGH and 3 CRITICAL findings, only 15 of them with an
upstream fix, so failing on them would stop releases without making anything
safer. Pinning digests and triaging the fixable ones is ROADMAP 2.6.
Both documents now carry a table of what gates and what does not, instead of
describing checks that did not exist.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
deploy/stack.yaml passes DATABASE_URL_FILE, AWS_ACCESS_KEY_ID_FILE,
OPERATOR_PASSWORD_FILE and the provider tokens as Swarm secret paths, but the
runtime only ever read the plain names. That stack could not start: the database
URL and R2 credentials were absent, and operator login raised KeyError, so it
returned 500 instead of the intended 503.
local/secrets.py resolves every <NAME>_FILE into <NAME> before configuration is
read, from the API, worker and bootstrap entrypoints. It fails closed on an
unreadable or empty secret and on a name supplied both directly and as a file,
because starting with a credential nobody intended is worse than not starting.
Only one trailing newline is stripped, so a generated password keeps any
whitespace that belongs to it, and no value reaches an error message.
The stack also passed OPERATOR_USER while the Kanban authenticates by email;
it now passes OPERATOR_EMAIL, matching the runtime.
The release gate checked this by searching local/secrets.py for the literal
"DATABASE_URL_FILE", which would pass for any file containing that string. It
now loads the module and makes it resolve every secret the stack declares, and
asserts it fails closed on a missing one. Four marker strings that stopped
matching when R2 support landed are removed rather than left to rot; the two
that still describe real blockers stay, so the gate continues to refuse a
release while payment and messaging adapters are fake.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The pipeline ran py_compile plus four unit tests, then built and called the
Portainer webhook. None of that starts the application, so a missing import in
local/auth.py passed every check and reached production, where it returned 500
on every session, login and registration.
Add an integration job that builds the localhost stack and runs the suites that
already existed but were never executed automatically: smoke, workflow,
security, scanning, retention, runtime security, and the two browser tests.
publish-and-deploy now depends on it, so a failure blocks the deploy instead of
shipping.
Verified by reintroducing the original defect: py_compile and the unit tests
still passed, and smoke_test failed on /session, which would have stopped the
release.
The browser tests need a real Chrome and are skipped with a warning when the
runner has none; installing google-chrome-stable or setting CHROME_BIN makes
them gate too. Every other suite gates unconditionally.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
uvicorn does not trust forwarded headers from a peer outside
forwarded_allow_ips, so request.client.host was the web gateway for every
request. The auth-source bucket therefore counted all customers together:
60 failed logins from one attacker locked out everyone. Security events
recorded the gateway address, which made the audit trail useless for
attribution.
The gateway now overwrites X-Forwarded-For with the peer address it observed
instead of appending to whatever the client sent, so the header carries one
value the client cannot choose, and client_ip() resolves it with a fallback to
the connection peer.
The guest-session limiter was keyed on the environment name, making it one
global bucket of 120 per 15 minutes: roughly eight new visitors a minute for
the whole site before legitimate traffic started receiving 429. It is now per
source, and the ceiling is deliberately generous because offices and mobile
carriers put many real customers behind a single address.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Findings from the 2026-09-18 audit, ordered by block, each with its acceptance
criterion and audit id. Block 0 is closed; the remaining blocks record security,
architecture, scale and maintenance work, including the decisions that need a
product answer before any code.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
docker-compose.yml became the production/R2 stack, but LOCAL_SETUP.md still
documented "docker compose up --build" against .env.example, which fails on
missing R2_ENDPOINT, SITE_DOMAIN and KANBAN_DOMAIN, and MinIO was gone.
Add compose.local.yaml: builds from source, MinIO storage, fake providers,
disposable credentials, an app database password distinct from the
administrator one, and published origins in ALLOWED_ORIGINS so browser writes
are not rejected. Correct the documented command and the Kanban login, which
listed a username the email-validated model rejects.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
docker-compose.yml passed POSTGRES_PASSWORD as APP_DB_PASSWORD, so the DML-only
dtf_app role and the owning administrator shared one credential and the
privilege separation bootstrap.py sets up was decorative.
APP_DB_PASSWORD is now its own required variable, and bootstrap refuses to run
when it matches the administrator password, in both the URL and discrete-field
configuration forms.
Deploying this requires APP_DB_PASSWORD to be set in the stack environment
first; db-init rotates the role to it on the same deploy.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Navigation links carried data-modo-cta and opened loose artwork directly, so
every entry point except the price card bypassed "Arquivo por metro". Dropping
a PNG or JPG from the by-metre editor then switched the order to loose artwork
and repriced it, while the same screen advertised PNG/JPG in its drop zone and
marked that path as the cheaper one.
Replace the guessing with a declaration: the product is the choice. #tipoEnvio
shows ready sheet and loose artwork side by side with both prices, in all four
modes, reversible until a file is attached and locked afterwards. sel() no
longer reassigns modo, so a file can never change product or price on its own.
Verify the declaration instead of trusting it. medirFolha returns dpiFolha, and
an image without the pixels to span the film width at DPI_RECUSA is refused as
a sheet, with the numbers shown and one click to send it as loose artwork.
Previously a small image declared as a sheet was billed by the metre.
Scope pintaCaminhos to #caminhos .cam: its global selector was clearing the new
control. Move the checkout status out of #carr, which the success path hides,
so a paid order still confirms itself to the customer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
local/auth.py used os.environ without importing os, so new_session raised
NameError. Every first visit to /api/session, every registration and every
login returned 500, which left Site checkout, cart recovery and the customer
portal unusable since the R2 stack change.
Read COOKIE_SECURE once as a module constant and share it with local/app.py
instead of resolving the same variable in two places.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>