feat: accept sheets of up to 5 GB end to end
Some checks failed
Build and deploy / Validate source (push) Successful in 12s
Build and deploy / Integration suite on a real stack (push) Failing after 2m36s
Build and deploy / Secret scan and release gate (push) Successful in 7s
Build and deploy / Publish images (push) Has been skipped

Sheets of several GB are the normal order. The upload limit is now 5 GB.
ClamAV scans files up to 2 GB; a larger file is released only when its
first bytes match the format its name claims, and a disguised file is
refused. The Site grades a sheet over 150 MB from the pixel size in its
PNG, JPEG or WebP header without decoding it, and reads large PDFs in
ranges. The worker never opens a source over 300 MB: a finished sheet
placed whole becomes its own print file, which the Kanban offers to approve
as the final, and anything else goes to hand preparation. Files start
uploading as they enter the cart, with progress in the summary, and each
part renews the reservation so slow uploads do not expire. Quotas grow to
50 GB per customer and 500 GB in total; the Swarm config for ClamAV is
renamed because a deployed config cannot change in place.

Verified locally with a 386 MB and a 1.8 GB PNG (scanned, paid, original
as print file), a 2.3 GB PNG (format check) and a disguised 2.3 GB file
(refused).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
Cauê Faleiros
2026-09-29 13:18:18 -03:00
parent 4437232d27
commit 5f2be7ea20
21 changed files with 329 additions and 54 deletions

View File

@@ -9,6 +9,12 @@ The result is an ordinary upload row owned by an identity derived from the
order, already marked clean: its only inputs are artwork that passed the
malware scan, and the bytes are written here. The operator still decides
whether it becomes the final file; generation never approves anything.
Sheets of several GB are the normal order, and decoding one would take more
memory than the worker has. A source above LARGE_SOURCE_BYTES is never
opened: a finished sheet placed whole on the film is already its own print
file, so the original becomes the print file; any other layout is prepared
by hand from the original.
"""
import logging
import os
@@ -25,6 +31,9 @@ from .printfile import Unsupported, render
CLAIM_TIMEOUT = timedelta(minutes=15)
MAX_ATTEMPTS = 3
LARGE_SOURCE_BYTES = int(os.environ.get('PRINT_DECODE_MAX_BYTES', str(300 * 1024 * 1024)))
# Formats the operator can import as they are.
PRINTABLE_ORIGINAL = ('.png', '.jpg', '.jpeg', '.tif', '.tiff', '.pdf')
def generated_identity(order_id):
@@ -60,13 +69,26 @@ def render_one(storage):
c.execute('''UPDATE dtf_local.print_files SET status='rendering', claimed_at=now(),
attempts=attempts+1 WHERE id=%s''', (job['id'],))
item = job['snapshot']['items'][job['item_index']]
uploads = c.execute('''SELECT id,name,object_key,scan_state,purged_at,expires_at,
uploads = c.execute('''SELECT id,name,size,object_key,scan_state,purged_at,expires_at,
(expires_at<=now()) AS expired FROM dtf_local.uploads WHERE id=ANY(%s)''',
([UUID(u) for u in item['uploads']],)).fetchall()
# Every generated file shares its order's artwork retention deadline.
expiry = c.execute('SELECT min(created_at)+interval \'30 days\' AS e FROM dtf_local.uploads WHERE id=ANY(%s)',
([UUID(u) for u in item['uploads']],)).fetchone()['e']
by_id = {str(row['id']): row for row in uploads}
if any(row['size'] > LARGE_SOURCE_BYTES for row in uploads):
original = whole_sheet(item, by_id)
if original:
with connect() as c:
c.execute('''UPDATE dtf_local.print_files SET status='ready', upload_id=%s, detail=%s,
finished_at=now(), claimed_at=NULL WHERE id=%s''',
(original['id'], Jsonb({'source': 'original', 'name': original['name']}), job['id']))
audit('print_file_original', order=str(job['order_id']), item=job['item_index'])
else:
finish(job, 'manual', {'reason': f'arquivo acima de {LARGE_SOURCE_BYTES // 1048576} MB: '
'monte a folha a partir do original'})
audit('print_file_manual', order=str(job['order_id']), item=job['item_index'])
return True
try:
result = produce(storage, job, item, by_id)
except Unsupported as reason:
@@ -96,6 +118,25 @@ def render_one(storage):
return True
def whole_sheet(item, uploads):
"""The original, when the item is one finished sheet placed whole, once,
unrotated and unmirrored, across the film: then it is the print file."""
spec = item.get('production') or {}
sources, placements = spec.get('sources') or [], spec.get('placements') or []
if len(item['uploads']) != 1 or len(sources) != 1 or len(placements) != 1:
return None
source, place = sources[0], placements[0]
row = uploads.get(item['uploads'][0])
if (not row or row['scan_state'] != 'clean' or row['purged_at'] or row['expired']
or source.get('kind') != 'sheet' or int(source.get('copies', 1)) != 1
or not row['name'].lower().endswith(PRINTABLE_ORIGINAL)):
return None
if (float(place['x_cm']) != 0 or float(place['y_cm']) != 0 or int(place['rotation_degrees']) != 0
or place['mirrored'] or abs(float(source['width_cm']) - float(spec['film_width_cm'])) > 0.5):
return None
return row
def produce(storage, job, item, uploads):
"""Fetch the item's artwork and render it. Returns (pdf path, name, size, evidence)."""
for upload_id in item['uploads']: