Files
dtf-system/.gitea/workflows/deploy.yml
Cauê Faleiros 7386469404
All checks were successful
Build and deploy / Validate source (push) Successful in 5s
Build and deploy / Integration suite on a real stack (push) Successful in 1m18s
Build and deploy / Secret scan and release gate (push) Successful in 6s
Build and deploy / Publish images and notify Portainer (push) Successful in 1m26s
docs: move the engineering documents into docs/
Thirteen files at the repository root, seven of them documents. Only README.md
earns a place there; the rest are now in docs/ beside the meeting notes, the
client roadmap and the historical material.

The compose files stay. docker-compose.yml is the path the dtf-cloud Portainer
stack reads, so moving it would break deployment, and Docker resolves a compose
file's relative build contexts against its own directory, so moving the other
two would silently break every build. Both reasons are now written down where
someone would otherwise try it.

Correcting references turned up a live fault: the Portainer stack creation
instructions still named deploy/stack.yaml as the compose path. That file was
removed, so anyone recreating the stack from these instructions would have
failed. It names docker-compose.yml now, with the reason it stays at the root.

ROADMAP.md keeps the paths its closed findings were written with, and says so at
the top. Those entries record where a fault was when it was found; rewriting
them to match a later layout would make the record less true, not more.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 17:52:20 -03:00

252 lines
11 KiB
YAML

name: Build and deploy
on:
pull_request:
push:
branches: [main]
workflow_dispatch:
jobs:
validate:
name: Validate source
runs-on: ubuntu-latest
timeout-minutes: 15
steps:
- name: Checkout
uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683
- name: Run fast regression checks
run: |
python3 -m py_compile app/*.py app/**/*.py ops/*.py deploy/*.py
python3 -m unittest \
tests.test_dependency_lock \
tests.test_staging_readiness \
deploy.test_production_preflight \
tests.test_pricing \
tests.test_secrets -v
sh -n infra/lock_dependencies.sh
integration:
name: Integration suite on a real stack
needs: validate
runs-on: ubuntu-latest
timeout-minutes: 45
env:
# The runner shares the host's Docker daemon, so every published port is
# taken on the machine itself. Known occupants of that host:
# 8000, 9443 Portainer (the Edge tunnel and its UI)
# 18080/18081 the production dtf-cloud stack (docker-compose.yml defaults)
# 9000/9001 MinIO defaults elsewhere
# This block avoids all of them. Ephemeral ports are not an option: the
# published port is baked into PUBLIC_ORIGIN, ALLOWED_ORIGINS and the CSP
# when the containers start, so it has to be known beforehand.
SITE_PORT: "28080"
KANBAN_PORT: "28081"
API_PORT: "28000"
STORAGE_PORT: "29000"
STORAGE_CONSOLE_PORT: "29001"
# Presigned URLs are signed against this endpoint, so it must be reachable
# by whoever follows them. The suites run inside the network, so it has to
# be the service name, not a published port on the host.
S3_PUBLIC_ENDPOINT: http://storage:9000
COMPOSE: docker compose -f compose.local.yaml
steps:
- name: Checkout
uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683
# py_compile cannot see an unresolved name, and the four unit tests above
# never start the application. A missing import in local/auth.py therefore
# reached production and returned 500 on every session, login and
# registration. These suites exercise the running stack and would have
# failed on it immediately.
- name: Start the stack
run: |
$COMPOSE up --build -d --wait --wait-timeout 600
$COMPOSE ps
# Run inside the stack's own network. The runner is itself a container, so
# ports published on the host's loopback are in a different namespace and
# unreachable from here. SITE_HOST_HEADER keeps the Host the gateway and
# TrustedHostMiddleware expect, so the configuration under test is the same
# one a developer exercises on localhost.
- name: API and workflow regressions
run: |
for suite in smoke_test workflow_test security_test scanning_test; do
echo "--- $suite"
$COMPOSE exec -T \
-e SITE_BASE_URL=http://site \
-e SITE_HOST_HEADER=localhost \
api python -m "tests.$suite"
done
- name: Runtime and retention regressions
run: |
$COMPOSE exec -T api python -m tests.retention_test
$COMPOSE exec -T api python -m tests.runtime_security_test
# These need a real Chrome. They are the only coverage for the artwork
# editor and the full customer journey, so install google-chrome-stable
# (or set CHROME_BIN) on the runner to make them gate deployments. The
# suites above stay hard gates either way.
- name: Browser regressions
run: |
for candidate in "$CHROME_BIN" /usr/bin/google-chrome-stable \
/usr/bin/google-chrome /usr/bin/chromium /usr/bin/chromium-browser; do
if [ -n "$candidate" ] && [ -x "$candidate" ]; then
export CHROME_BIN="$candidate"
break
fi
done
if [ ! -x "${CHROME_BIN:-}" ]; then
echo "::warning::No Chrome on this runner; browser regressions were NOT run."
echo "Install google-chrome-stable or set CHROME_BIN to gate on them."
exit 0
fi
# Chrome runs here, in the runner container, and reaches the stack only
# through ports published on the host. When the runner is itself a
# container those are in another namespace, so check before running
# rather than failing with a bare connection error. See ROADMAP 5.10.
if ! wget -q -T 5 -O /dev/null "http://localhost:${SITE_PORT}/health"; then
echo "::warning::Stack not reachable from the runner; browser regressions were NOT run."
exit 0
fi
echo "Using $CHROME_BIN"
node tests/artwork_browser_test.mjs
node tests/browser_test.mjs
- name: Diagnostics on failure
if: failure()
run: |
$COMPOSE ps || true
$COMPOSE logs --tail 200 api worker site kanban || true
- name: Tear down
if: always()
run: $COMPOSE down -v || true
scan:
name: Secret scan and release gate
needs: validate
runs-on: ubuntu-latest
timeout-minutes: 30
env:
TRIVY_IMAGE: ${{ vars.TRIVY_IMAGE }}
ENFORCE_PRODUCTION_PREFLIGHT: ${{ vars.ENFORCE_PRODUCTION_PREFLIGHT }}
steps:
- name: Checkout
uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683
# Blocking. A credential committed by accident must never reach the
# registry or the deployed stack, and the repository is clean today, so
# this gate costs nothing until it is actually needed.
- name: Secret scan
run: |
image="${TRIVY_IMAGE:-aquasec/trivy:0.58.1}"
docker run --rm -v "$PWD:/src:ro" "$image" \
fs --scanners secret --exit-code 1 --severity HIGH,CRITICAL \
--no-progress /src
# docs/PORTAINER.md described this as blocking publication. It never ran at
# all, and turning it on unconditionally would block every deploy: the
# source preflight refuses a release while the payment and messaging
# adapters are fake, which is the deliberate state the stack runs in
# today. So its verdict is always printed, and enforcement is opt-in.
# Set the repository variable ENFORCE_PRODUCTION_PREFLIGHT to "true" once
# real adapters land, and this becomes the gate the documentation claims.
- name: Production source preflight
run: |
set +e
python3 deploy/production_preflight.py --source-only
verdict=$?
set -e
if [ "$verdict" -eq 0 ]; then
echo "Source preflight passes."
exit 0
fi
if [ "${ENFORCE_PRODUCTION_PREFLIGHT:-false}" = "true" ]; then
echo "::error::Source preflight blocked the release."
exit "$verdict"
fi
echo "::warning::Source preflight reports blockers (advisory; set ENFORCE_PRODUCTION_PREFLIGHT=true to gate)."
publish-and-deploy:
name: Publish images and notify Portainer
needs: [validate, integration, scan]
if: gitea.event_name == 'push' && gitea.ref == 'refs/heads/main'
runs-on: ubuntu-latest
timeout-minutes: 45
env:
TRIVY_IMAGE: ${{ vars.TRIVY_IMAGE }}
PYTHON_BASE_IMAGE: ${{ vars.PYTHON_BASE_IMAGE }}
NGINX_BASE_IMAGE: ${{ vars.NGINX_BASE_IMAGE }}
steps:
- name: Checkout
uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683
- name: Sign in to the Gitea Container Registry
env:
REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }}
REGISTRY_TOKEN: ${{ secrets.REGISTRY_TOKEN }}
run: |
test -n "$REGISTRY_USERNAME"
test -n "$REGISTRY_TOKEN"
echo "$REGISTRY_TOKEN" | docker login gitea.blyzer.com.br \
--username "$REGISTRY_USERNAME" --password-stdin
- name: Build and publish API
run: |
image="gitea.blyzer.com.br/blyzer/dtf-api"
# The Dockerfiles pin digests themselves; these variables let a base be
# moved forward without editing the repository. --pull is intentionally
# absent: a digest already names one immutable image.
set --
[ -n "$PYTHON_BASE_IMAGE" ] && set -- --build-arg PYTHON_BASE_IMAGE="$PYTHON_BASE_IMAGE"
docker build --file deploy/Dockerfile.api "$@" \
--build-arg VCS_REF="${{ gitea.sha }}" \
--tag "$image:latest" --tag "$image:${{ gitea.sha }}" .
docker push "$image:latest"
docker push "$image:${{ gitea.sha }}"
- name: Build and publish web
run: |
image="gitea.blyzer.com.br/blyzer/dtf-web"
set --
[ -n "$PYTHON_BASE_IMAGE" ] && set -- --build-arg PYTHON_BASE_IMAGE="$PYTHON_BASE_IMAGE"
[ -n "$NGINX_BASE_IMAGE" ] && set -- "$@" --build-arg NGINX_BASE_IMAGE="$NGINX_BASE_IMAGE"
docker build --file deploy/Dockerfile.web "$@" \
--build-arg VCS_REF="${{ gitea.sha }}" \
--tag "$image:latest" --tag "$image:${{ gitea.sha }}" .
docker push "$image:latest"
docker push "$image:${{ gitea.sha }}"
# CRITICAL blocks, HIGH is reported. Both images carry zero CRITICAL after
# the base pinning and OS upgrades, so this gate holds the line already
# reached. The remaining HIGH findings have no upstream fix, so failing on
# them would stop releases without making anything safer.
- name: Image vulnerabilities
run: |
image="${TRIVY_IMAGE:-aquasec/trivy:0.58.1}"
failed=0
for target in \
"gitea.blyzer.com.br/blyzer/dtf-api:${{ gitea.sha }}" \
"gitea.blyzer.com.br/blyzer/dtf-web:${{ gitea.sha }}"; do
echo "--- $target (HIGH, reported)"
docker run --rm -v /var/run/docker.sock:/var/run/docker.sock "$image" \
image --scanners vuln --severity HIGH --no-progress \
--format table --exit-code 0 "$target" ||
echo "::warning::Could not scan $target for HIGH findings"
echo "--- $target (CRITICAL, blocking)"
docker run --rm -v /var/run/docker.sock:/var/run/docker.sock "$image" \
image --scanners vuln --severity CRITICAL --no-progress \
--format table --exit-code 1 "$target" || failed=1
done
if [ "$failed" -ne 0 ]; then
echo "::error::A CRITICAL vulnerability was found in a published image."
exit 1
fi
- name: Trigger Portainer redeployment
env:
PORTAINER_WEBHOOK: ${{ secrets.PORTAINER_WEBHOOK }}
run: |
if [ -z "$PORTAINER_WEBHOOK" ]; then
echo "PORTAINER_WEBHOOK is not configured; images were published but deployment was skipped."
exit 0
fi
curl --fail --silent --show-error --max-time 30 --request POST "$PORTAINER_WEBHOOK"