Alpine ships BusyBox ps (no -p PID, no pgrep), which failed CI on the
first run of this workflow with "ps: unrecognized option: p". Replace
the pgrep-then-ps chain with `ps -o user,comm | awk '$2=="node"'`
which works on both BusyBox (Alpine, in the container) and procps
(the GitHub runner host, though we don't use it here).
The fresh-install restart loop reported by @MrGabri (and confirmed by
@AloePacci with the user:0:0 workaround) had a clear root cause:
- Dockerfile pinned USER nodejs (UID 1001) before the entrypoint
ran, so the existing chown branch in init-production.sh:13 was
dead code.
- wait-for-db.sh (the actual entrypoint, not init-production.sh)
silently swallowed mkdir/EACCES on bind mounts with || true,
then a downstream migration error surfaced as the visible failure.
- Net effect on a typical Linux host where the bind-mount dir is
owned by UID 1000: container can't write, exits non-zero,
restarts forever with no clear error.
Switch to the standard Docker drop-privileges pattern:
1. Install su-exec, drop `USER nodejs` from the Dockerfile —
container now starts as root.
2. wait-for-db.sh: if running as root, chown /app/storage,
/app/data, /app/logs to nodejs and re-exec self via
su-exec nodejs:nodejs. App still ends up running as UID 1001.
3. Preflight check for non-root invocations (compose `user:`
overrides): verify the bind mounts are actually writable
before continuing. If not, exit 1 immediately with an
actionable error pointing at the docs — no more silent
restart loops.
Also:
- Delete backend/init-production.sh. It was an orphan — no caller
in the Dockerfile, compose, or anywhere else. Its chown logic
looked authoritative enough that @MrGabri ran it manually trying
to debug, which is what finally surfaced the EACCES.
- docker-compose.yml: drop user: + PUID/PGID env. The pattern-B
UID-matching workaround they implemented is obsolete now that
pattern A (root-then-drop) is in place.
- .env.example + README: drop PUID/PGID documentation.
- Add fresh-install smoke test workflow. Boots backend + postgres
against bind mounts owned by UID 1000 (the GitHub runner UID,
and the common-mismatch case on Linux hosts) and verifies:
+ container reaches healthy without restart-looping
+ chown happened (dirs now owned by 1001 inside the container)
+ node runs as nodejs, not root (su-exec drop worked)
+ /health returns status:ok
+ with --user 5005:5005 + unwritable mounts, preflight exits
loud with the expected error string
Verified locally end-to-end against a fresh Postgres + UID-501-owned
bind mount: backend reaches healthy in ~20s, chown applied, node
runs as nodejs, no restart loop. Docs in picpeak-docs cover the new
behavior + a Troubleshooting section for the install-path bugs
fixed in #484/#494/#511/#488.
Refs: #484
PR #477 moved Trivy from the merge-* job into the per-arch build-*
matrix scanning by digest. The amd64 leg works; the arm64 leg
crashes with:
remote error: no child with platform linux/amd64 in index
ghcr.io/.../<image>@sha256:<digest>
Root cause: docker/build-push-action wraps every push in an OCI
index — the actual image manifest sits next to a SLSA provenance
attestation manifest as siblings under the digest. Trivy's remote
backend defaults to linux/amd64 when resolving an index, so:
- amd64 leg → looks for amd64 child → finds the amd64 image → ok.
- arm64 leg → looks for amd64 child → finds NO amd64 child
(the only platform child is arm64) → fails.
Fix: set TRIVY_PLATFORM = ${{ matrix.platform }} on each leg's
Trivy step. Each scanner then asks for its own arch and finds it.
SLSA provenance attestation stays attached to the per-arch images
— a real win for supply-chain visibility we'd lose if we'd
disabled provenance instead.
amd64 was the only thing keeping CI partly green; this restores
full green across both legs without touching the build artifact
shape.
Initial pinning shipped a tag that doesn't exist in the
aquasecurity/trivy-action repo. Workflow run failed with:
Unable to resolve action 'aquasecurity/trivy-action@0.28.0',
unable to find version '0.28.0'
The repo's tags use a v prefix (v0.36.0, v0.35.0, …). Bumping
both occurrences (build-backend and build-frontend matrix jobs)
to v0.36.0, which is the latest stable as of 2026-04-22.
Resolves the intermittent "no child with platform linux/amd64 in
index" failure on the merge-backend job — and fixes the same latent
bug on merge-frontend before it surfaces.
Two compounding root causes per Luca's diagnosis:
1. aquasecurity/trivy-action@master was unpinned, so the action and
its bundled Trivy binary float on every CI run. A green build
could flip red overnight without a single repo change.
2. Trivy was asked to scan a multi-platform OCI index by tag (the
merge-* jobs ran AFTER manifest creation). Its remote resolver
cannot reliably pick the right per-arch child out of an index
reference — it needs a single-platform reference (digest, or a
--platform flag).
Fix:
- Move the Trivy + upload-sarif steps OUT of merge-backend /
merge-frontend and INTO the per-arch build-backend / build-frontend
matrix jobs. Each leg scans the image it just pushed by its
sha256 digest (`...@${{ steps.build.outputs.digest }}`), which is
always single-platform by construction.
- Pin aquasecurity/trivy-action@0.28.0 (was @master).
- Distinct SARIF category per arch
(`backend-vulnerabilities-linux-amd64`, …-arm64) so an
amd64-only finding in a base layer doesn't get masked by the
arm64 scan in the Security tab.
- Move security-events: write down to the build-* jobs (where the
scan now runs) and remove it from the merge-* jobs (which only
publish the manifest now).
Out of scope: flipping `exit-code: '1'` to actually gate CI on
findings. Worth doing as a separate follow-up after an audit pass —
landing it here would surprise beta with a red build for any
pre-existing CRITICAL/HIGH in current images. Inline TODO in the
workflow notes the deferral.
Triage of an external SAST/SCA scan run on 2026-05-06. Most loud findings
were already resolved by PR #412 (the 18-CVE backport); this PR addresses
the residual real items:
* Drop unused `handlebars` from backend deps. The runtime require was
removed in PR #367 (#367) but the package.json line stayed. handlebars
was the source of two flagged criticals (CVE-2026-33937 RCE,
GHSA-2w6w-674q-4c4q AST injection) plus 8 highs — all now gone.
* `npm audit fix` on backend + frontend. Bumps transitive picomatch,
flatted, postcss, brace-expansion via lockfile, and direct dompurify,
lodash, vite, i18next-http-backend within their existing semver ranges.
Both audits now report 0 vulnerabilities.
* Add `event.origin === window.location.origin` check to the THEME_PREVIEW
message listener in PreviewPage. The branding page posts from the same
origin, so nothing legitimate is rejected; without the check, any third
party that window.open()'d the preview could push arbitrary
branding/theme payloads (semgrep
insufficient-postmessage-origin-validation).
* nginx: `proxy_hide_header` for X-Frame-Options, X-Content-Type-Options,
Referrer-Policy, Content-Security-Policy, Permissions-Policy,
Strict-Transport-Security at server level. nginx adds these itself, but
helmet on the backend was also emitting them — clients were seeing
duplicates (testssl flagged "Multiple X-Frame-Options / CSP /
Permissions-Policy / Referrer-Policy headers" on the live origin).
Single source of truth now.
* Dockerfile hardening (checkov):
- HEALTHCHECK on backend/Dockerfile, backend/Dockerfile.dev,
frontend/Dockerfile.dev. Frontend production Dockerfile already had
one.
- USER node in frontend/Dockerfile.dev (was running as root).
* GitHub Actions docker-build.yml: explicit top-level
`permissions: contents: read`. Per-job blocks already declare
`packages: write` where needed; this stops future steps from
inheriting unintended privileges (CKV2_GHA_1).
Backend npm audit: 4 vulns -> 0.
Frontend npm audit: 6 vulns -> 0.
Backend unit tests: 13 suites, 131/132 passing (1 pre-existing skip).
Frontend type-check + lint: clean.
The pre-existing integration-test failures (live DB / S3 required) and
the ThemeCustomizerEnhanced QueryClientProvider failures are unrelated
and reproduce on origin/beta without these changes.
Remove sync-versions job that fails on protected branches.
Instead, use Release Please's extra-files feature to update
package.json versions as part of the release PR.
QEMU emulation of ARM64 on x86 GitHub runners is too slow and
unreliable for npm operations, causing builds to hang or crash
with "Illegal instruction" errors.
Changed platform detection logic to:
- Tagged releases (v*.*.*): Build both amd64 and arm64
- All other builds (branches, PRs): Build amd64 only
This ensures fast CI feedback during development while still
providing multi-arch images for production releases.
Sharp library native binaries cause QEMU 'Illegal instruction' errors during
ARM64 emulation. This change builds only amd64 for PR checks (faster, reliable)
while maintaining multi-arch (amd64+arm64) builds for main/develop/tags.
The workflow was generating invalid Docker tags with format ':-3b251d7'
due to empty branch names in PR contexts.
Problem:
- Tag config: type=sha,prefix={{branch}}-,format=short
- For PRs: {{branch}} is empty → results in ':-3b251d7' (invalid)
- Docker doesn't allow tags starting with hyphen
Solution:
- Changed to: type=sha,format=short
- Now generates: '3b251d7' (valid) without branch prefix
- Works correctly for PRs, branches, and tags
Valid tag examples now:
- PRs: pr-44, 3b251d7
- Branches: main, 3b251d7
- Tags: v1.0.0, 1.0, 1, 3b251d7
The publish-manifest job was failing because it tried to create manifests
from non-existent architecture-specific tags (latest-amd64, latest-arm64).
docker/build-push-action@v5 already creates multi-arch manifests automatically
when building for multiple platforms, making this job redundant.
The workflow now correctly builds and pushes multi-arch images in a single
step with proper manifest lists included.
🤖 Generated with [Claude Code](https://claude.ai/code)
Co-Authored-By: Claude <noreply@anthropic.com>