Two follow-up fixes inside the same install-experience surface as
the previous commit:
1. **Removed `docker compose exec -T backend npm run migrate`** in
both install_docker and update_docker_installation. The backend
container's wait-for-db.sh already runs `npm run migrate:safe`
on startup; the script was racing it with a separate (and
non-safe) `npm run migrate`. That race is the most likely
actual mechanism behind #484's "relation 'photos' does not
exist" error on the second install attempt — partial schema
visible to one of the two parallel migrators. Replaced with a
bounded wait for the backend container to become healthy
(Docker healthcheck reports green only after wait-for-db.sh
finishes its migration pass).
2. **Added the missing frontend container** to the script-generated
compose. The script previously generated a postgres + redis +
backend stack with no frontend at all (backend on host port
3001), while the documented production install
(docker-compose.production.yml) ships postgres + redis +
backend + frontend (nginx /api proxy on host port 3000). That
shape divergence is half of issue B in #484 — script-installed
admins had no frontend container and were left wondering where
the UI lived. Aligning both compose files on the same shape
eliminates the divergence; the frontend uses curl in its
healthcheck (frontend/Dockerfile explicitly `apk add curl`)
unlike the backend.
The remaining piece of issue B — picking ONE canonical install
path (build-from-source script vs. prebuilt-image production
compose) and deprecating the other — is a deployment-strategy
call that deserves its own design pass. Both paths now produce
architecturally-equivalent stacks.
Three install-experience bugs that compounded into MrGabri's "fresh
install fails" report:
1. **postgres healthcheck noise.** `pg_isready -U <user>` without
-d defaults to probing a database whose name matches the user.
Since DB_NAME defaults to picpeak_prod (not picpeak), every
healthcheck interval logged
FATAL: database "picpeak" does not exist
into postgres logs even though the install was working
correctly. Reporter saw the FATAL, assumed broken, restarted
with DB_NAME=picpeak, hit a tainted-state migration error on
the second try, filed a bug. Fixed in both
docker-compose.production.yml and the inline compose generated
by scripts/picpeak-setup.sh — pin -d to ${DB_NAME} so the
probe hits the real database.
2. **backend container shows perpetually `unhealthy`.** Both
compose files used `curl -f` for the backend healthcheck, but
backend/Dockerfile only installs dumb-init + postgresql-client +
ffmpeg — no curl. Switch to wget --no-verbose --tries=1 --spider
to match what backend/Dockerfile's own HEALTHCHECK already
does. Now docker ps, docker compose ps, and the backend image's
built-in healthcheck all agree.
3. **stale separate `workers` container.** scripts/picpeak-setup.sh
still generated a second container running `npm run workers`
alongside the backend, but workers (fileWatcher,
expirationChecker, emailQueueProcessor, backgroundProcessor,
webhookWorker) have been started by server.js in-process for
a while — see the comment at line ~895 of the same script for
the systemd-side cleanup. The duplicate container caused two
file watchers and two expiration checkers to compete for the
same DB rows. Removed from the generated compose; install +
upgrade paths now stop and rm any pre-existing picpeak-workers
container.
Doesn't address issue B's bigger architecture mismatch (the script
generates a build-from-source compose with no frontend container,
while docker-compose.production.yml uses prebuilt images with a
separate frontend container). That deserves its own design pass
to pick a canonical install path and align — out of scope here.
The gallery promotional banner (#440) read as visually offset from
the gallery footer because:
- Footer used `container text-center px-4` (full container width,
centered text).
- Promo block used `container py-4 sm:py-6` with an inner
`max-w-3xl mx-auto` wrapper holding left-aligned text — a
narrower column with left-aligned content sitting in the
middle of the page.
Two issues compounded: the column was narrower than the footer AND
its text alignment differed. Reported by Rekoo-PS in #482 with a
screenshot showing the misalignment, with a request for an admin
alignment option.
Fix:
- Drop the inner max-w-3xl wrapper. Promo content now spans the
same .container width as the footer, eliminating the
narrower-column visual.
- Default text alignment changed from left → center to match the
footer.
- New `branding_promo_alignment` setting ('left' | 'center' | 'right',
default 'center'). Surfaced as a dropdown next to the existing
Position dropdown on the BrandingPage. Live preview block on the
BrandingPage mirrors the gallery render so admins see what
guests will see.
- Also replaced the no-op `prose-sm` prose-modifier with a real
`prose prose-sm` outer class so the existing `prose-a:text-accent`
modifier actually takes effect (it didn't before — modifiers
without an outer .prose are silently ignored by Tailwind
Typography).
Migration 103 seeds the new setting at 'center' so existing
installs that have a promo banner today see the corrected
alignment immediately on next deploy.
i18n: en + de hand-translated; nl/pt/ru/fr machine-translated and
flagged for native review per project convention.
PR #477 moved Trivy from the merge-* job into the per-arch build-*
matrix scanning by digest. The amd64 leg works; the arm64 leg
crashes with:
remote error: no child with platform linux/amd64 in index
ghcr.io/.../<image>@sha256:<digest>
Root cause: docker/build-push-action wraps every push in an OCI
index — the actual image manifest sits next to a SLSA provenance
attestation manifest as siblings under the digest. Trivy's remote
backend defaults to linux/amd64 when resolving an index, so:
- amd64 leg → looks for amd64 child → finds the amd64 image → ok.
- arm64 leg → looks for amd64 child → finds NO amd64 child
(the only platform child is arm64) → fails.
Fix: set TRIVY_PLATFORM = ${{ matrix.platform }} on each leg's
Trivy step. Each scanner then asks for its own arch and finds it.
SLSA provenance attestation stays attached to the per-arch images
— a real win for supply-chain visibility we'd lose if we'd
disabled provenance instead.
amd64 was the only thing keeping CI partly green; this restores
full green across both legs without touching the build artifact
shape.
Initial pinning shipped a tag that doesn't exist in the
aquasecurity/trivy-action repo. Workflow run failed with:
Unable to resolve action 'aquasecurity/trivy-action@0.28.0',
unable to find version '0.28.0'
The repo's tags use a v prefix (v0.36.0, v0.35.0, …). Bumping
both occurrences (build-backend and build-frontend matrix jobs)
to v0.36.0, which is the latest stable as of 2026-04-22.
Resolves the intermittent "no child with platform linux/amd64 in
index" failure on the merge-backend job — and fixes the same latent
bug on merge-frontend before it surfaces.
Two compounding root causes per Luca's diagnosis:
1. aquasecurity/trivy-action@master was unpinned, so the action and
its bundled Trivy binary float on every CI run. A green build
could flip red overnight without a single repo change.
2. Trivy was asked to scan a multi-platform OCI index by tag (the
merge-* jobs ran AFTER manifest creation). Its remote resolver
cannot reliably pick the right per-arch child out of an index
reference — it needs a single-platform reference (digest, or a
--platform flag).
Fix:
- Move the Trivy + upload-sarif steps OUT of merge-backend /
merge-frontend and INTO the per-arch build-backend / build-frontend
matrix jobs. Each leg scans the image it just pushed by its
sha256 digest (`...@${{ steps.build.outputs.digest }}`), which is
always single-platform by construction.
- Pin aquasecurity/trivy-action@0.28.0 (was @master).
- Distinct SARIF category per arch
(`backend-vulnerabilities-linux-amd64`, …-arm64) so an
amd64-only finding in a base layer doesn't get masked by the
arm64 scan in the Security tab.
- Move security-events: write down to the build-* jobs (where the
scan now runs) and remove it from the merge-* jobs (which only
publish the manifest now).
Out of scope: flipping `exit-code: '1'` to actually gate CI on
findings. Worth doing as a separate follow-up after an audit pass —
landing it here would surprise beta with a red build for any
pre-existing CRITICAL/HIGH in current images. Inline TODO in the
workflow notes the deferral.
Background: galleryOgService already serves OG/Twitter Card meta tags
to social-crawler User-Agents (WhatsApp, Facebook, Slack, Telegram,
Discord, ~21 in total) for /gallery/:slug URLs. Today the og:image
is always the brand logo with the inline rationale "no protected
photo content".
#474 asked for a hero/cover photo preview. The trade-off is that any
URL embedded in og:image is fetched unauthenticated by every
link-preview crawler — so an opted-in image is effectively public
to anyone the gallery URL is shared to. Ship as a per-event boolean,
default FALSE, so existing galleries never start surfacing photos
without explicit admin intent.
Schema (migration 102):
- events.og_image_share_enabled BOOLEAN NOT NULL DEFAULT FALSE.
Backend:
- galleryOgService.buildOgMetadata: when opt-in is on AND a
hero_photo_id is set AND the photo has a generated thumbnail,
emit og:image as /og/gallery/:slug/cover. Falls back to the
brand logo on any miss (deleted hero, missing thumbnail, no
opt-in) so a half-configured gallery still gets a polished
preview rather than a broken-image src.
- galleryOgService.handleGalleryOgCover: new public endpoint that
streams the hero thumbnail. Validates slug shape, checks the
opt-in flag + hero presence + thumbnail existence; returns 404
on any failure. ETag = thumbnail mtime + photo id so a
regenerated thumb busts crawler caches. Cache-Control:
public, max-age=300 (short — admins shouldn't wait an hour for
a cover swap to land in chat previews).
- server.js: mount the new GET /og/gallery/:slug/cover route. The
existing nginx ^~ /og/gallery/ proxy block already covers it.
- adminEvents.js: validator + persistence on POST + PUT.
formatBoolean coercion so SQLite (0/1) and Postgres (boolean)
both behave correctly.
Frontend:
- Event type + UpdateEventData carry og_image_share_enabled.
- EventDetailsPage adds a checkbox under the HeroPhotoSelector,
disabled when no hero photo is picked. Help text deliberately
spells out the public-by-design consequence — admins shouldn't
flip this on for a sensitive gallery without realising what
they're sharing with link-preview crawlers.
Tests: 8 new in galleryOgService.shareImage.test.js — pin the
cover-vs-logo decision contract (3 cases) plus the defensive
fallbacks (deleted hero, missing thumbnail) and the 404 contract
on the cover endpoint (4 cases). The 404 tests assert that
ensureThumbnail() is NOT called when opt-in is off, so a future
refactor can't accidentally widen the unauthenticated cover
endpoint to expose a hero the admin hasn't shared.
i18n: en + de hand-translated; nl + pt + ru + fr machine-translated
and flagged for native review per project convention.
The trigger: PR #458 mounted requireCustomerPortalEnabled which
410'd every /api/customer/* + /api/admin/customers/* request when
the master toggle was off. Some browsers cached that 410 (no
Cache-Control header was set, so heuristic freshness applied —
the wrong default for an authenticated/sensitive surface).
PR #470 reverted the middleware, but a customer whose tab cached
the 410 still saw 410s until they hard-refreshed.
Add noStoreCache middleware and mount it in front of both route
groups. Every response (200, 4xx, 5xx) now carries
`Cache-Control: no-store, no-cache, must-revalidate, private`
plus the HTTP/1.0 Pragma + Expires fallbacks. Any future
transient error from these endpoints can no longer get pinned in
browser or proxy caches and outlive its cause.
Cost is one setHeader per request; applied per route group rather
than globally so static assets + galleries keep their own caching
strategy unchanged.
Includes a dedicated unit test pinning the header set so a future
cleanup pass can't quietly drop it and re-introduce the bug.
4 unit tests pinning the contract of the customer-minted JWT
re-check added in #470:
- via='customer' + customerId, assignment present → next() runs.
- via='customer' + customerId, assignment removed → 403 with
CUSTOMER_ASSIGNMENT_REVOKED code.
- customerId in payload but `via` claim missing → no re-check
(defends against a future refactor accidentally widening the
gate to match every legacy session that happens to carry a
customerId field).
- per-event-password JWT (no via, no customerId) → no
event_customer_assignments query at all (asserted by counting
db() invocations — a regression that quietly added a re-check
here would 403 every guest the moment any unrelated customer
was unassigned from any event).
Same mock pattern as customerAuth.middleware.test.js. The re-check
is the load-bearing piece behind the "Manage galleries" dialog
UX promise — these tests guard it explicitly.
5 new tests covering the diff math (added/removed), the
archived-event filter, the no-op short-circuit when wanted equals
existing, and the type-coercion of the wanted-list input. Mirrors
the existing setAssignmentsForEvent suite shape so the inverse-
direction service function carries equivalent regression coverage.
This function is the writer behind the "Manage galleries" dialog
and the verifyGalleryAccess re-check together form the access-
control story for the whole feature — getting the diff math
wrong here means assignments don't actually revoke, which is the
entire promise of the new UI.
The Dashboard "Recent Activity" widget and the header notifications
dropdown both rendered raw activity-type strings (e.g. the literal
"feature_flags_updated") for any type missing from their lookup
maps — including everything emitted by the recently-added customer
portal (#354), webhooks (#327), API tokens (#322), event types,
event-publish flow, admin user management (#350), and the
feature-flags reorg itself.
Two coordinated changes:
1. Smart formatter for feature_flags_updated. The backend writes
`metadata.changed = { [flagKey]: { from, to } }` on every save.
New formatFeatureFlagsChanged() helper in admin.service.ts reads
that diff and renders:
- 1 change → "Customer Portal enabled"
- N changes → "3 features updated: Customer Portal enabled,
Calendar disabled, Quotes enabled"
Per-flag display labels source from `settings.features.<key>.title`
so they stay in sync with the Features tab. Unknown flag keys
fall through to a humanised version of the key.
2. 33 missing activity types added to BOTH renderers and to the
`admin.activities.*` + `admin.notificationMessages.*` i18n
namespaces across all six locales. Coverage groups: customer
portal (13 types), admin user management (6), webhooks (3),
API tokens (2), event types (4), event publish/logo (3), bulk
delete (1), and assorted post-merge surfaces (4).
The notifications.service.ts switch + admin.service.ts fallback
message map are still duplicated; consolidating them into a
single source of truth is a follow-up worth doing before the
next significant addition. For now both stay in sync via this PR.
en + de hand-translated. nl + pt + ru + fr machine-translated and
flagged for native review per project convention.
Settings → Features showed the customer-portal toggle as "Accounts"
("Konten" in DE, "Comptes" in FR, etc.) — the deeper sub-nav label
inside ClientsLayout — while the prominent menu-bar entry the admin
actually clicks first reads "Clients" / "Kunden". The mismatch was
confusing on first encounter ("which one do I look for?").
Align the Features tab card title and the "Sidebar:" callout with
the menu-bar wording (`navigation.clients`) across all six locales.
The sub-nav inside ClientsLayout keeps its own "Accounts" label —
that one matches the /admin/clients/accounts URL and is correct.