fix(install): silence pg healthcheck noise + drop legacy workers container (#484)

Three install-experience bugs that compounded into MrGabri's "fresh
install fails" report:

1. **postgres healthcheck noise.** `pg_isready -U <user>` without
   -d defaults to probing a database whose name matches the user.
   Since DB_NAME defaults to picpeak_prod (not picpeak), every
   healthcheck interval logged
     FATAL: database "picpeak" does not exist
   into postgres logs even though the install was working
   correctly. Reporter saw the FATAL, assumed broken, restarted
   with DB_NAME=picpeak, hit a tainted-state migration error on
   the second try, filed a bug. Fixed in both
   docker-compose.production.yml and the inline compose generated
   by scripts/picpeak-setup.sh — pin -d to ${DB_NAME} so the
   probe hits the real database.

2. **backend container shows perpetually `unhealthy`.** Both
   compose files used `curl -f` for the backend healthcheck, but
   backend/Dockerfile only installs dumb-init + postgresql-client +
   ffmpeg — no curl. Switch to wget --no-verbose --tries=1 --spider
   to match what backend/Dockerfile's own HEALTHCHECK already
   does. Now docker ps, docker compose ps, and the backend image's
   built-in healthcheck all agree.

3. **stale separate `workers` container.** scripts/picpeak-setup.sh
   still generated a second container running `npm run workers`
   alongside the backend, but workers (fileWatcher,
   expirationChecker, emailQueueProcessor, backgroundProcessor,
   webhookWorker) have been started by server.js in-process for
   a while — see the comment at line ~895 of the same script for
   the systemd-side cleanup. The duplicate container caused two
   file watchers and two expiration checkers to compete for the
   same DB rows. Removed from the generated compose; install +
   upgrade paths now stop and rm any pre-existing picpeak-workers
   container.

Doesn't address issue B's bigger architecture mismatch (the script
generates a build-from-source compose with no frontend container,
while docker-compose.production.yml uses prebuilt images with a
separate frontend container). That deserves its own design pass
to pick a canonical install path and align — out of scope here.
This commit is contained in:
Paul Nothaft
2026-05-14 20:31:25 +02:00
parent 409ddf9c93
commit 0b0b1bb2d5
2 changed files with 57 additions and 21 deletions
+14 -3
View File
@@ -15,7 +15,13 @@ services:
- picpeak-network
restart: unless-stopped
healthcheck:
test: ["CMD-SHELL", "pg_isready -U ${DB_USER:-picpeak}"]
# `pg_isready -U <user>` without -d defaults to probing a database
# whose name matches the user — postgres then logs constant
# `FATAL: database "picpeak" does not exist` even though the
# actual DB is `picpeak_prod`. Pinning -d to DB_NAME makes the
# probe hit the real database and silences the log noise that
# made #484's reporter think the install was broken.
test: ["CMD-SHELL", "pg_isready -U ${DB_USER:-picpeak} -d ${DB_NAME:-picpeak}"]
interval: 10s
timeout: 5s
retries: 5
@@ -64,8 +70,13 @@ services:
condition: service_healthy
restart: unless-stopped
healthcheck:
# Backend exposes /health on internal port 3000
test: ["CMD", "curl", "-f", "http://localhost:3000/health"]
# Backend exposes /health on internal port 3000.
# The backend image only ships wget (Alpine base) — using curl
# here makes `docker ps` show the container as `unhealthy`
# indefinitely even when /health responds. Mirrors the wget-based
# HEALTHCHECK already declared in backend/Dockerfile so docker
# compose, plain `docker run`, and `docker ps` all agree.
test: ["CMD", "wget", "--no-verbose", "--tries=1", "--spider", "http://localhost:3000/health"]
interval: 30s
timeout: 10s
retries: 3