0874a30ac9
* feat(docker): add all-in-one image — backend + frontend in one container (#1042) One container, one Node process, SQLite by default: `docker run` with no compose file, no nginx, no supervisor, no bundled Postgres/Redis. - Dockerfile.aio (repo-root context): frontend build stage + backend deps stage + a runtime stage mirroring backend/Dockerfile's production stage, with the built SPA copied to /app/frontend/dist and SERVE_FRONTEND=true. DATABASE_CLIENT=sqlite3 and STORAGE_PATH=/app/storage are pinned explicitly — the storage fallback resolves to container-root /storage, which EACCESes after the su-exec drop. - server.js: the SERVE_FRONTEND block now does what the nginx image did — renders ${BRAND_TITLE}/${BRAND_DESCRIPTION} into index.html once at boot, serves that rendered shell on /index.html and every SPA route, caches hashed /assets/* immutably while the shell revalidates, and gzips the bundle via compression() mounted after all /api routers. express.static now runs with index:false so `/` keeps flowing to handlePublicSiteRequest — its default index option was shadowing the landing page on native installs. - wait-for-db.sh: skip the Postgres readiness wait when DATABASE_CLIENT is sqlite3. The engine resolver still runs, still logs, and still refuses the populated-both conflict (#1038). - .dockerignore: **/node_modules, so the root-context build can't pick up host deps from backend/ or frontend/. - docker-build.yml: build-aio / merge-aio follow the same per-arch build → digest-merge → per-version tag scheme as backend/frontend (GHCR only for now; the Docker Hub mirror is wired once the Hub repo exists), plus a smoke-aio job that boots the image on every PR and asserts /health, the SPA shell, the rendered brand title, immutable asset caching and the SQLite engine resolution. Pointing DB_HOST/DB_USER/DB_PASSWORD + DATABASE_CLIENT=pg at an external Postgres works exactly like the backend image. * fix(ci): correct three smoke-aio assertions that would fail a green image (#1042) Found by running the smoke job locally against a real build — the image passed every behavioral check, but three assertions were wrong: - `/` asserts 200, but handlePublicSiteRequest 302s to /admin/login while the public landing site is disabled, which is the state of the fresh install the smoke container always is. Assert the redirect target instead — that still proves express.static's index option is not shadowing the handler, which is the thing the check exists for. - The placeholder-leak grep matched index.html's explanatory comment, which mentions BRAND_TITLE in prose and survives into the built shell. Match the literal ${BRAND_TITLE}/${BRAND_DESCRIPTION} tokens with -F, and cover the description token too. - Add a gzip assertion, probing with GET: the compression middleware skips bodyless responses, so a HEAD probe reports no Content-Encoding even when compression is active. Verified locally on linux/arm64: image builds clean, boots to healthy in ~8s on the SQLite default, and 25/25 checks pass (SPA shell, rendered brand title, immutable+gzipped assets, no-store shell, SPA fallbacks, npm removed, su-exec drop to nodejs, no errors in the boot log). The DATABASE_CLIENT=pg override was exercised against a real Postgres too — the readiness wait still runs and the engine resolves to postgres. * fix(server): serve the SPA for every client route, not just /admin and /gallery (#1042) nginx did `try_files $uri $uri/ /index.html`, so behind compose every client-side route survived a direct hit or a refresh and the short `['/admin', '/admin/*', '/gallery/*']` list was never exercised. Without nginx that list is the whole contract, and everything outside it 404'd: /setup /customer /impressum /datenschutz /payment-check /quote/:token /contract/:token /invite/:token /transfer/:token /transfer-upload/:token /setup is the first URL a new install visits, so the all-in-one image was unusable from a cold start. The catch-all is registered after `app.use('/api', notFoundHandler)`, so an unknown /api route still answers JSON instead of being handed the HTML shell, and after the /s/:shortSlug resolver, so a typo'd short URL still 404s (#699). It is GET-only — a stray POST keeps 404ing rather than getting a 200 page back. The handler is hoisted out of the SERVE_FRONTEND block via `spaCatchAll` because that block runs before the API 404 handler is registered. Verified on the built image: all ten routes above now 200, /api/nope still returns JSON 404, /s/nonexistent still returns 404, / still 302s to /admin/login, and the smoke suite is 25/25. Both boundaries are now asserted in the smoke-aio job. * docs(readme): document the single-container install (#1042) The README had no mention of the all-in-one image, so the only way to discover it was reading the workflow file. Adds a Quick Start subsection with the one-line `docker run` and the `docker exec … cat SETUP_TOKEN` step, plus a row in the documentation table. Deliberately does not sell it as the default: the note says the compose stack is still the right choice for anything busier, gives the reason (SQLite takes one writer at a time), and points at the `.picpeak` restore as the way out, so nobody picks it and then finds themselves stuck. Full details live at docs.picpeak.app/deployment/single-container (PicPeak/docs#8). * feat(docker): fold #1067's items into the all-in-one image (#1042) Consolidating the two parallel AIO branches into this one. This PR's approach is kept wherever the two differed on design — in particular the in-process brand render, `index: false` (which fixes express.static shadowing handlePublicSiteRequest, a bug #1067 had), the compression middleware, and the smoke-aio job. What follows is what #1067 had that this branch did not. Layout — the issue asks for a single mountable root, and this moves to one: /data/db picpeak.db (+ -wal/-shm) and SETUP_TOKEN /data/storage originals, thumbnails, archives /data/logs application logs /data/backup built-in backup output; /backup symlinks here `-v picpeak:/data` and nothing else to remember. README and the smoke job's database-path assertion follow the new layout. Correctness items: - sqlite CLI. DatabaseBackupService SPAWNS `sqlite3` for `.backup` and PRAGMA integrity_check; the npm module does not ship that binary. backend/Dockerfile omits it because compose always runs Postgres — this image defaults to SQLite, so every database backup failed with ENOENT. - /backup wired in. Migrations 029 + 030 seed /backup/picpeak and /backup/database as the backup destinations; nothing created or mounted them, so backups had nowhere to write and anything written would die with the container. Symlinked into the volume, subdirectories created at startup (a bind mount hides the tree baked into the image), and adopted only when BACKUP_DIR is set so it never gates boot for compose deployments that do not mount it. - logger.js honours LOG_DIR. It hard-coded <backend>/logs, so logs could not leave the container. Unset keeps the old path for every existing install. - wait-for-db.sh derives its writable roots from STORAGE_PATH / DATA_DIR / LOG_DIR instead of hard-coded /app paths, and mkdir -p's them before chown — a bind-mounted /data hides the image's tree, and chown against a missing path reports "the filesystem rejects chown", which is both wrong and a dead end. - .dockerignore excludes backend/-prefixed runtime data. Docker reads only the root file, so the unprefixed data/*.db, logs/* and storage/* rules missed backend/data, backend/logs and backend/storage entirely; a checkout used to run PicPeak would bake its database, photos, logs and SETUP_TOKEN into a published layer. - HEALTHCHECK follows $PORT rather than a hard-coded 3000. - --max-http-header-size=32768 matches nginx's large_client_header_buffers 4 32k; Node's 16 KiB default would reject a guest carrying several per-gallery JWT cookies. docs/single-container.md is added as the in-repo reference the README links to. The smoke job gains four assertions for the above: the one-volume layout and writable backup destinations, the sqlite3 CLI, logs landing on the volume, and the image carrying no runtime data from the build context. Verified on a built image — named volume, bind mount and PORT=8080 all healthy; every existing smoke assertion still passes, including / -> 302 /admin/login, the rendered BRAND_TITLE, immutable assets, gzip and /s/<unknown> -> 404. Co-authored-by: Luca-Timo <102960244+Luca-Timo@users.noreply.github.com> * fix(docker): restore the SPA-fallback exclusions and close the build-context leak (#1042) Both found by external review of the consolidated branch. - The SPA catch-all had no backend-owned exclusions. This was a regression I introduced while merging: #1067 carried a BACKEND_OWNED prefix list, and taking this branch's server.js wholesale (correctly — its index:false and in-process brand render are the better design) dropped it. /photos, /thumbnails, /uploads and /fonts are static mounts whose middleware calls next() on a miss, so the catch-all was answering 200 text/html under image and font URLs instead of 404. nginx gave each of those its own location block, so try_files never applied to them. - backend/data is now excluded wholesale rather than by suffix. The suffix list (*.db, *.db-wal, *.db-shm, SETUP_TOKEN) let real secrets through: a used checkout carries ADMIN_CREDENTIALS.txt next to the database, plus -journal files and any DATABASE_PATH not ending in .db. Since Dockerfile.aio builds from the repository root and COPYs backend/ wholesale, any of those would be baked into a published layer. The directory holds only runtime state and is already gitignored in full. smoke-aio gains an assertion that the backend static routes still 404, so the exclusion cannot be dropped again silently. Verified on a built image: /photos, /thumbnails, /fonts and /uploads misses all 404; /setup, /impressum, /gallery/x, /admin/login still 200; / still 302s to /admin/login; /api/nope still answers JSON; /s/<unknown> still 404s; and the image carries no *.db, ADMIN_CREDENTIALS.txt, logs or storage from the context. * fix(aio): three failures that only surface outside a dev laptop (#1042) Backups aborted on SQLite. getTableChecksums() built its digest with `CAST(t.* AS TEXT)`, which is Postgres row-to-text syntax; SQLite parses `*` there as a syntax error, so every backup threw before reaching the .backup call. Since the all-in-one image ships SQLite by default, that is every AIO install. Enumerate the columns via columnInfo() and sum their lengths instead. The shared /data mount root was never adopted. wait-for-db.sh chowned the children it creates but not the mount point itself, so a host directory arriving as 0700 with a foreign owner stayed untraversable by UID 1001 after the su-exec drop. Docker Desktop's permissive bind mounts hide this completely, which is why local testing passed; a NAS share does not. DATA_ROOT is now adopted first. Maintenance mode locked the admin out of the box. The middleware runs at server.js:493, long before the static block at 891, and exempted the auth endpoints but not the page that calls them. With the backend serving the frontend, /admin/login and /assets/* returned 503 JSON, so an admin who enabled maintenance mode could never load the UI to turn it off. nginx serves those paths in the compose stack, which is why it never surfaced there. Guest and API surfaces stay gated. Verified on a built image: checksums compute across all 95 tables; a bind mount created 0700/4000:4000 boots healthy and ends up 1001:1001; with general_maintenance_mode=true, /admin/login, /admin and /assets/* return 200 while /gallery/* and /api/gallery/* return 503 — and 503 across all three once the exemption is removed again. Claude-Session: https://claude.ai/code/session_01Ra4hcsYiKuQLbbRsg6EjAc * fix(aio): stop leaking .env into the image, fix the broken checksum test (#1042) The Jest suite was red: mocking db.raw is no longer enough now that the SQLite checksum branch asks the query builder for its column list, so db(table) came back undefined and getTableChecksums failed on every PR. The production code is right; the fixture needed to know about the call. backend/.env was landing in the published layer. The root ignore file's `.env`, `.env.*` and `data/*.db` rules read as unanchored but Docker matches them from the context root, so they catch ./.env and never backend/.env — and `COPY backend/ .` then puts a real JWT_SECRET at /app/.env. Matched at any depth instead, the way **/node_modules in the same file already is. Confirmed by building from a checkout carrying a planted secret: before, `cat /app/.env` printed it back. Business documents wrote outside the volume. quoteService, invoice sending/reminders and contract signatures build paths from process.cwd()/storage and never read STORAGE_PATH; compose hides it by setting STORAGE_PATH=/app/storage with WORKDIR /app so the two are the same directory. Here they are not, and /app is root-owned, so a quote or invoice PDF failed to write as UID 1001 — and would not survive the container if it had. Symlinked /app/storage into the volume, matching the /backup symlink beside it. Teaching those services STORAGE_PATH is the real fix and wants its own change. Two smaller ones: the mount root is now chowned shallow rather than recursively, since every child below it is already walked recursively and a NAS-sized photo library should not be traversed twice on each restart; and /assets/ joins the backend-owned prefixes, so a stale hashed chunk requested by a tab left open across an upgrade gets a 404 instead of index.html served with 200 under a .js URL. Verified on a built image: planted backend/.env and backend/probe.db are absent; /app/storage resolves to /data/storage and a business-doc write as UID 1001 appears on the host; a 0700 bind mount owned by 4000:4000 boots healthy; a missing /assets chunk 404s while the real bundle still serves 200 as application/javascript. The databaseBackup suite is green again, and the branch adds no failing suite that origin/main does not already fail on the same machine. Claude-Session: https://claude.ai/code/session_01Ra4hcsYiKuQLbbRsg6EjAc * test(aio): teach the leak assertion about the storage symlink (#1042) The previous check listed /app/storage/events and treated a hit as a leak. That was true while /app/storage was either absent or a copied directory; now it is a symlink into the volume, so the check followed it and found the empty tree the image itself creates — a false positive on its own design. Check the shape instead: /app/storage must be a symlink pointing at /data/storage, and the volume's photo tree must contain no files on a fresh install. A real directory there now fails loudly, which is the condition the assertion was always trying to catch. Also extended the path list to /app/.env and loose database files, matching the .dockerignore rules added alongside. Claude-Session: https://claude.ai/code/session_01Ra4hcsYiKuQLbbRsg6EjAc * fix(aio): show the maintenance screen instead of raw JSON to guests (#1042) The previous commit exempted the admin shell so an admin could still reach the switch they had just flipped. Guests had the same problem for the same reason: with no nginx in front, /gallery/<slug> reaches this middleware long before the static block, so a visitor during maintenance got a 503 JSON body where every other deployment shows the branded maintenance screen the frontend already ships. Replaced the two path-specific exemptions with the rule they were both special cases of: a GET that is not an API call and not a backend-owned content mount is the SPA shell, and the shell is inert HTML — it boots, reads /api/public/settings (already exempt) and renders MaintenanceMode on its own. Everything that carries real data stays gated: /api/*, /photos/, /thumbnails/, /fonts/, and any non-GET. Compose is untouched by construction, since nginx answers those paths and they never arrive here. Verified on a built image with the flag on: /gallery/x, /customer/x, /admin and /admin/login return 200 text/html while /api/gallery/x/verify, /photos/x.jpg and /thumbnails/x.jpg return 503 and a POST to a public API still returns 503; with the flag off the same paths go back to 404. Added a middleware test over that exemption matrix — over-exemption is the real risk in this change, so it asserts the gated half too. It fails on five cases without the fix. Claude-Session: https://claude.ai/code/session_01Ra4hcsYiKuQLbbRsg6EjAc * fix(aio): stop the shell exemption from un-gating /og and the public CMS (#1042) The previous commit exempted "any GET that is not an API call". That negative rule reads as safe and is not: /og/gallery/<slug> and its /cover render the event name and the hero thumbnail, /s/<code> renders short-link previews, and `/` is handed to the public CMS. All four are proxy_passed to the backend by nginx, so they were gated before this PR in every deployment — the rule un-gated them, and for compose too, not just the new image. A site switched to maintenance would have kept publishing gallery metadata. Replaced the guess with the split nginx already defines: exempt what the frontend container answers itself, gate what it proxies. That is the same rule the all-in-one image needs by definition, since its whole job is to be both halves of that stack, and it now matches compose in both directions rather than only in the direction the last commit tested. Verified on a built image with the flag on: /admin/login, /gallery/<slug> and /customer/* return 200, while /, /og/gallery/x, /og/gallery/x/cover, /s/abc, /robots.txt, /api/* and /photos/* return 503; with the flag off all of them behave normally again. The middleware test grew the gated cases — it now covers 21, most of them asserting what must NOT be exempt. Claude-Session: https://claude.ai/code/session_01Ra4hcsYiKuQLbbRsg6EjAc * fix(aio): give the image a FRONTEND_URL default so share links are absolute (#1042) getFrontendBaseUrl() reads FRONTEND_URL, falls back to the general_site_url setting, and otherwise returns an empty string — which makes share_url come back as a bare "/gallery/<slug>/<token>". Compose defaults the variable to http://localhost:3000, but the documented one-liner for this image passes only JWT_SECRET, so every fresh single-container install handed out relative links in API responses, QR codes and emails. Defaulted to the same value compose uses; -e FRONTEND_URL=https://... overrides it, as does the site URL field in Settings. Found by pointing tests/e2e/local at a running AIO container: auth/06-api-tokens asserts share_url matches /^https?:\/\//, and it was the one spec that failed for a product reason rather than a harness one. It passes now, and the suite is 19/20 against the image — the remaining failure is smoke/02-auth-flow, whose seed helper shells out to a hard-coded `docker exec picpeak-backend`, so it cannot arrange its precondition against any other container. Claude-Session: https://claude.ai/code/session_01Ra4hcsYiKuQLbbRsg6EjAc * feat(aio): mark the image so face recognition stays off (#1042, #1074) Face recognition needs a separate ML container this image does not contain, and enabling it here would add a second image-processing pipeline competing with Sharp for the CPU and memory of a container sized for one photographer plus guests browsing. The failure mode would not be a clear error — just a slow install that looks broken. The backend gate for this lands in #1075 and keys on PICPEAK_SINGLE_CONTAINER. Without this line the guard never triggers on an actual all-in-one build, so the two changes have to arrive together: whichever merges second completes the pair. Verified against this file's exact value — isFeatureEnabled() returns false with it set. An explicit marker rather than inferring from SERVE_FRONTEND or the SQLite path, because legitimate multi-container deployments do both of those and should keep the feature. Also adds it to the Limits section of docs/single-container.md, next to the SQLite and Redis constraints, since that is where someone will look before choosing this image. --------- Co-authored-by: Paul Nothaft <paul@MacStudio-von-Paul.local> Co-authored-by: the-luap <paul-nothaft@hotmail.de>
237 lines
11 KiB
Bash
Executable File
237 lines
11 KiB
Bash
Executable File
#!/bin/sh
|
|
# wait-for-db.sh - Wait for PostgreSQL to be ready before starting the application
|
|
|
|
set -e
|
|
|
|
# Machine secrets (JWT/DB/Redis): if not supplied via the environment, read them
|
|
# from the generated secret files that the compose `secrets-init` service writes
|
|
# to /run/secrets. Explicit env ALWAYS wins, so installs that set
|
|
# JWT_SECRET/DB_PASSWORD/REDIS_PASSWORD in .env are unaffected. Runs before the
|
|
# root -> nodejs re-exec so the exported values survive su-exec.
|
|
for _pair in JWT_SECRET:jwt_secret DB_PASSWORD:db_password REDIS_PASSWORD:redis_password; do
|
|
_var="${_pair%%:*}"
|
|
_file="/run/secrets/${_pair##*:}"
|
|
eval "_cur=\${$_var:-}"
|
|
if [ -z "$_cur" ] && [ -s "$_file" ]; then
|
|
export "$_var=$(cat "$_file")"
|
|
fi
|
|
done
|
|
unset _pair _var _file _cur
|
|
|
|
# Permission handling (#484): the image starts as root so this script can
|
|
# chown bind-mounted host volumes to UID 1001 (nodejs) before dropping
|
|
# privileges via su-exec. This avoids the fresh-install restart loop where
|
|
# the host directory's UID (commonly 1000) didn't match the container's
|
|
# hard-coded nodejs user. Compose deployments that pin `user:` to something
|
|
# other than root skip this branch — they own permissions themselves and hit
|
|
# the preflight check below instead.
|
|
# The writable roots. Defaults are the compose layout; the all-in-one image
|
|
# (#1042) points all of them under one mounted volume, so these must follow the
|
|
# same env vars the app itself reads rather than hard-coding /app.
|
|
DATA_DIRS="${STORAGE_PATH:-/app/storage} ${DATA_DIR:-/app/data} ${LOG_DIR:-/app/logs}"
|
|
|
|
# The backup root is adopted when explicitly configured, but never gates boot:
|
|
# docker-compose.production.yml does not mount /backup, so a hardened non-root
|
|
# deployment would fail `mkdir -p /backup` against a root-owned / and refuse to
|
|
# start over a directory it never needed.
|
|
if [ -n "${BACKUP_DIR:-}" ]; then
|
|
DATA_DIRS="$DATA_DIRS $BACKUP_DIR"
|
|
fi
|
|
|
|
# When every root lives under ONE mounted volume (the all-in-one image, #1042),
|
|
# the mount point itself must be adopted too. Chowning only the children leaves
|
|
# a host directory created with 0700/0750 and a foreign owner untraversable by
|
|
# UID 1001 after the su-exec drop, so the preflight below rejects children the
|
|
# script just created. Docker Desktop's permissive bind mounts hide this; a NAS
|
|
# share does not.
|
|
#
|
|
# It is deliberately kept out of DATA_DIRS: everything below it is already
|
|
# chowned recursively, so adding it there would walk the whole photo library a
|
|
# second time on every restart — minutes of startup delay on exactly the large
|
|
# NAS libraries this image targets. The mount point needs its own ownership
|
|
# fixed, nothing more, so it gets a shallow chown of its own below.
|
|
DATA_ROOT_DIR="${DATA_ROOT:-}"
|
|
|
|
# Create the roots before touching them. With the compose layout each is its own
|
|
# mount point so they always exist — but the AIO image mounts ONE volume at
|
|
# /data, and a bind-mounted host directory hides the tree baked into the image.
|
|
# chown would then fail on paths that do not exist and report "the filesystem
|
|
# rejects chown", which is both wrong and a dead end for NAS users.
|
|
# shellcheck disable=SC2086 — intentional word-splitting over the roots
|
|
mkdir -p $DATA_ROOT_DIR $DATA_DIRS 2>/dev/null || true
|
|
|
|
if [ "$(id -u)" = "0" ]; then
|
|
if [ -n "$DATA_ROOT_DIR" ] && ! chown nodejs:nodejs "$DATA_ROOT_DIR" 2>/dev/null; then
|
|
echo "ERROR: failed to chown $DATA_ROOT_DIR to nodejs (UID 1001)." >&2
|
|
echo " The mounted volume root must be traversable by UID 1001 after the privilege drop." >&2
|
|
echo " Workaround: chown 1001:1001 the host directory you mounted at $DATA_ROOT_DIR." >&2
|
|
exit 1
|
|
fi
|
|
if ! chown -R nodejs:nodejs $DATA_DIRS 2>/dev/null; then
|
|
echo "ERROR: failed to chown $DATA_DIRS to nodejs (UID 1001)." >&2
|
|
echo " This usually means the host filesystem rejects chown (e.g. NFS without root squash" >&2
|
|
echo " disabled, or a SELinux/AppArmor policy blocking the operation)." >&2
|
|
echo " Workaround: pre-chown the host directories to 1001:1001 and pin 'user: \"1001:1001\"'" >&2
|
|
echo " in your compose file so this script never tries to chown them itself." >&2
|
|
echo " See https://docs.picpeak.app/deployment/docker#permissions" >&2
|
|
exit 1
|
|
fi
|
|
exec su-exec nodejs:nodejs "$0" "$@"
|
|
fi
|
|
|
|
# Belt-and-suspenders: if we got here as non-root (compose `user:` override),
|
|
# verify the bind mounts are actually writable before proceeding. Failing
|
|
# loud here beats the previous behavior — silent mkdir-||-true at line 69
|
|
# followed by a confusing migration error and a restart loop.
|
|
_uid="$(id -u)"
|
|
_gid="$(id -g)"
|
|
for _dir in $DATA_ROOT_DIR $DATA_DIRS; do
|
|
if [ ! -w "$_dir" ]; then
|
|
echo "ERROR: $_dir is not writable by UID $_uid." >&2
|
|
echo " Either drop the 'user:' override from your compose file so the container starts as" >&2
|
|
echo " root and can self-fix permissions, or run on the host:" >&2
|
|
echo " chown -R $_uid:$_gid <host-mount-for-$_dir>" >&2
|
|
echo " See https://docs.picpeak.app/deployment/docker#permissions" >&2
|
|
exit 1
|
|
fi
|
|
done
|
|
|
|
# Explicit SQLite boots (DATABASE_CLIENT=sqlite3 — the all-in-one image's
|
|
# default, #1042) have no Postgres to wait for: skip the whole readiness/
|
|
# create/verify section below. The engine resolver further down still runs,
|
|
# still logs the resolved engine, and still refuses the populated-both
|
|
# conflict (#1038). Compose deployments pin DATABASE_CLIENT=pg and knexfile's
|
|
# production block defaults to pg when unset, so nothing changes for them.
|
|
if [ "${DATABASE_CLIENT:-}" != "sqlite3" ]; then
|
|
|
|
host="${DB_HOST:-postgres}"
|
|
port="${DB_PORT:-5432}"
|
|
user="${DB_USER:-picpeak}"
|
|
target_db="${DB_NAME:-picpeak}"
|
|
|
|
# Hand the app EXACTLY the connection this script verified. knexfile's
|
|
# production block defaults DB_HOST to `db` while this script defaults to
|
|
# `postgres`, so a bare `docker run` with no DB_HOST would have had the
|
|
# readiness check pass against one host and the app then dial another (#1038
|
|
# review). Compose sets DB_HOST explicitly and is unaffected.
|
|
export DB_HOST="$host"
|
|
export DB_PORT="$port"
|
|
export DB_USER="$user"
|
|
export DB_NAME="$target_db"
|
|
# Use target database for checks - the picpeak user may not have access to 'postgres' database
|
|
default_db="${DB_CHECK_DB:-$target_db}"
|
|
|
|
sanitize_identifier() {
|
|
printf '%s' "$1" | sed "s/'/''/g"
|
|
}
|
|
|
|
echo "Waiting for PostgreSQL at $host:$port..."
|
|
|
|
# First, wait for PostgreSQL server to be reachable
|
|
max_attempts=30
|
|
attempt=0
|
|
while [ $attempt -lt $max_attempts ]; do
|
|
if PGPASSWORD="$DB_PASSWORD" psql -h "$host" -p "$port" -U "$user" -d "$target_db" -c '\q' >/dev/null 2>&1; then
|
|
>&2 echo "PostgreSQL is up - database \"$target_db\" is accessible."
|
|
break
|
|
fi
|
|
|
|
# If target DB doesn't work, try connecting to 'postgres' or 'template1' to create it
|
|
if PGPASSWORD="$DB_PASSWORD" psql -h "$host" -p "$port" -U "$user" -d "template1" -c '\q' >/dev/null 2>&1; then
|
|
>&2 echo "PostgreSQL is up - checking if database \"$target_db\" needs to be created..."
|
|
|
|
# Check if database exists
|
|
db_exists=$(PGPASSWORD="$DB_PASSWORD" psql -h "$host" -p "$port" -U "$user" -d "template1" -tAc "SELECT 1 FROM pg_database WHERE datname = '$(sanitize_identifier "$target_db")'" 2>/dev/null || echo 0)
|
|
|
|
if [ "$db_exists" != "1" ]; then
|
|
>&2 echo "Database \"$target_db\" not found. Attempting to create..."
|
|
if PGPASSWORD="$DB_PASSWORD" psql -h "$host" -p "$port" -U "$user" -d "template1" -c "CREATE DATABASE \"$target_db\";" >/dev/null 2>&1; then
|
|
>&2 echo "Database \"$target_db\" created successfully."
|
|
else
|
|
>&2 echo "Warning: Could not create database. It may already exist or user lacks permissions."
|
|
fi
|
|
fi
|
|
break
|
|
fi
|
|
|
|
attempt=$((attempt + 1))
|
|
>&2 echo "PostgreSQL is unavailable - sleeping (attempt $attempt/$max_attempts)"
|
|
sleep 2
|
|
done
|
|
|
|
if [ $attempt -eq $max_attempts ]; then
|
|
>&2 echo "Failed to connect to PostgreSQL after $max_attempts attempts."
|
|
exit 1
|
|
fi
|
|
|
|
# Final verification - wait for target database to accept connections
|
|
until PGPASSWORD="$DB_PASSWORD" psql -h "$host" -p "$port" -U "$user" -d "$target_db" -c '\q' >/dev/null 2>&1; do
|
|
>&2 echo "Waiting for database \"$target_db\" to accept connections..."
|
|
sleep 2
|
|
done
|
|
|
|
>&2 echo "Target database \"$target_db\" is ready."
|
|
|
|
else
|
|
>&2 echo "DATABASE_CLIENT=sqlite3 — skipping the PostgreSQL readiness wait."
|
|
fi # end Postgres wait (skipped for explicit sqlite3 boots)
|
|
|
|
# Ensure storage directories exist with proper permissions (Issue #67 fix)
|
|
# When host directories are bind-mounted, the container's built-in directories are overridden
|
|
# This ensures the required directory structure exists before the application starts
|
|
echo "Ensuring storage directories exist..."
|
|
STORAGE_BASE="${STORAGE_PATH:-/app/storage}"
|
|
mkdir -p "$STORAGE_BASE/events/active" "$STORAGE_BASE/events/archived" "$STORAGE_BASE/thumbnails" 2>/dev/null || true
|
|
|
|
# Backup destinations seeded by migrations 029 + 030 (/backup/picpeak and
|
|
# /backup/database). Creating the root alone is not enough: on a bind mount the
|
|
# subdirectories baked into the image are hidden and the backup services do not
|
|
# create them, so a backup would fail with ENOENT.
|
|
BACKUP_BASE="${BACKUP_DIR:-/backup}"
|
|
if [ -d "$BACKUP_BASE" ]; then
|
|
mkdir -p "$BACKUP_BASE/picpeak" "$BACKUP_BASE/database" 2>/dev/null || true
|
|
fi
|
|
|
|
# Resolve which database engine this boot should use (#1038) BEFORE migrations
|
|
# run, while the Postgres target is still untouched. An install that has been
|
|
# unknowingly running on SQLite (the image used to leave NODE_ENV unset, so
|
|
# knexfile.js fell back to sqlite3 and ignored DB_HOST/DB_USER/DB_PASSWORD)
|
|
# keeps serving from its SQLite file instead of coming up against an empty
|
|
# Postgres. The exported value survives the `exec` below, so the migration
|
|
# runner and the server agree on the engine.
|
|
RESOLVED_DB_CLIENT="$(node scripts/resolve-db-engine.js)"
|
|
RESOLVER_STATUS=$?
|
|
# Exit 3 means two populated databases with no record of which is current
|
|
# (#1038). Starting either would hide the other's data, so stop here — the
|
|
# resolver has already printed what to do.
|
|
if [ "$RESOLVER_STATUS" = "3" ]; then
|
|
exit 1
|
|
fi
|
|
# Validate rather than trust: anything unexpected on stdout (a stray log line
|
|
# from a library that writes to the console) must not become DATABASE_CLIENT,
|
|
# which would break knexfile for every process that follows.
|
|
case "$RESOLVED_DB_CLIENT" in
|
|
pg|sqlite3)
|
|
export DATABASE_CLIENT="$RESOLVED_DB_CLIENT"
|
|
;;
|
|
"")
|
|
>&2 echo "Database engine resolver returned nothing; falling back to the configured client."
|
|
;;
|
|
*)
|
|
>&2 echo "Database engine resolver returned an unexpected value; ignoring it and falling back to the configured client."
|
|
;;
|
|
esac
|
|
|
|
# Run migrations (use safe runner in production). Invoked via node directly —
|
|
# the runtime image no longer ships npm (see Dockerfile: its bundled deps kept
|
|
# tripping CVE scanners while npm itself never runs in production).
|
|
echo "Running database migrations..."
|
|
if [ "$NODE_ENV" = "production" ]; then
|
|
node migrations/run-migrations-safe.js
|
|
else
|
|
node migrations/run-migrations.js
|
|
fi
|
|
|
|
# Execute the main command
|
|
exec "$@"
|