* fix(external-media): one row per external file per event (#1162) Stable twin of #1167. Two overlapping import-external runs against the same event inserted every file twice. The route checked for an existing external_relpath and then inserted, with an fs.stat and a sharp().metadata() read sitting in between — a window wide enough for both runs to see "not there". A reporter's event held 8004 rows for 6012 distinct paths. Nothing at the storage layer stopped it: migration 041 created only a NON-unique (event_id, source_origin) index. - migration 176 removes the existing duplicates and adds a partial unique index on (event_id, external_relpath), verified against the catalog afterwards — a failed CREATE INDEX raises 23505 on Postgres, which run-migrations-safe treats as "schema already exists" and would record as applied on an install that never got the index. - dependent rows are removed explicitly rather than by cascade: PicPeak never sets `PRAGMA foreign_keys = ON`, so on SQLite the declared CASCADE is inert and a bare delete strands feedback and access-log rows. Guest feedback moves to the survivor instead of being discarded, keyed on guest identity the way feedbackService defines it, and the survivor's denormalized counters are recomputed. - the route treats a unique violation as a skip, so a writer this process cannot see converges instead of duplicating, and a second import while one is running gets a 409. - a .picpeak taken before migration 176 carries exactly these duplicates, and suspending FK enforcement does not suspend a unique index — so the restore drops the index for the load and rebuilds it after running the same dedupe. Divergences from the main twin, both because the feature is absent here: faces (no faceProcessor, so no purgePhotoFaces reconciliation — the rows are still deleted so nothing dangles), admin marks, transfer membership, and photos.view_count/download_count. The service guards each on hasTable / hasColumn, so those branches simply do not fire. Verified on this branch: 36 new tests pass; full suite leaves the same 5 pre-existing failures as origin/stable, unchanged. * fix(external-media): invalidate the download zip when duplicates are removed (#1162) External review. Same fix as the main twin. The pre-built "download everything" archive still contained the duplicate rows the dedupe had just deleted, so guests kept receiving them. Every ordinary photo-deletion path calls downloadZipService.invalidate for exactly this reason. The columns are cleared rather than the service being called: that service carries debounce timers and a regeneration queue, which a migration should not start. getZipInfo already treats a cleared record as a cache miss and rebuilds on the next request. The stale object is left in storage, as elsewhere. --------- Co-authored-by: Paul Nothaft <[email protected]>
This commit is contained in:
co-authored by
Paul Nothaft
parent
f83d144f28
commit
e9fcf4960e
@@ -0,0 +1,60 @@
|
||||
/**
|
||||
* Migration 176: one row per external file per event (#1162).
|
||||
*
|
||||
* The import route checked for an existing external_relpath and then inserted,
|
||||
* with an fs.stat and a sharp().metadata() call sitting in between — a window
|
||||
* wide enough that two overlapping imports of the same folder each see "not
|
||||
* there" and both insert. Nothing at the storage layer stopped them: 041
|
||||
* created only a NON-unique (event_id, source_origin) index. A reporter's
|
||||
* event ended up holding 8004 rows for 6012 distinct paths.
|
||||
*
|
||||
* So this does two things: clear the duplicates that already exist, and add
|
||||
* the constraint that makes the race unwinnable from here on.
|
||||
*
|
||||
* The work — which row survives, what happens to the guest feedback and admin
|
||||
* marks hanging off the loser, and why the dependent rows are deleted by hand
|
||||
* rather than left to ON DELETE CASCADE — lives in
|
||||
* services/externalPhotoDedupe.js, because a .picpeak restore has to run it
|
||||
* too: the archive carries the photos table verbatim, so a pre-#1162 backup
|
||||
* would otherwise hit the unique index mid-restore and roll the whole thing
|
||||
* back.
|
||||
*
|
||||
* Irreversible by design: down() drops the index but cannot resurrect the
|
||||
* deleted rows. They were never distinct data — the same file counted twice.
|
||||
*
|
||||
* What it does NOT do is delete the duplicates' thumbnail files. Those are
|
||||
* `ext<id>_<name>` keys under the thumbnail root, and a migration is the wrong
|
||||
* place to reach into storage — the backend may be pointed at S3, and a failed
|
||||
* object delete must not fail the schema change. They are left behind as
|
||||
* unreferenced bytes; the storage figures on the dashboard count them, which
|
||||
* is the correct answer to "what is on the disk".
|
||||
*/
|
||||
|
||||
const {
|
||||
dedupeExternalPhotos,
|
||||
createExternalRelpathIndex,
|
||||
dropExternalRelpathIndex,
|
||||
} = require('../../src/services/externalPhotoDedupe');
|
||||
|
||||
exports.up = async function(knex) {
|
||||
if (!(await knex.schema.hasTable('photos'))) return;
|
||||
if (!(await knex.schema.hasColumn('photos', 'external_relpath'))) return;
|
||||
|
||||
const removed = await dedupeExternalPhotos(knex);
|
||||
if (removed) {
|
||||
console.log(`176_external_relpath_unique: removed ${removed} duplicate external photo row(s)`);
|
||||
}
|
||||
|
||||
// Deliberately unguarded. Recording this migration as applied without the
|
||||
// index would leave the install permanently racy — the in-flight set only
|
||||
// covers one process, and the route's unique-violation path cannot converge
|
||||
// without a constraint to violate — with nothing to trigger a retry. A
|
||||
// failure here means the dedupe above did not achieve uniqueness, which is
|
||||
// worth stopping the upgrade for.
|
||||
await createExternalRelpathIndex(knex);
|
||||
};
|
||||
|
||||
exports.down = async function(knex) {
|
||||
if (!(await knex.schema.hasTable('photos'))) return;
|
||||
await dropExternalRelpathIndex(knex);
|
||||
};
|
||||
Reference in New Issue
Block a user