DEDUP: real-server breakage campaign + bit-rot scrub before the photo import #231

Closed
opened 2026-09-27 11:57:46 +00:00 by kayg · 12 comments
Owner

Owner question (2026-09-27): "how robust is deduplication? has this been stress / breakage tested?" Production (calternal.cloud) will soon hold the owner's real photo library, so dedup must be proven before the import.

Current state (checked)

  • DESIGN §5: instance-wide BLAKE3 CAS in .cas/<2>/<digest>, hardlinks (reflink when supported), files 0444, dedup after upload completes, quota counts logical size.
  • calternal-fs unit tests cover: concurrent identical writes, wrong-bytes blob replaced, copy reuses blob, orphan GC, GC waiting for in-flight links across roots, journal-step crash recovery.
  • NOT covered: real-server breakage of dedup. The adversarial probe only checks that .cas is unreachable by path.

Build

  1. Real-server dedup campaign in tests/adversarial/ (new section dedup), run against a real local server:
    • storms of identical and near-identical uploads from several users and sessions at once, including sync clients and WebDAV if present;
    • kill -9 of the server at random points during upload/dedup/GC; after restart, every file must have the right bytes (BLAKE3 against the source) and journals must be empty;
    • delete / trash / restore / version / rename / move / share across users that share one blob; GC must never delete a blob that is still linked; no user sees another user's file;
    • quota: logical size per user stays exact through all of the above;
    • edit a deduped file through every write path (editor, upload overwrite, sync, WebDAV); no other user's copy may change (the server must never write in place into a shared inode);
    • an outside process changes a file in a home in place on disk (simulate a stray tool); detect and contain it, do not spread it to other links;
    • timing: prove no dedup timing oracle (upload of a known file vs unknown file has the same latency distribution within noise).
  2. Bit-rot scrub: a scheduled low-priority job (idle IO class, nice) that re-hashes .cas entries against their names, rate-limited, resumable. On mismatch: quarantine the entry, record which user files are affected, raise an admin notice, and repair from any intact copy if one exists (for example a version or a non-deduped copy). Admin setting for the schedule; default weekly. Expose last scrub status in admin.
  3. Fix every finding. Blocks: data loss/corruption, cross-user leak, crash. Regression test for each.
  4. Update DESIGN §5 (ASD-STE100) with the scrub and the guarantees proven.
Owner question (2026-09-27): "how robust is deduplication? has this been stress / breakage tested?" Production (calternal.cloud) will soon hold the owner's real photo library, so dedup must be proven before the import. ## Current state (checked) - DESIGN §5: instance-wide BLAKE3 CAS in `.cas/<2>/<digest>`, hardlinks (reflink when supported), files 0444, dedup after upload completes, quota counts logical size. - calternal-fs unit tests cover: concurrent identical writes, wrong-bytes blob replaced, copy reuses blob, orphan GC, GC waiting for in-flight links across roots, journal-step crash recovery. - NOT covered: real-server breakage of dedup. The adversarial probe only checks that `.cas` is unreachable by path. ## Build 1. Real-server dedup campaign in `tests/adversarial/` (new section `dedup`), run against a real local server: - storms of identical and near-identical uploads from several users and sessions at once, including sync clients and WebDAV if present; - `kill -9` of the server at random points during upload/dedup/GC; after restart, every file must have the right bytes (BLAKE3 against the source) and journals must be empty; - delete / trash / restore / version / rename / move / share across users that share one blob; GC must never delete a blob that is still linked; no user sees another user's file; - quota: logical size per user stays exact through all of the above; - edit a deduped file through every write path (editor, upload overwrite, sync, WebDAV); no other user's copy may change (the server must never write in place into a shared inode); - an outside process changes a file in a home in place on disk (simulate a stray tool); detect and contain it, do not spread it to other links; - timing: prove no dedup timing oracle (upload of a known file vs unknown file has the same latency distribution within noise). 2. Bit-rot scrub: a scheduled low-priority job (idle IO class, nice) that re-hashes `.cas` entries against their names, rate-limited, resumable. On mismatch: quarantine the entry, record which user files are affected, raise an admin notice, and repair from any intact copy if one exists (for example a version or a non-deduped copy). Admin setting for the schedule; default weekly. Expose last scrub status in admin. 3. Fix every finding. Blocks: data loss/corruption, cross-user leak, crash. Regression test for each. 4. Update DESIGN §5 (ASD-STE100) with the scrub and the guarantees proven.
Author
Owner

Starting Forgejo #231 on branch job/dedup-break, based on dev at 1701cef5bd8936076522eb18039c3c8caed66299. I am tracing the existing calternal-fs dedup and maintenance paths, then I will add the real-server campaign and bit-rot scrub with regression coverage.

Starting Forgejo #231 on branch `job/dedup-break`, based on `dev` at `1701cef5bd8936076522eb18039c3c8caed66299`. I am tracing the existing calternal-fs dedup and maintenance paths, then I will add the real-server campaign and bit-rot scrub with regression coverage.
Author
Owner

Implementation decisions not set by DESIGN §5:

  • The default schedule is Sunday 03:00 UTC. The cron crate in this repo numbers Sunday as 1.
  • The scrub checks 64 Blob store entries per page and caps aggregate reads at 1 MiB/s.
  • A dedicated worker thread uses nice value 10 and Linux idle I/O priority. It stores resumable status in .system/dedup-scrub.json through Root.
  • Admins view status and set the schedule under Admin → Backups. A scrub with any finding shows attention, including findings that were repaired; the status lists affected, repaired and unrepaired paths.
Implementation decisions not set by DESIGN §5: - The default schedule is Sunday 03:00 UTC. The cron crate in this repo numbers Sunday as 1. - The scrub checks 64 Blob store entries per page and caps aggregate reads at 1 MiB/s. - A dedicated worker thread uses nice value 10 and Linux idle I/O priority. It stores resumable status in `.system/dedup-scrub.json` through Root. - Admins view status and set the schedule under Admin → Backups. A scrub with any finding shows attention, including findings that were repaired; the status lists affected, repaired and unrepaired paths.
Author
Owner

#231 implementation finished on job/dedup-break; HEAD 24db6fb2 (includes merge 2a31238c of current dev). The branch is pushed.

Built the real-server dedup breakage campaign, CAS bit-rot scrub and resumable checkpoints, low-priority scheduler and admin status/schedule/trigger API, Admin → Backups UI, OpenAPI/client updates, regression coverage, and the DESIGN §5 update. The scrub uses existing Root helpers for .cas access.

The post-merge real-server campaign ended with:

server alive at end: True

==== ROUND 2 FINDINGS 0

==== ROUND 2 SLOW 1
 - dedup GC trash orphans :: 7.2s status 200

The first rerun classified the oversized JSON request as no response because Python http.client raised BrokenPipe while still sending the body. The server had rejected the oversized body before the client read the response. The probe now reads an early HTTP response after a write error; the rerun above passed. The only reported item is SLOW GC load.

Files changed: Cargo.lock; apps/web/src/routes/settings/admin/AdminSection.svelte; apps/web/src/routes/settings/admin/BackupsGroup.svelte; contracts/openapi.json; crates/calternal-fs/src/blob.rs; crates/calternal-fs/src/lib.rs; crates/calternal-fs/src/root.rs; crates/calternal-fs/tests/storage.rs; crates/calternal-server/Cargo.toml; crates/calternal-server/src/wire.rs; docs/DESIGN.md; packages/api-client/src/generated.ts; tests/adversarial/attack2.py; tests/adversarial/run.sh.

Gates after merging dev:

  • cargo fmt --check: exit 0; no output.
  • cargo clippy --all-targets -- -D warnings: exit 0. Output: Finished \dev` profile [unoptimized + debuginfo] target(s) in 2m 48s`
  • cargo test: exit 0; 72 suites, 1259 passed, 0 failed, 12 ignored.
  • bun run --cwd apps/web check: exit 0. Output:
    Loading svelte-check in workspace: /home/kayg/Developer/calternal-wt/dedup-break/apps/web
    Getting Svelte diagnostics...
    
    svelte-check found 0 errors and 0 warnings
    
  • bun run --cwd apps/web test: exit 0:
    Test Files  80 passed (80)
         Tests  585 passed (585)
      Start at  17:05:53
      Duration  49.46s (transform 66%, import 13%, environment 11%, tests 7%, setup 2%)
    

Known gap: the server has no file WebDAV route; its DAV routes cover CalDAV, so WebDAV file-write dedup was not applicable. The campaign exercises upload overwrite, copy, rename, trash/restore, share, and sync import/recovery.

Decisions where DESIGN §5 was silent: default schedule 0 0 3 * * 1 (Sunday at 03:00); 64 CAS entries per page; 1 MiB/s pacing; process nice 10 plus idle I/O priority; private checkpoint .system/dedup-scrub.json; and scrub controls/status in Admin → Backups. Production-build screenshot captured at /home/kayg/Developer/calternal-wt/artifacts/dedup-break-admin-backups.png.

#231 implementation finished on `job/dedup-break`; HEAD `24db6fb2` (includes merge `2a31238c` of current `dev`). The branch is pushed. Built the real-server dedup breakage campaign, CAS bit-rot scrub and resumable checkpoints, low-priority scheduler and admin status/schedule/trigger API, Admin → Backups UI, OpenAPI/client updates, regression coverage, and the DESIGN §5 update. The scrub uses existing `Root` helpers for `.cas` access. The post-merge real-server campaign ended with: ```text server alive at end: True ==== ROUND 2 FINDINGS 0 ==== ROUND 2 SLOW 1 - dedup GC trash orphans :: 7.2s status 200 ``` The first rerun classified the oversized JSON request as no response because Python `http.client` raised `BrokenPipe` while still sending the body. The server had rejected the oversized body before the client read the response. The probe now reads an early HTTP response after a write error; the rerun above passed. The only reported item is SLOW GC load. Files changed: `Cargo.lock`; `apps/web/src/routes/settings/admin/AdminSection.svelte`; `apps/web/src/routes/settings/admin/BackupsGroup.svelte`; `contracts/openapi.json`; `crates/calternal-fs/src/blob.rs`; `crates/calternal-fs/src/lib.rs`; `crates/calternal-fs/src/root.rs`; `crates/calternal-fs/tests/storage.rs`; `crates/calternal-server/Cargo.toml`; `crates/calternal-server/src/wire.rs`; `docs/DESIGN.md`; `packages/api-client/src/generated.ts`; `tests/adversarial/attack2.py`; `tests/adversarial/run.sh`. Gates after merging `dev`: - `cargo fmt --check`: exit 0; no output. - `cargo clippy --all-targets -- -D warnings`: exit 0. Output: `Finished \`dev\` profile [unoptimized + debuginfo] target(s) in 2m 48s` - `cargo test`: exit 0; 72 suites, 1259 passed, 0 failed, 12 ignored. - `bun run --cwd apps/web check`: exit 0. Output: ```text Loading svelte-check in workspace: /home/kayg/Developer/calternal-wt/dedup-break/apps/web Getting Svelte diagnostics... svelte-check found 0 errors and 0 warnings ``` - `bun run --cwd apps/web test`: exit 0: ```text Test Files 80 passed (80) Tests 585 passed (585) Start at 17:05:53 Duration 49.46s (transform 66%, import 13%, environment 11%, tests 7%, setup 2%) ``` Known gap: the server has no file WebDAV route; its DAV routes cover CalDAV, so WebDAV file-write dedup was not applicable. The campaign exercises upload overwrite, copy, rename, trash/restore, share, and sync import/recovery. Decisions where DESIGN §5 was silent: default schedule `0 0 3 * * 1` (Sunday at 03:00); 64 CAS entries per page; 1 MiB/s pacing; process nice 10 plus idle I/O priority; private checkpoint `.system/dedup-scrub.json`; and scrub controls/status in Admin → Backups. Production-build screenshot captured at `/home/kayg/Developer/calternal-wt/artifacts/dedup-break-admin-backups.png`.
Author
Owner

Formatting correction to the clippy line in my previous report. The exact output was:

Finished `dev` profile [unoptimized + debuginfo] target(s) in 2m 48s

The exit status and all other gate results in the report are unchanged.

Formatting correction to the clippy line in my previous report. The exact output was: ```text Finished `dev` profile [unoptimized + debuginfo] target(s) in 2m 48s ``` The exit status and all other gate results in the report are unchanged.
Author
Owner

Starting the Admin → Backups UI rebuild and schedule representation on branch job/dedup-break, based on dev at 24db6fb2849268d9908b26021456304d4450d3d6. The backend campaign and scrub are accepted. I will reuse the existing Settings controls, render schedule values in the Instance timezone, and verify the real production build at desktop and phone sizes in light and dark themes, including a problem state.

Starting the Admin → Backups UI rebuild and schedule representation on branch `job/dedup-break`, based on `dev` at `24db6fb2849268d9908b26021456304d4450d3d6`. The backend campaign and scrub are accepted. I will reuse the existing Settings controls, render schedule values in the Instance timezone, and verify the real production build at desktop and phone sizes in light and dark themes, including a problem state.
Author
Owner

Correction to my start comment: 24db6fb2 is the continuation head, not the branch base. The original #231 report records base 1701cef5 and a merge of dev at 2a31238c. Current local dev is 895d3479; I will merge that before final gates as requested.

Correction to my start comment: `24db6fb2` is the continuation head, not the branch base. The original #231 report records base `1701cef5` and a merge of `dev` at `2a31238c`. Current local `dev` is `895d3479`; I will merge that before final gates as requested.
Author
Owner

Finding: BackupsGroup.svelte labels the scrub schedule UTC and exposes a raw cron field. CronScheduler::enqueue_due evaluates all expressions with DateTime<Utc>, and the default test checks only the literal string. GET /api/v1/admin/config has no Instance time-zone value. I will expose the server's detected IANA zone as the Instance time zone, parse the friendly frequency/day/time input into the internal cron expression on the server, and evaluate that expression in the same zone so the wall-clock schedule survives daylight-saving changes. The default will use the named SUN day token and a test of its actual next occurrence.

Finding: `BackupsGroup.svelte` labels the scrub schedule UTC and exposes a raw cron field. `CronScheduler::enqueue_due` evaluates all expressions with `DateTime<Utc>`, and the default test checks only the literal string. `GET /api/v1/admin/config` has no Instance time-zone value. I will expose the server's detected IANA zone as the Instance time zone, parse the friendly frequency/day/time input into the internal cron expression on the server, and evaluate that expression in the same zone so the wall-clock schedule survives daylight-saving changes. The default will use the named `SUN` day token and a test of its actual next occurrence.
Author
Owner

The accepted scrub status records affected Home paths but no stable item IDs. That cannot support the §33 deep-link rule: a path link can break on rename, while an Index lookup can preserve the item's identity. I added a parameterized, batched Files Index read that returns only existing IDs and does not assign an ID during a status read. Regression evidence: cargo test -p calternal-plugin-files passed (102 passed, 0 failed).

The accepted scrub status records affected Home paths but no stable item IDs. That cannot support the §33 deep-link rule: a path link can break on rename, while an Index lookup can preserve the item's identity. I added a parameterized, batched Files Index read that returns only existing IDs and does not assign an ID during a status read. Regression evidence: `cargo test -p calternal-plugin-files` passed (102 passed, 0 failed).
Author
Owner

The admin status API now returns only the fields the screen needs: Instance timezone, checked file count, byte count, completion time, and affected files with stable IDs and repair state. The existing checkpoint still stores hashes and scan state privately. The adversarial probe now checks that storage fields stay out of the response and verifies affected file links from the damage fixture. Server evidence: cargo test -p calternal-server passed (46 passed, 2 ignored).

The admin status API now returns only the fields the screen needs: Instance timezone, checked file count, byte count, completion time, and affected files with stable IDs and repair state. The existing checkpoint still stores hashes and scan state privately. The adversarial probe now checks that storage fields stay out of the response and verifies affected file links from the damage fixture. Server evidence: `cargo test -p calternal-server` passed (46 passed, 2 ignored).
Author
Owner

Web gate finding, fixed: the first full bun run test reported 1 failure in shared date formatter reuse gate because BackupsGroup.svelte used Intl.DateTimeFormat directly. A new timezone boundary check also showed formatWeekday(2026-09-26T22:30:00Z, ..., "UTC") returned Sun on the Europe/Berlin host. The UI now uses the shared cached weekday and time formatters, and the shared weekday formatter accepts an explicit timezone. The focused test passes (14/14), bun run check reports 0 errors and 0 warnings, and the rerun of the full web suite passes (81 files, 590 tests).

Web gate finding, fixed: the first full `bun run test` reported 1 failure in `shared date formatter reuse gate` because `BackupsGroup.svelte` used `Intl.DateTimeFormat` directly. A new timezone boundary check also showed `formatWeekday(2026-09-26T22:30:00Z, ..., "UTC")` returned `Sun` on the Europe/Berlin host. The UI now uses the shared cached weekday and time formatters, and the shared weekday formatter accepts an explicit timezone. The focused test passes (14/14), `bun run check` reports 0 errors and 0 warnings, and the rerun of the full web suite passes (81 files, 590 tests).
Author
Owner

Final report

Head: 2cb1ced36643fcf60f4c3472c627d422e0649409 on job/dedup-break. Pushed to origin/job/dedup-break (Everything up-to-date). Merged the latest dev at 3f0f6f64 before the final gates.

Built

  • Admin → Backups now uses the settings controls for Weekly / Daily / Off, a weekday selector for weekly checks, and a time input. It shows the Instance timezone. The server converts the friendly schedule to cron for internal storage. The default is Sunday at 03:00 in the Instance timezone.
  • The “File integrity” group uses plain language, a readable last-check line, and affected-file links that use stable Item IDs. Back up now and Check now use the settings button style.
  • The status line uses the shared cached weekday and time formatters. The weekday formatter accepts an explicit timezone.
  • The real-server dedup probe ended with 0 findings and 0 slow findings. The post-merge Event tag probe passed its Unicode/bidi, 65,536-byte category, malformed-input, and parallel-read checks.

Files for this issue

  • apps/web/src/routes/settings/admin/AdminSection.svelte
  • apps/web/src/routes/settings/admin/BackupsGroup.svelte
  • apps/web/src/lib/time.test.ts, packages/ui/src/time.ts
  • crates/calternal-db/src/cron.rs, crates/calternal-server/src/wire.rs, crates/plugins/files/src/lib.rs
  • contracts/openapi.json, packages/api-client/src/generated.ts, Cargo.lock
  • tests/adversarial/attack2.py, tests/adversarial/run.sh

Production screenshots

All four screenshots show the persisted “problems found” state and are attached to this issue. They are in the ignored artifacts/ directory and are not committed.

Desktop light, problems found
Desktop dark, problems found
Phone light, problems found
Phone dark, problems found

Gates

$ cargo fmt --check
(no output; exit 0)

$ cargo clippy --all-targets -- -D warnings
Finished `dev` profile [unoptimized + debuginfo] target(s) in 19m 55s

$ cargo test
test result: ok. 10 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 5.72s
test result: ok. 106 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 54.28s
test result: ok. 49 passed; 0 failed; 2 ignored; 0 measured; 0 filtered out; finished in 10.81s
Aggregate: 1,290 passed; 0 failed; 12 ignored across 72 test-result summaries.

$ bun run check
svelte-check found 0 errors and 0 warnings

$ bun run test
 Test Files  81 passed (81)
      Tests  590 passed (590)

The web test run also printed Not implemented: Window's scrollTo() method from the test DOM environment; Vitest exited 0.

Known gap

If a problem path has no existing Files Index Item ID, the status shows the path without a link. It does not create an ID or use a rename-sensitive path as a link.

Decisions not stated in DESIGN.md

  • The Instance timezone comes from the server host’s detected IANA timezone. It falls back to UTC if detection fails; this branch has no separate Instance timezone setting.
  • The status line uses a short English weekday and the user’s selected clock style, formatted in the Instance timezone.
## Final report Head: `2cb1ced36643fcf60f4c3472c627d422e0649409` on `job/dedup-break`. Pushed to `origin/job/dedup-break` (`Everything up-to-date`). Merged the latest `dev` at `3f0f6f64` before the final gates. ### Built - Admin → Backups now uses the settings controls for Weekly / Daily / Off, a weekday selector for weekly checks, and a time input. It shows the Instance timezone. The server converts the friendly schedule to cron for internal storage. The default is Sunday at 03:00 in the Instance timezone. - The “File integrity” group uses plain language, a readable last-check line, and affected-file links that use stable Item IDs. Back up now and Check now use the settings button style. - The status line uses the shared cached weekday and time formatters. The weekday formatter accepts an explicit timezone. - The real-server dedup probe ended with 0 findings and 0 slow findings. The post-merge Event tag probe passed its Unicode/bidi, 65,536-byte category, malformed-input, and parallel-read checks. ### Files for this issue - `apps/web/src/routes/settings/admin/AdminSection.svelte` - `apps/web/src/routes/settings/admin/BackupsGroup.svelte` - `apps/web/src/lib/time.test.ts`, `packages/ui/src/time.ts` - `crates/calternal-db/src/cron.rs`, `crates/calternal-server/src/wire.rs`, `crates/plugins/files/src/lib.rs` - `contracts/openapi.json`, `packages/api-client/src/generated.ts`, `Cargo.lock` - `tests/adversarial/attack2.py`, `tests/adversarial/run.sh` ### Production screenshots All four screenshots show the persisted “problems found” state and are attached to this issue. They are in the ignored `artifacts/` directory and are not committed. ![Desktop light, problems found](https://git.kayg.org/attachments/3b6c8bbc-4dd3-448f-a62a-4542f9b8d776) ![Desktop dark, problems found](https://git.kayg.org/attachments/22e0cdab-3538-4671-a843-d07a2cad4d34) ![Phone light, problems found](https://git.kayg.org/attachments/1ee3833f-e88e-4ea0-a853-fc120893685c) ![Phone dark, problems found](https://git.kayg.org/attachments/a69c03fa-2a96-4002-b64f-9e7b7a74f767) ### Gates ```text $ cargo fmt --check (no output; exit 0) $ cargo clippy --all-targets -- -D warnings Finished `dev` profile [unoptimized + debuginfo] target(s) in 19m 55s $ cargo test test result: ok. 10 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 5.72s test result: ok. 106 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 54.28s test result: ok. 49 passed; 0 failed; 2 ignored; 0 measured; 0 filtered out; finished in 10.81s Aggregate: 1,290 passed; 0 failed; 12 ignored across 72 test-result summaries. $ bun run check svelte-check found 0 errors and 0 warnings $ bun run test Test Files 81 passed (81) Tests 590 passed (590) ``` The web test run also printed `Not implemented: Window's scrollTo() method` from the test DOM environment; Vitest exited 0. ### Known gap If a problem path has no existing Files Index Item ID, the status shows the path without a link. It does not create an ID or use a rename-sensitive path as a link. ### Decisions not stated in DESIGN.md - The Instance timezone comes from the server host’s detected IANA timezone. It falls back to UTC if detection fails; this branch has no separate Instance timezone setting. - The status line uses a short English weekday and the user’s selected clock style, formatted in the Instance timezone.
Author
Owner

Merged in 19e65b22 (test-name conflict resolved; Files tests 110/110; generated client OK). The real-server dedup campaign, scrub and readable Admin → Backups UI are in dev.

Merged in 19e65b22 (test-name conflict resolved; Files tests 110/110; generated client OK). The real-server dedup campaign, scrub and readable Admin → Backups UI are in dev.
kayg referenced this issue from a commit 2026-09-27 18:39:21 +00:00
kayg closed this issue 2026-09-27 18:39:21 +00:00
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#231
No description provided.