Adversarial dedup crash/GC probes cannot finish in the full runner (scrub pacing vs shared CAS size) #972

Open
opened 2026-10-03 05:54:43 +00:00 by kayg · 4 comments
Owner

Found by the merge-round-7a full adversarial runs (#427, round 5). Test-infrastructure gap; no product data loss was seen.

Problem

In the full tests/adversarial/run.sh, the round-2 dedup crash and GC probes (attack2.py section dedup) cannot finish. The CAS scrub is paced at 1 MiB/s by design (CAS_SCRUB_BYTES_PER_SECOND in crates/calternal-server/src/wire.rs). After round 1 and the earlier round-2 sections the shared server holds about 100 MB in .cas, so one scrub pass takes minutes. The probe waits 120 s for the collector to start deleting its injected garbage and 180 s for the resumed scrub, then reports:

dedup scrub repair :: scrub did not report repaired corruption: {'phase': 'scanning', ... 'checked_files': 64, 'bytes_checked': 105082960, 'problems': [...]}
dedup GC crash :: the collector did not begin deleting CAS garbage
dedup GC crash :: the server restart coordinator did not resume
dedup GC resume :: scrub did not complete after restart: {'phase': 'scanning', 'checked_files': 128, ...}
dedup GC resume :: orphan blob <hash> remained after collection resumed   (x32)

The scrub is progressing (checked_files 64 to 128, no has_error), and the orphans remain only because collection never started inside the window. Focused ADVERSARIAL_DEDUP_ONLY=1 runs start with a small CAS.

Related product fix in this round: before commit 05979db32 the scrub also stopped permanently with CAS scrub scan failed: invalid relative path when a Home held a name with a control character; that is fixed with a regression test (blob_scrub_repair_skips_unaddressable_names_in_homes).

Expected

The full runner exercises the dedup crash and GC paths to completion. Options: run the dedup section on its own fresh data directory (as ADVERSARIAL_DEDUP_ONLY does), or add a test-only scrub pacing override that the runner sets, or make the probe wait in proportion to the CAS size it measures.

Regression

tests/adversarial/run.sh on merge-round-7a reports no dedup findings while the rest of round 2 runs on the shared server.

Found by the merge-round-7a full adversarial runs (#427, round 5). Test-infrastructure gap; no product data loss was seen. ## Problem In the full `tests/adversarial/run.sh`, the round-2 dedup crash and GC probes (`attack2.py` section `dedup`) cannot finish. The CAS scrub is paced at 1 MiB/s by design (`CAS_SCRUB_BYTES_PER_SECOND` in `crates/calternal-server/src/wire.rs`). After round 1 and the earlier round-2 sections the shared server holds about 100 MB in `.cas`, so one scrub pass takes minutes. The probe waits 120 s for the collector to start deleting its injected garbage and 180 s for the resumed scrub, then reports: ``` dedup scrub repair :: scrub did not report repaired corruption: {'phase': 'scanning', ... 'checked_files': 64, 'bytes_checked': 105082960, 'problems': [...]} dedup GC crash :: the collector did not begin deleting CAS garbage dedup GC crash :: the server restart coordinator did not resume dedup GC resume :: scrub did not complete after restart: {'phase': 'scanning', 'checked_files': 128, ...} dedup GC resume :: orphan blob <hash> remained after collection resumed (x32) ``` The scrub is progressing (`checked_files` 64 to 128, no `has_error`), and the orphans remain only because collection never started inside the window. Focused `ADVERSARIAL_DEDUP_ONLY=1` runs start with a small CAS. Related product fix in this round: before commit 05979db32 the scrub also stopped permanently with `CAS scrub scan failed: invalid relative path` when a Home held a name with a control character; that is fixed with a regression test (`blob_scrub_repair_skips_unaddressable_names_in_homes`). ## Expected The full runner exercises the dedup crash and GC paths to completion. Options: run the dedup section on its own fresh data directory (as `ADVERSARIAL_DEDUP_ONLY` does), or add a test-only scrub pacing override that the runner sets, or make the probe wait in proportion to the CAS size it measures. ## Regression `tests/adversarial/run.sh` on merge-round-7a reports no dedup findings while the rest of round 2 runs on the shared server.
Author
Owner

Merge-round 7a dedup recovery evidence: after the probe corrupted a shared CAS blob and triggered the scrub, the 90-second poll ended while status remained phase=scanning (checked_files=64, bytes_checked=107053213). The affected-file status still had repaired=false; the probe reported that scrub did not report repaired corruption. This run had high shared-host load. The CAS repair was not confirmed within the probe bound; see #427's full adversarial log target/tmp/verify-7a-adversarial.log.

Merge-round 7a dedup recovery evidence: after the probe corrupted a shared CAS blob and triggered the scrub, the 90-second poll ended while status remained `phase=scanning` (`checked_files=64`, `bytes_checked=107053213`). The affected-file status still had `repaired=false`; the probe reported that scrub did not report repaired corruption. This run had high shared-host load. The CAS repair was not confirmed within the probe bound; see #427's full adversarial log `target/tmp/verify-7a-adversarial.log`.
Author
Owner

The same run's dedup/GC phase also timed out both POST /api/v1/files/trash for 32 newly uploaded orphan files and POST /api/v1/files/trash/empty after the 32 CAS orphans were created. The run then refreshed the owner assertion and continued toward the GC restart probe. These are non-SLOW request failures during a high-load run; full output remains in #427's worktree log.

The same run's dedup/GC phase also timed out both `POST /api/v1/files/trash` for 32 newly uploaded orphan files and `POST /api/v1/files/trash/empty` after the 32 CAS orphans were created. The run then refreshed the owner assertion and continued toward the GC restart probe. These are non-SLOW request failures during a high-load run; full output remains in #427's worktree log.
Author
Owner

Continuation of the merge-round 7a dedup/GC probe: after the two orphan trash requests timed out, the scrub trigger did not begin deleting any of the 512 marker entries within its 120-second wait (dedup GC crash: the collector did not begin deleting CAS garbage). This prevented the restart barrier from being reached. The subsequent bounded status poll is still running.

Continuation of the merge-round 7a dedup/GC probe: after the two orphan trash requests timed out, the scrub trigger did not begin deleting any of the 512 marker entries within its 120-second wait (`dedup GC crash: the collector did not begin deleting CAS garbage`). This prevented the restart barrier from being reached. The subsequent bounded status poll is still running.
Author
Owner

Final result of the merge-round 7a dedup crash/GC probe: after restart, the scrub still reported phase=scanning at its 180-second bound (checked_files=128, bytes_checked=111007156), with the damaged shared file still repaired=false. The probe found 32 test orphan blobs plus CAS marker garbage remained. The issue is a data-integrity/collection follow-up; no unrelated cross-User content change was seen. The shared host remained heavily loaded. Full details are in #427's local run log.

Final result of the merge-round 7a dedup crash/GC probe: after restart, the scrub still reported `phase=scanning` at its 180-second bound (`checked_files=128`, `bytes_checked=111007156`), with the damaged shared file still `repaired=false`. The probe found 32 test orphan blobs plus CAS marker garbage remained. The issue is a data-integrity/collection follow-up; no unrelated cross-User content change was seen. The shared host remained heavily loaded. Full details are in #427's local run log.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#972
No description provided.