Adversarial attack2-only run uses deleted users and skips restart coordination #285

Closed
opened 2026-09-28 03:02:03 +00:00 by kayg · 8 comments
Owner

Evidence

The single bounded API round used ADVERSARIAL_ATTACK2_ONLY=1. It reported late dedup failures that came from test setup and orchestration:

  • attack2.py runs deletion_probes() before section("dedup"). The deletion probe removes the invited fixture Users. The later dedup section still uses those Users. Their file calls returned 401 Authentication required, and the share call returned 400 invalid path or request; the dependent byte and scrub assertions then failed.
  • The attack2-only branch in tests/adversarial/run.sh runs attack2.py without restart_on_request. The dedup probe waits for restart markers 4, 5, and 6. The run reported server restart coordinator did not resume, then the interrupted upload was visible and could not be resumed.

The server stayed alive. These results do not establish a Files isolation or crash-recovery defect. The dedicated ADVERSARIAL_DEDUP_ONLY=1 branch provisions the restart coordinator, but it was not run because this job was limited to one adversarial round. Please move dedup before destructive user deletion or provision fresh users, and keep restart probes on the coordinated branch.

## Evidence The single bounded API round used `ADVERSARIAL_ATTACK2_ONLY=1`. It reported late `dedup` failures that came from test setup and orchestration: - `attack2.py` runs `deletion_probes()` before `section("dedup")`. The deletion probe removes the invited fixture Users. The later dedup section still uses those Users. Their file calls returned `401 Authentication required`, and the share call returned `400 invalid path or request`; the dependent byte and scrub assertions then failed. - The attack2-only branch in `tests/adversarial/run.sh` runs `attack2.py` without `restart_on_request`. The dedup probe waits for restart markers 4, 5, and 6. The run reported `server restart coordinator did not resume`, then the interrupted upload was visible and could not be resumed. The server stayed alive. These results do not establish a Files isolation or crash-recovery defect. The dedicated `ADVERSARIAL_DEDUP_ONLY=1` branch provisions the restart coordinator, but it was not run because this job was limited to one adversarial round. Please move dedup before destructive user deletion or provision fresh users, and keep restart probes on the coordinated branch.
Author
Owner

Gates on dev f9c0a066 (2026-09-28): the full adversarial round reported 186 findings, and nearly all are this harness bug. Every dedup check (campaign folder, storm, download, quota, overwrite read, scrub member, crash publish) got 401 unauthenticated, so the follow-on assertions ('completed upload bytes differ', 'overwrite isolation: User … changed after another User overwrote a linked file', 'link count expected 13 found 4', 'crash upload visibility 404 vs 200', 'scrub did not report repaired corruption') are unverified, not proven wrong. This is serious because cross-user dedup isolation is a data-integrity guarantee and the probe currently cannot test it. Priority: raised. Fix the runner so each dedup phase uses live users with fresh sessions (re-authenticate before each phase, create users per phase, never reuse users deleted by the purge phase; fold in #272's late-dedup preparation). Then rerun and report every dedup finding that remains as a real bug with a Rust regression test. Other non-dedup findings in the same run (for their owners): POST /api/v1/admin/search/rebuild timed out (#258 step C), appearance oversized body 502 (#282/#283), write during user purge took 13 s (expected <2 s).

Gates on dev f9c0a066 (2026-09-28): the full adversarial round reported **186 findings, and nearly all are this harness bug**. Every dedup check (campaign folder, storm, download, quota, overwrite read, scrub member, crash publish) got **401 unauthenticated**, so the follow-on assertions ('completed upload bytes differ', 'overwrite isolation: User … changed after another User overwrote a linked file', 'link count expected 13 found 4', 'crash upload visibility 404 vs 200', 'scrub did not report repaired corruption') are **unverified, not proven wrong**. This is serious because **cross-user dedup isolation is a data-integrity guarantee and the probe currently cannot test it.** Priority: raised. Fix the runner so each dedup phase uses live users with fresh sessions (re-authenticate before each phase, create users per phase, never reuse users deleted by the purge phase; fold in #272's late-dedup preparation). Then rerun and report every dedup finding that remains as a real bug with a Rust regression test. Other non-dedup findings in the same run (for their owners): `POST /api/v1/admin/search/rebuild` timed out (#258 step C), appearance oversized body 502 (#282/#283), write during user purge took 13 s (expected <2 s).
Author
Owner

Starting work on job/adv-harness at f0c13fdb7977939bceb2b3a39cc59e7ce9a201fa; current dev base is 6c1e5062c781feb87a85a70825a69661239bc855.

I read #285 and #272, including their comments. The full run deletes fixture Users B–D before the dedup campaign, and the full-run shell does not coordinate dedup restart barriers 4–6. I will run non-dedup round-2 probes first, reset the local server and auth limiter, provision a live dedup fixture set with fresh installation sessions, and then run the dedup section with all required restart barriers coordinated. I will report any remaining dedup finding with evidence and add a Rust regression test plus fix for any real calternal-fs/dedup defect.

Starting work on `job/adv-harness` at `f0c13fdb7977939bceb2b3a39cc59e7ce9a201fa`; current `dev` base is `6c1e5062c781feb87a85a70825a69661239bc855`. I read #285 and #272, including their comments. The full run deletes fixture Users B–D before the dedup campaign, and the full-run shell does not coordinate dedup restart barriers 4–6. I will run non-dedup round-2 probes first, reset the local server and auth limiter, provision a live dedup fixture set with fresh installation sessions, and then run the dedup section with all required restart barriers coordinated. I will report any remaining dedup finding with evidence and add a Rust regression test plus fix for any real calternal-fs/dedup defect.
Author
Owner

The first local round stopped before round 2. run.sh timed out after 180 seconds waiting for restart-ready-4, while search_chaos.py was still in its watcher/indexing phase. The probe can spend up to 240 seconds waiting for each of three indexed markers before it requests that restart. I increased the coordinator window to 30 minutes to cover those bounded waits under load; this was a harness timeout, not an API or dedup finding. I will now run the corrected full round once.

The first local round stopped before round 2. `run.sh` timed out after 180 seconds waiting for `restart-ready-4`, while `search_chaos.py` was still in its watcher/indexing phase. The probe can spend up to 240 seconds waiting for each of three indexed markers before it requests that restart. I increased the coordinator window to 30 minutes to cover those bounded waits under load; this was a harness timeout, not an API or dedup finding. I will now run the corrected full round once.
Author
Owner

The full round reached the late fixture step: it provisioned an Owner session and three fresh member Users after deletion. The subsequent dedup section got proxy 502 responses before its normal campaign requests. The retained work directory had restart-ready-4 and restart-go-4 from the earlier search-chaos restart, and dedup also used barrier 4 for its interrupted upload. The full-run coordinator therefore consumed the stale search marker and restarted the server at the start of dedup. This makes the 125 dedup findings from that phase setup/restart cascades, not evidence of a dedup storage defect. I moved dedup crash barriers to 5–7, leaving search chaos on 4, and will run the dedicated dedup campaign with the corrected barriers.

The full round reached the late fixture step: it provisioned an Owner session and three fresh member Users after deletion. The subsequent dedup section got proxy 502 responses before its normal campaign requests. The retained work directory had `restart-ready-4` and `restart-go-4` from the earlier search-chaos restart, and dedup also used barrier 4 for its interrupted upload. The full-run coordinator therefore consumed the stale search marker and restarted the server at the start of dedup. This makes the 125 dedup findings from that phase setup/restart cascades, not evidence of a dedup storage defect. I moved dedup crash barriers to 5–7, leaving search chaos on 4, and will run the dedicated dedup campaign with the corrected barriers.
Author
Owner

The corrected dedup campaign passed on the local server: server alive at end: True, ==== ROUND 2 FINDINGS 0, and one SLOW entry (dedup GC trash orphans, 5.5s, HTTP 200). Its timing check measured known-content uploads at 0.539s and new-content uploads at 0.721s (ratio 1.34). No dedup storage finding remains, so no calternal-fs regression test was needed.

The broad full round did record non-dedup failures that need owner triage: 128 concurrent /api/v1/files/recent?limit=500 requests timed out at 30s; 17 of 24 Journal PATCH race requests timed out, with later Journal and Bookmark requests also timing out; the Calendar burst expected 120 Photos and saw 0; Bookmark Calendar projection timed out and its Note was missing on the saved date; the auth matrix fixture upload returned no response; and the oversized appearance request got proxy 502. The deletion availability check also measured a 14.91s upload during another User's purge against its 2s limit. The server was alive at the end of round 1. These are not dedup findings and remain unassigned here.

The first broad run's dedup results were inconclusive because search chaos and dedup reused restart barrier 4. Dedup now uses barriers 5–7. No request expectation was changed.

The corrected dedup campaign passed on the local server: `server alive at end: True`, `==== ROUND 2 FINDINGS 0`, and one SLOW entry (`dedup GC trash orphans`, 5.5s, HTTP 200). Its timing check measured known-content uploads at 0.539s and new-content uploads at 0.721s (ratio 1.34). No dedup storage finding remains, so no calternal-fs regression test was needed. The broad full round did record non-dedup failures that need owner triage: 128 concurrent `/api/v1/files/recent?limit=500` requests timed out at 30s; 17 of 24 Journal PATCH race requests timed out, with later Journal and Bookmark requests also timing out; the Calendar burst expected 120 Photos and saw 0; Bookmark Calendar projection timed out and its Note was missing on the saved date; the auth matrix fixture upload returned no response; and the oversized appearance request got proxy 502. The deletion availability check also measured a 14.91s upload during another User's purge against its 2s limit. The server was alive at the end of round 1. These are not dedup findings and remain unassigned here. The first broad run's dedup results were inconclusive because search chaos and dedup reused restart barrier 4. Dedup now uses barriers 5–7. No request expectation was changed.
Author
Owner

The full #303 adversarial run reproduced the stale restart-marker setup described here. The dedup section's wait_for_restart(4) returned using an old restart-go-4 marker from the earlier search-rebuild restart; no new server restart occurred before the incomplete-upload assertions. Evidence: restart-ready-4 was touched at 12:28:27, while restart-go-4 was last updated at 11:24:19. The round also reached dedup after deleting/expiring its fixture users, so the calls returned 401. The resulting dedup crash/isolation failures are invalid probes, not evidence of a product defect. The final restart probe reported zero findings.

The full #303 adversarial run reproduced the stale restart-marker setup described here. The dedup section's `wait_for_restart(4)` returned using an old `restart-go-4` marker from the earlier search-rebuild restart; no new server restart occurred before the incomplete-upload assertions. Evidence: `restart-ready-4` was touched at 12:28:27, while `restart-go-4` was last updated at 11:24:19. The round also reached dedup after deleting/expiring its fixture users, so the calls returned 401. The resulting dedup crash/isolation failures are invalid probes, not evidence of a product defect. The final restart probe reported zero findings.
Author
Owner

Final report: #285 and #272

Branch: job/adv-harness; base dev SHA: 6c1e5062c781feb87a85a70825a69661239bc855; merged dev once in 92410adf; final HEAD: 781b14c13f64e41ff59ba3562a9a889f2701d45e.

Changes

  • Added a late dedup preparation phase that logs the persisted Owner in again, issues fresh installation sessions, and replaces member fixtures with three fresh Users. This avoids sessions and Users invalidated by earlier deletion probes.
  • Split full-round non-dedup and dedup attack2 sections around a server restart. The runner now coordinates dedup restart barriers 5, 6, and 7, separate from Search chaos barrier 4. Increased the Search restart-marker wait to 30 minutes because its probe waits for three markers, each with a four-minute window.
  • Files: tests/adversarial/dedup_setup.mjs, tests/adversarial/attack2.py, tests/adversarial/run.sh.

Adversarial evidence

The corrected dedicated dedup campaign completed with a live server and ==== ROUND 2 FINDINGS 0. Known/new timing medians were 0.539s/0.721s (ratio 1.34). The only ROUND 2 SLOW entry was dedup GC trash orphans at 5.5s with status 200. No dedup finding remains, so no calternal-fs/dedup regression was needed.

The full broad round completed but reported non-dedup failures; SLOW-only latency markers are excluded. Evidence: 128 concurrent Files recent requests timed out; 17/24 Journal PATCH requests timed out and follow-up Journal operations also timed out; the Calendar burst expected 120 Photos and saw 0; Bookmark Calendar projections timed out and a saved Note was absent on its date; PDF text search and a semantic recall check missed expected results; Appearance probes returned proxy 502 local adversarial server is unavailable; the browser editor undo/redo equality check changed text spacing; the authz matrix upload fixture returned -1 before the authorization assertion (inconclusive); and upload during User purge took 14.91s against the 2s probe limit. Related existing trackers: #267, #269, #271, #283, #314, #265, #51, #58, and #59. These are outside the dedup scope and remain open for their owners.

Gates

  • cargo fmt --check: exit 0, no output.
  • cargo clippy --all-targets -- -D warnings: exit 0. Verbatim output:
    Finished `dev` profile [unoptimized + debuginfo] target(s) in 40m 11s
  • cargo test: exit 0. Final captured doc-test output (verbatim):
running 0 tests

test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s
  • bun run check: exit 0. Verbatim result: svelte-check found 0 errors and 0 warnings.
  • bun run test: exit 0. Verbatim summary:
 Test Files  107 passed (107)
      Tests  701 passed (701)
   Start at  12:53:38
   Duration  145.24s (transform 59%, environment 16%, import 13%, tests 9%, setup 3%)

cargo clean removed 41,301 files / 13.4 GiB. Generated apps/web/build and apps/web/.svelte-kit output was removed. The job branch was pushed; git push origin job/adv-harness reported Everything up-to-date.

Decisions not specified in DESIGN

I prepared dedup fixtures after destructive probes so the roster and sessions are live, assigned barriers 5–7 to dedup while preserving barrier 4 for Search chaos, and set the Search crash-barrier wait to 30 minutes to cover the three 240-second marker windows.

## Final report: #285 and #272 Branch: `job/adv-harness`; base dev SHA: `6c1e5062c781feb87a85a70825a69661239bc855`; merged dev once in `92410adf`; final HEAD: `781b14c13f64e41ff59ba3562a9a889f2701d45e`. ### Changes - Added a late dedup preparation phase that logs the persisted Owner in again, issues fresh installation sessions, and replaces member fixtures with three fresh Users. This avoids sessions and Users invalidated by earlier deletion probes. - Split full-round non-dedup and dedup attack2 sections around a server restart. The runner now coordinates dedup restart barriers 5, 6, and 7, separate from Search chaos barrier 4. Increased the Search restart-marker wait to 30 minutes because its probe waits for three markers, each with a four-minute window. - Files: `tests/adversarial/dedup_setup.mjs`, `tests/adversarial/attack2.py`, `tests/adversarial/run.sh`. ### Adversarial evidence The corrected dedicated dedup campaign completed with a live server and `==== ROUND 2 FINDINGS 0`. Known/new timing medians were 0.539s/0.721s (ratio 1.34). The only ROUND 2 SLOW entry was `dedup GC trash orphans` at 5.5s with status 200. No dedup finding remains, so no `calternal-fs/dedup` regression was needed. The full broad round completed but reported non-dedup failures; SLOW-only latency markers are excluded. Evidence: 128 concurrent Files recent requests timed out; 17/24 Journal PATCH requests timed out and follow-up Journal operations also timed out; the Calendar burst expected 120 Photos and saw 0; Bookmark Calendar projections timed out and a saved Note was absent on its date; PDF text search and a semantic recall check missed expected results; Appearance probes returned proxy 502 `local adversarial server is unavailable`; the browser editor undo/redo equality check changed text spacing; the authz matrix upload fixture returned -1 before the authorization assertion (inconclusive); and upload during User purge took 14.91s against the 2s probe limit. Related existing trackers: #267, #269, #271, #283, #314, #265, #51, #58, and #59. These are outside the dedup scope and remain open for their owners. ### Gates - `cargo fmt --check`: exit 0, no output. - `cargo clippy --all-targets -- -D warnings`: exit 0. Verbatim output: ```text Finished `dev` profile [unoptimized + debuginfo] target(s) in 40m 11s ``` - `cargo test`: exit 0. Final captured doc-test output (verbatim): ```text running 0 tests test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s ``` - `bun run check`: exit 0. Verbatim result: `svelte-check found 0 errors and 0 warnings`. - `bun run test`: exit 0. Verbatim summary: ```text Test Files 107 passed (107) Tests 701 passed (701) Start at 12:53:38 Duration 145.24s (transform 59%, environment 16%, import 13%, tests 9%, setup 3%) ``` `cargo clean` removed 41,301 files / 13.4 GiB. Generated `apps/web/build` and `apps/web/.svelte-kit` output was removed. The job branch was pushed; `git push origin job/adv-harness` reported `Everything up-to-date`. ### Decisions not specified in DESIGN I prepared dedup fixtures after destructive probes so the roster and sessions are live, assigned barriers 5–7 to dedup while preserving barrier 4 for Search chaos, and set the Search crash-barrier wait to 30 minutes to cover the three 240-second marker windows.
Author
Owner

Verification qualification

The full broad run described above used the runner before commit 781b14c1. Its dedup coordinator reused Search chaos barrier 4, saw the already-present restart-ready-4, and restarted the server too early. That made the dedup portion of that broad run inconclusive and caused proxy 502s. Commit 781b14c1 moved dedup restart barriers to 5, 6, and 7. The corrected dedicated dedup campaign then completed with ROUND 2 FINDINGS 0 and a live server. I did not repeat the combined full round after this correction, so the final full-mode handoff across Search chaos into dedup remains unverified. The broad run non-dedup evidence and the corrected standalone dedup result should be read separately.

## Verification qualification The full broad run described above used the runner before commit `781b14c1`. Its dedup coordinator reused Search chaos barrier 4, saw the already-present `restart-ready-4`, and restarted the server too early. That made the dedup portion of that broad run inconclusive and caused proxy 502s. Commit `781b14c1` moved dedup restart barriers to 5, 6, and 7. The corrected dedicated dedup campaign then completed with `ROUND 2 FINDINGS 0` and a live server. I did not repeat the combined full round after this correction, so the final full-mode handoff across Search chaos into dedup remains unverified. The broad run non-dedup evidence and the corrected standalone dedup result should be read separately.
kayg closed this issue 2026-09-28 11:05:52 +00:00
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#285
No description provided.