Sync: adversarial round 2 reported a lost concurrent edit and file/folder flips under heavy load #88

Closed
opened 2026-09-25 13:33:23 +00:00 by kayg · 2 comments
Owner

During the brand-assets job (2026-09-25) tests/adversarial/run.sh round 2 reported 5 sync-client findings (file/folder flips, a lost concurrent edit) while the build VM ran 3 Codex jobs (load ~10-15). Two reruns of the same code at lower load (4-8) were clean, and dev was clean twice at 41d12bd. A lost edit only under load is still a sync collision, which blocks merges by the owner's rule, so treat it as a real race until proven otherwise.

Do: rerun the sync section of round 2 in a loop under synthetic CPU load (e.g. stress-ng or parallel cargo builds) until it reproduces; capture the exact finding text and client/server logs; find the ordering assumption (timeouts, mtime granularity, rename detection window, conflict detection); fix with a regression test. Related: #75 (sync lost update, fixed), #46 (random sync p95).

During the brand-assets job (2026-09-25) tests/adversarial/run.sh round 2 reported 5 sync-client findings (file/folder flips, a lost concurrent edit) while the build VM ran 3 Codex jobs (load ~10-15). Two reruns of the same code at lower load (4-8) were clean, and dev was clean twice at 41d12bd. A lost edit only under load is still a sync collision, which blocks merges by the owner's rule, so treat it as a real race until proven otherwise. Do: rerun the sync section of round 2 in a loop under synthetic CPU load (e.g. stress-ng or parallel cargo builds) until it reproduces; capture the exact finding text and client/server logs; find the ordering assumption (timeouts, mtime granularity, rename detection window, conflict detection); fix with a regression test. Related: #75 (sync lost update, fixed), #46 (random sync p95).
Author
Owner

Result (branch job/sync-load, not merged)

Reproduced. It is a real bug, and load is not the cause. On the first loop run at load 8, round 2 gave the same 5 sync findings (file→directory, folder collision setup, remote file→local folder, changing file, concurrent edit). The daemon log showed one error on every pass, from the flip step to the end of the run:

sync reconcile deferred: server returned HTTP 409

A focused loop of only the flip step failed on iteration 1 or 2 on the old code.

Root cause

The local file identity was dev:ino. File systems reuse a freed inode number at once. After rm flip; mkdir flip; echo > flip/inner.txt, the new flip/inner.txt sometimes got the inode number of the deleted flip. follow_renames took that as a rename flip -> flip/inner.txt and asked the server to move the file into a folder with its own name. The server refused (mkdir over a file = 409), and the daemon sent the same move again on every pass. The pair stopped syncing. The 4 later findings were effects of this: the local edit in "concurrent edit" was still on disk but was never synced. Load only changes how often the inode number is reused.

Fix

  • 8c22c73 fix(sync): the local id is dev:ino:birth (birth time from statx / st_birthtime). same_file_identity still matches journals that have dev:ino ids, so an upgrade does not make spurious conflicts.
  • 71511c5 fix(sync): a file never moves into a path below itself. If the server refuses a followed rename (409/404/412), the engine plans each path separately and does not retry the move. A new remote file whose local path is a folder goes to a conflict name. Before, the pair also stalled on every pass in this case.
  • 7bc5a24 test(sync): crates/calternal-sync/type_flip_campaign.py, a deterministic regression test. It writes the reused dev:ino into the journal baseline. It fails on the old code (409 on every pass) and passes on the fix.
  • f1bdc91 test(adversarial): daemon output now goes to log files, not an undrained pipe. A sync finding prints the daemon log tails.
  • e46b9fe docs §24.

Verification

  • Round 2 sync section: 50/50 under 10 extra busy-loop CPU hogs (load average 19–33, 8 CPUs). Flip-only loop: 150/150 (old code: failed within 2).
  • #75 stress: PASS 200 daemon-restart collisions; lost_updates=0. pending_rename_campaign.py passes.
  • Gates: cargo fmt --check clean; cargo clippy --workspace --all-targets -- -D warnings clean; cargo test --workspace 1003 passed, 0 failed; tests/adversarial/run.sh: FINDINGS 0, ROUND 2 FINDINGS 0.

Found on the way (not fixed, does not block)

  • The server keeps a tus upload that failed with 412 at the final PATCH as "active" until its 24 h TTL. There is a limit of 16 active uploads per user. The sync client terminates its own stale session, so sync does not leak. But the round-2 probe leaves its two deliberate 412 sessions open, so running the sync section more than about 7 times against one server gives 429 too many active uploads. Other abandoned uploads, for example from browser tabs, count against the same limit.
  • SyncEngine::download scans the full local tree after each download to verify one file. The cost is O(downloads × tree), and a change to any other file fails the pass. This probably contributes to #46.
## Result (branch `job/sync-load`, not merged) **Reproduced.** It is a real bug, and load is not the cause. On the first loop run at load 8, round 2 gave the same 5 sync findings (file→directory, folder collision setup, remote file→local folder, changing file, concurrent edit). The daemon log showed one error on every pass, from the `flip` step to the end of the run: ``` sync reconcile deferred: server returned HTTP 409 ``` A focused loop of only the `flip` step failed on iteration 1 or 2 on the old code. ### Root cause The local file identity was `dev:ino`. File systems reuse a freed inode number at once. After `rm flip; mkdir flip; echo > flip/inner.txt`, the new `flip/inner.txt` sometimes got the inode number of the deleted `flip`. `follow_renames` took that as a rename `flip -> flip/inner.txt` and asked the server to move the file into a folder with its own name. The server refused (mkdir over a file = 409), and the daemon sent the same move again on every pass. The pair stopped syncing. The 4 later findings were effects of this: the local edit in "concurrent edit" was still on disk but was never synced. Load only changes how often the inode number is reused. ### Fix - `8c22c73` fix(sync): the local id is `dev:ino:birth` (birth time from statx / st_birthtime). `same_file_identity` still matches journals that have `dev:ino` ids, so an upgrade does not make spurious conflicts. - `71511c5` fix(sync): a file never moves into a path below itself. If the server refuses a followed rename (409/404/412), the engine plans each path separately and does not retry the move. A new remote file whose local path is a folder goes to a conflict name. Before, the pair also stalled on every pass in this case. - `7bc5a24` test(sync): `crates/calternal-sync/type_flip_campaign.py`, a deterministic regression test. It writes the reused `dev:ino` into the journal baseline. It fails on the old code (409 on every pass) and passes on the fix. - `f1bdc91` test(adversarial): daemon output now goes to log files, not an undrained pipe. A sync finding prints the daemon log tails. - `e46b9fe` docs §24. ### Verification - Round 2 sync section: **50/50** under 10 extra busy-loop CPU hogs (load average 19–33, 8 CPUs). Flip-only loop: 150/150 (old code: failed within 2). - #75 stress: `PASS 200 daemon-restart collisions; lost_updates=0`. `pending_rename_campaign.py` passes. - Gates: `cargo fmt --check` clean; `cargo clippy --workspace --all-targets -- -D warnings` clean; `cargo test --workspace` 1003 passed, 0 failed; `tests/adversarial/run.sh`: `FINDINGS 0`, `ROUND 2 FINDINGS 0`. ### Found on the way (not fixed, does not block) - The server keeps a tus upload that failed with 412 at the final PATCH as "active" until its 24 h TTL. There is a limit of 16 active uploads per user. The sync client terminates its own stale session, so sync does not leak. But the round-2 probe leaves its two deliberate 412 sessions open, so running the sync section more than about 7 times against one server gives `429 too many active uploads`. Other abandoned uploads, for example from browser tabs, count against the same limit. - `SyncEngine::download` scans the full local tree after each download to verify one file. The cost is O(downloads × tree), and a change to any other file fails the pass. This probably contributes to #46.
kayg closed this issue 2026-09-25 15:36:38 +00:00
Author
Owner

Reproduced the existing #88 root-deletion finding during the merged-dev adversarial run (job/analytics, HEAD 1109748cbe5ac209eb5f945165060bc0d8f84438). The sync round deleted its local root, recreated it empty, and then found two remote files missing: keep/k0.txt, keep/k1.txt. The server remained alive. This confirms that the checked-out dev base still has the reported root-deletion behavior; the known repair is on job/sync-load per the earlier #88 report and is not part of this branch. I have not changed Sync here.

Reproduced the existing #88 root-deletion finding during the merged-dev adversarial run (`job/analytics`, HEAD `1109748cbe5ac209eb5f945165060bc0d8f84438`). The sync round deleted its local root, recreated it empty, and then found two remote files missing: `keep/k0.txt`, `keep/k1.txt`. The server remained alive. This confirms that the checked-out `dev` base still has the reported root-deletion behavior; the known repair is on `job/sync-load` per the earlier #88 report and is not part of this branch. I have not changed Sync here.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#88
No description provided.