SYNC: a file that keeps changing during upload ends with stale remote bytes #220

Closed
opened 2026-09-27 08:20:44 +00:00 by kayg · 7 comments
Owner

Adversarial round on dev cf142cb4 (tests/adversarial/attack2.py, 'sync changing file'): while the daemon runs, a 3 MiB local file is rewritten repeatedly (appends of 4 KiB every 50 ms), then left alone. 30 s later the remote copy still differs from the final local bytes (len 3145728 remote). The final state must converge: after the last local write settles, the next pass must upload the final bytes. Likely causes to check: change detection keyed on mtime/size that misses a write within the same timestamp granularity, a hash taken while the file was still changing and then trusted, or the change-feed path (#123) recording the remote version as the synced baseline for the local state it did not match.
Fix in crates/calternal-sync with a focused reproduction test (a file that changes during upload converges to its final bytes on the remote, including same-second writes and same-size rewrites) and a sync campaign case. Blocks merges (sync correctness).

Adversarial round on dev cf142cb4 (tests/adversarial/attack2.py, 'sync changing file'): while the daemon runs, a 3 MiB local file is rewritten repeatedly (appends of 4 KiB every 50 ms), then left alone. 30 s later the remote copy still differs from the final local bytes (len 3145728 remote). The final state must converge: after the last local write settles, the next pass must upload the final bytes. Likely causes to check: change detection keyed on mtime/size that misses a write within the same timestamp granularity, a hash taken while the file was still changing and then trusted, or the change-feed path (#123) recording the remote version as the synced baseline for the local state it did not match. Fix in crates/calternal-sync with a focused reproduction test (a file that changes during upload converges to its final bytes on the remote, including same-second writes and same-size rewrites) and a sync campaign case. Blocks merges (sync correctness).
Author
Owner

Starting work on job/sync-changing from dev at f170d4e972. I will reproduce the changing-file upload failure with a focused test before fixing it.

Starting work on job/sync-changing from dev at f170d4e97231032d4ebf4021ab3bcb278a453829. I will reproduce the changing-file upload failure with a focused test before fixing it.
Author
Owner

Reproduced a stable-read detection gap in crates/calternal-sync/src/local.rs: a same-size rewrite with mtime restored changes ctime, but the current read-stability predicate accepts the before/after metadata as equal. The focused Rust test fails at the assertion that these metadata snapshots differ for stable-read purposes. The live-server campaign and adversarial sync section passed once each; I am fixing the concrete ctime omission and keeping the deterministic campaign as regression coverage.

Reproduced a stable-read detection gap in crates/calternal-sync/src/local.rs: a same-size rewrite with mtime restored changes ctime, but the current read-stability predicate accepts the before/after metadata as equal. The focused Rust test fails at the assertion that these metadata snapshots differ for stable-read purposes. The live-server campaign and adversarial sync section passed once each; I am fixing the concrete ctime omission and keeping the deterministic campaign as regression coverage.
Author
Owner

The sync campaign matrix found a non-SLOW convergence failure: workload_campaign.py exited 1 after 49.7 s at its concurrent-edit step (baseline-1.txt edited on both clients). It did not converge with both byte streams within 30 s. I am rerunning this campaign separately to capture daemon state and determine the cause before final gates.

The sync campaign matrix found a non-SLOW convergence failure: `workload_campaign.py` exited 1 after 49.7 s at its concurrent-edit step (`baseline-1.txt` edited on both clients). It did not converge with both byte streams within 30 s. I am rerunning this campaign separately to capture daemon state and determine the cause before final gates.
Author
Owner

Confirmed and corrected two pre-existing campaign false failures. The daemon excludes .calternal-sync-root-id from sync, but the workload and random local inventories included that per-client marker. After filtering client state through a shared campaign helper, workload passed and random completed 30 exact convergence rounds. Its edit-to-visible p95 was 2.171 s versus the 1.000 s target while several other jobs were active; all rounds converged. I am treating that SLOW-only result as shared-host load and will retry after the host clears.

Confirmed and corrected two pre-existing campaign false failures. The daemon excludes `.calternal-sync-root-id` from sync, but the `workload` and `random` local inventories included that per-client marker. After filtering client state through a shared campaign helper, `workload` passed and `random` completed 30 exact convergence rounds. Its edit-to-visible p95 was 2.171 s versus the 1.000 s target while several other jobs were active; all rounds converged. I am treating that SLOW-only result as shared-host load and will retry after the host clears.
Author
Owner

The retried feed_rescan case passed. restart_collisions ran 70/200 rounds, then its 10-second HTTP request for a tiny remote replacement timed out while the campaign's intentional CPU load was active. This is a SLOW-only timeout on the shared host, not a reported convergence mismatch. I will retry when other active jobs clear.

The retried `feed_rescan` case passed. `restart_collisions` ran 70/200 rounds, then its 10-second HTTP request for a tiny remote replacement timed out while the campaign's intentional CPU load was active. This is a SLOW-only timeout on the shared host, not a reported convergence mismatch. I will retry when other active jobs clear.
Author
Owner

Finished — Forgejo #220

Branch: job/sync-changing
Head: d4c8228aa6345c35c5ab289a2f3538c30a27d48f
Merged current dev at 2a379b185677fd6d85980900c5e757af990d9527 before final verification.

Built

  • Stable local reads now compare the full StatKey before and after hashing. This includes ctime, so a same-size rewrite with restored mtime cannot reuse a stale revision.
  • Added a regression test and real-server changing-file campaign for same-size edits during upload.
  • Campaign inventories now ignore the daemon's private root marker and valid temporary marker/download names. Proxy teardown also treats peer disconnects as expected.

Job files:

  • crates/calternal-sync/src/local.rs
  • crates/calternal-sync/changing_file_campaign.py
  • crates/calternal-sync/run_campaigns.py
  • crates/calternal-sync/integration_campaign.py
  • crates/calternal-sync/workload_campaign.py
  • crates/calternal-sync/random_campaign.py

Validation

  • Sync campaign matrix completed. The changing-file probe passed; the 200-round restart-collision probe reported lost_updates=0. An earlier random campaign run converged in all 30 rounds but reported p95 2.171s under shared-host load; it was classified as SLOW-only.
  • Latest real-server adversarial sync section:
---------- sync ----------
server alive at end: True

==== ROUND 2 FINDINGS 0
  • cargo fmt --check: exit 0, no output.
  • cargo clippy --all-targets -- -D warnings output:
    Checking calternal-collab v0.0.1 (/home/kayg/Developer/calternal-wt/sync-changing/crates/calternal-collab)
   Compiling calternal-server v0.0.1 (/home/kayg/Developer/calternal-wt/sync-changing/crates/calternal-server)
    Finished `dev` profile [unoptimized + debuginfo] target(s) in 25.41s
  • The first unfiltered cargo test hit this shared-host timing assertion in unchanged calternal-notes-core:
[perf] parse_task_runs=8205µs  parse_entry_runs=111µs  ("p1 !!! every Jan 15 by 2026-06-30 on 2026-07-01 #a #b #c @x +y p1 !!! every Jan 15 by 2026-06-30 on 2026-07-01 #a #b #c @x +y p1 !!! every Jan 15 by 2026-06-30 on 2026-07-01 #a #b #c @x +y p1 !!! every Jan 15 by 2026-06-30 on 2026-07-01 #a #b #c @x +y ")
parse_task_runs "p1 !!! every Jan 15 by 2026-06-30 on 2026-07-01 #a #b #c @x +y p1 !!! every Jan 15 by 2026-06-30 on 2026-07-01 #a #b #c @x +y p1 !!! every Jan 15 by 2026-06-30 on 2026-07-01 #a #b #c @x +y p1 !!! every Jan 15 by 2026-06-30 on 2026-07-01 #a #b #c @x +y ": 8205µs ≥ 8000µs budget
test result: FAILED. 476 passed; 1 failed; 0 ignored; 0 measured; 0 filtered out; finished in 2.40s
  • On the latest merge, the isolated timing test passed, then the full workspace suite with only that timing assertion filtered exited 0: 1,240 passed, 12 ignored, 1 filtered, 0 failed. Its notes-core summary was:
test result: ok. 476 passed; 0 failed; 0 ignored; 0 measured; 1 filtered out; finished in 1.70s
  • bun run check: svelte-check found 0 errors and 0 warnings.
  • bun run test: Test Files 78 passed (78); Tests 579 passed (579).
  • bun run build: ✓ built in 1m 22s; adapter-static wrote the site to build.

Known gap

The unfiltered workspace test command had one transient 205 µs overage against an 8 ms performance budget in an unrelated notes-core test. The isolated test passed on the latest merge. All other workspace tests passed in the complete rerun. The random campaign's earlier 2.171 s p95 was SLOW-only under shared-host load; all 30 rounds converged.

Decisions outside DESIGN.md

  • Use ctime in the stable-read revision comparison because StatKey already uses it for cache validity.
  • Exclude only the daemon's reserved marker and well-formed private temporary names from test inventories because these files do not enter the remote namespace.
  • Treat BrokenPipeError and ConnectionResetError during proxy shutdown as expected teardown after a peer closes first.
# Finished — Forgejo #220 Branch: `job/sync-changing` Head: `d4c8228aa6345c35c5ab289a2f3538c30a27d48f` Merged current `dev` at `2a379b185677fd6d85980900c5e757af990d9527` before final verification. ## Built - Stable local reads now compare the full `StatKey` before and after hashing. This includes ctime, so a same-size rewrite with restored mtime cannot reuse a stale revision. - Added a regression test and real-server changing-file campaign for same-size edits during upload. - Campaign inventories now ignore the daemon's private root marker and valid temporary marker/download names. Proxy teardown also treats peer disconnects as expected. Job files: - `crates/calternal-sync/src/local.rs` - `crates/calternal-sync/changing_file_campaign.py` - `crates/calternal-sync/run_campaigns.py` - `crates/calternal-sync/integration_campaign.py` - `crates/calternal-sync/workload_campaign.py` - `crates/calternal-sync/random_campaign.py` ## Validation - Sync campaign matrix completed. The changing-file probe passed; the 200-round restart-collision probe reported `lost_updates=0`. An earlier random campaign run converged in all 30 rounds but reported p95 `2.171s` under shared-host load; it was classified as SLOW-only. - Latest real-server adversarial sync section: ```text ---------- sync ---------- server alive at end: True ==== ROUND 2 FINDINGS 0 ``` - `cargo fmt --check`: exit 0, no output. - `cargo clippy --all-targets -- -D warnings` output: ```text Checking calternal-collab v0.0.1 (/home/kayg/Developer/calternal-wt/sync-changing/crates/calternal-collab) Compiling calternal-server v0.0.1 (/home/kayg/Developer/calternal-wt/sync-changing/crates/calternal-server) Finished `dev` profile [unoptimized + debuginfo] target(s) in 25.41s ``` - The first unfiltered `cargo test` hit this shared-host timing assertion in unchanged `calternal-notes-core`: ```text [perf] parse_task_runs=8205µs parse_entry_runs=111µs ("p1 !!! every Jan 15 by 2026-06-30 on 2026-07-01 #a #b #c @x +y p1 !!! every Jan 15 by 2026-06-30 on 2026-07-01 #a #b #c @x +y p1 !!! every Jan 15 by 2026-06-30 on 2026-07-01 #a #b #c @x +y p1 !!! every Jan 15 by 2026-06-30 on 2026-07-01 #a #b #c @x +y ") parse_task_runs "p1 !!! every Jan 15 by 2026-06-30 on 2026-07-01 #a #b #c @x +y p1 !!! every Jan 15 by 2026-06-30 on 2026-07-01 #a #b #c @x +y p1 !!! every Jan 15 by 2026-06-30 on 2026-07-01 #a #b #c @x +y p1 !!! every Jan 15 by 2026-06-30 on 2026-07-01 #a #b #c @x +y ": 8205µs ≥ 8000µs budget test result: FAILED. 476 passed; 1 failed; 0 ignored; 0 measured; 0 filtered out; finished in 2.40s ``` - On the latest merge, the isolated timing test passed, then the full workspace suite with only that timing assertion filtered exited 0: 1,240 passed, 12 ignored, 1 filtered, 0 failed. Its notes-core summary was: ```text test result: ok. 476 passed; 0 failed; 0 ignored; 0 measured; 1 filtered out; finished in 1.70s ``` - `bun run check`: `svelte-check found 0 errors and 0 warnings`. - `bun run test`: `Test Files 78 passed (78)`; `Tests 579 passed (579)`. - `bun run build`: `✓ built in 1m 22s`; adapter-static wrote the site to `build`. ## Known gap The unfiltered workspace test command had one transient 205 µs overage against an 8 ms performance budget in an unrelated notes-core test. The isolated test passed on the latest merge. All other workspace tests passed in the complete rerun. The random campaign's earlier 2.171 s p95 was SLOW-only under shared-host load; all 30 rounds converged. ## Decisions outside DESIGN.md - Use ctime in the stable-read revision comparison because `StatKey` already uses it for cache validity. - Exclude only the daemon's reserved marker and well-formed private temporary names from test inventories because these files do not enter the remote namespace. - Treat `BrokenPipeError` and `ConnectionResetError` during proxy shutdown as expected teardown after a peer closes first.
kayg referenced this issue from a commit 2026-09-27 11:12:54 +00:00
Author
Owner

Merged in 2d84737b: stable reads now compare the full StatKey (incl. ctime); changing-file campaign added.

Merged in 2d84737b: stable reads now compare the full StatKey (incl. ctime); changing-file campaign added.
kayg closed this issue 2026-09-27 11:13:13 +00:00
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#220
No description provided.