Sync: three campaigns broken on base (local_failure, mass_deletion start without a token; feed_rescan checks before commit) #121

Closed
opened 2026-09-25 21:00:49 +00:00 by kayg · 1 comment
Owner

Found by job/sync-download (#97): crates/calternal-sync campaigns local_failure and mass_deletion start the daemon without CALTERNAL_TOKEN_SERVER so it has no credential and the campaign cannot test anything; feed_rescan asserts the journal before the pass commits. They fail on the base commit too. Fix the harnesses so each campaign actually exercises its scenario, then confirm they pass (a campaign that cannot fail is worse than none).

Found by job/sync-download (#97): crates/calternal-sync campaigns local_failure and mass_deletion start the daemon without CALTERNAL_TOKEN_SERVER so it has no credential and the campaign cannot test anything; feed_rescan asserts the journal before the pass commits. They fail on the base commit too. Fix the harnesses so each campaign actually exercises its scenario, then confirm they pass (a campaign that cannot fail is worse than none).
Author
Owner

Report: job/sync-campaigns (not pushed or merged)

Fixes

  • Shared harness (integration_campaign.py): added boot, client_env, spawn_daemon, start_daemon(preexec_fn=), cli, journal_paths and pair_config. Every campaign now uses them instead of its own copy of the setup. client_env always sets CALTERNAL_TOKEN_SERVER.
  • local_failure: now uses the shared env, so the daemon has a credential. The old 64 KiB RLIMIT_FSIZE made the journal fail first (disk I/O error), so the download was never tried. The limit is now 8 MiB with a 24 MiB download. The campaign now also proves that the limited daemon really syncs: a small download and a local upload go through while the big download fails. On every poll it checks that no partial file exists at the final path. After a restart, temp recovery and the exact bytes are checked.
  • mass_deletion: now uses the shared env. After --confirm, the daemon restarts on the saved config. The campaign checks that the server keeps the file for 1.5 s after the guard trips (a deletion that gets through reaches the server well inside that time). It also checks that the confirmation is used up after one pass.
  • feed_rescan: now waits until both journals have committed. This applies to the baseline cursor and also to the cursor after the 410, which must be at or past the feed floor.
  • One entry point: python3 crates/calternal-sync/run_campaigns.py (--list, or give names to run only some). A campaign that ends without a PASS line counts as failed. The default integration run now prints a PASS line. DESIGN §24 lists all 13 campaigns and this command.

Proof that each campaign can fail (each break was temporary and is not committed)

Break Result
Base campaigns, unchanged all three FAIL (File too large missing / mass-deletion state did not advance / assert cursor >= head)
Guard skipped (if false && deletion_guard(..)) mass_deletion: FAIL … did not advance: guard pauses the pair
Daemon never clears allow_mass_delete_once mass_deletion: FAIL … did not advance: confirmation is used once
410 fails the pass instead of falling through to the full listing feed_rescan: FAIL … file did not converge: late.txt
Commit keeps the expired cursor after a 410 feed_rescan: FAIL … journals did not commit: cursor at or past the feed floor
Failed download renames its temp file to the final path local_failure: FAIL … partial download published
Deferral fix disabled (base engine) local_failure: FAIL … limited daemon did not sync the small file

Real bug found and fixed

When one remote file could not be written locally (full disk, file size limit, folder with no write permission), every pass failed. Then nothing else in the pair synced, and local edits were not uploaded either. Now Error::Io from a download defers only that path. The path keeps its stored journal row, so the next pass plans it again from the same baseline and nothing is lost or overwritten. The rest of the pass commits, and the daemon retries the deferred path with backoff (100 ms up to 5 s). Remote errors, destination races and hash errors still fail the pass. The mass-deletion guard runs before any action, so this change does not affect it. The local_failure campaign is the regression test, together with the unit test only_a_local_write_failure_defers_a_download. DESIGN §24 now documents the rule.

Commits

  • f6f62f4 fix(sync): a local write failure defers only its own download
  • d8987ba test(sync): make the broken campaigns able to fail; share their setup
  • 52cea20 docs: list the sync campaigns and their entry point in §24
  • e58da8f test(sync): local_failure checks for a published partial file on every poll
  • 2c23e93 test(sync): the campaign runner leaves no pycache in the tree

Gates

integration: PASS 30 two-client convergence rounds; latency p50=1.411s p95=3.118s p99=3.474s
kill_rounds: PASS 30 process-kill rounds; seed=17
precondition_race: PASS stale listing race; fresh stat preserved both revisions without a stale upload
precondition_race: PASS precondition race; stale upload rejected with HTTP 412; conflict resolved before another tree scan
workload: PASS edit, rename, folder move, concurrent conflict; two clients and server have identical bytes
random: FAIL (exit 1, 85.7s)
latency p50=1.049s p95=2.221s p99=3.781s target-p95<1.000s
AssertionError: edit-to-visible p95 2.221s was not below 1.000s
mid_transfer: PASS 64 MiB tus transfer resumed after server SIGKILL mid-upload; exact bytes on both clients
mid_transfer: PASS local edit during tus upload kept the first revision in server versions and converged to the edit
pending_rename: PASS pending uploads followed case-only rename chain and folder move
type_flip: PASS reused inode identity: type flip and refused rename converge
download_race: PASS unrelated edit during download; committed in 0.116s; other edits sync while a file keeps changing
download_race: PASS same-file edit during download; kept as a conflict copy
local_failure: PASS local write failure left no partial file, did not stop the pair, and a restart recovered exact remote bytes
mass_deletion: PASS mass-deletion guard paused; --confirm allowed exactly one pass
feed_rescan: PASS feed 410 caused full remote rescan and advanced journal cursor
restart_collisions: PASS 200 daemon-restart collisions; lost_updates=0; seed=17
  • cargo fmt --check: exit 0, no output
  • cargo clippy --workspace --all-targets -- -D warnings: Finished \dev` profile [unoptimized + debuginfo] target(s) in 29.63s`
  • cargo test -p calternal-sync: test result: ok. 37 passed; 0 failed / ok. 2 passed / ok. 0 passed

random fails only its latency target (p95 < 1 s). All 30 steps converge. The host load average was 20 to 50 on 8 cores because other jobs were running. The base daemon misses the target in the same way. Interleaved A/B runs gave p95 of base 2.696 s / 3.655 s against fix 2.358 s / 3.389 s. So this is host load, not this change. To check the target, run run_campaigns.py random on an idle host.

## Report: job/sync-campaigns (not pushed or merged) ### Fixes - **Shared harness** (`integration_campaign.py`): added `boot`, `client_env`, `spawn_daemon`, `start_daemon(preexec_fn=)`, `cli`, `journal_paths` and `pair_config`. Every campaign now uses them instead of its own copy of the setup. `client_env` always sets `CALTERNAL_TOKEN_SERVER`. - **local_failure**: now uses the shared env, so the daemon has a credential. The old 64 KiB `RLIMIT_FSIZE` made the journal fail first (`disk I/O error`), so the download was never tried. The limit is now 8 MiB with a 24 MiB download. The campaign now also proves that the limited daemon really syncs: a small download and a local upload go through while the big download fails. On every poll it checks that no partial file exists at the final path. After a restart, temp recovery and the exact bytes are checked. - **mass_deletion**: now uses the shared env. After `--confirm`, the daemon restarts on the saved config. The campaign checks that the server keeps the file for 1.5 s after the guard trips (a deletion that gets through reaches the server well inside that time). It also checks that the confirmation is used up after one pass. - **feed_rescan**: now waits until both journals have committed. This applies to the baseline cursor and also to the cursor after the 410, which must be at or past the feed floor. - **One entry point**: `python3 crates/calternal-sync/run_campaigns.py` (`--list`, or give names to run only some). A campaign that ends without a PASS line counts as failed. The default `integration` run now prints a PASS line. DESIGN §24 lists all 13 campaigns and this command. ### Proof that each campaign can fail (each break was temporary and is not committed) | Break | Result | | --- | --- | | Base campaigns, unchanged | all three FAIL (`File too large` missing / `mass-deletion state did not advance` / `assert cursor >= head`) | | Guard skipped (`if false && deletion_guard(..)`) | `mass_deletion: FAIL … did not advance: guard pauses the pair` | | Daemon never clears `allow_mass_delete_once` | `mass_deletion: FAIL … did not advance: confirmation is used once` | | 410 fails the pass instead of falling through to the full listing | `feed_rescan: FAIL … file did not converge: late.txt` | | Commit keeps the expired cursor after a 410 | `feed_rescan: FAIL … journals did not commit: cursor at or past the feed floor` | | Failed download renames its temp file to the final path | `local_failure: FAIL … partial download published` | | Deferral fix disabled (base engine) | `local_failure: FAIL … limited daemon did not sync the small file` | ### Real bug found and fixed When one remote file could not be written locally (full disk, file size limit, folder with no write permission), every pass failed. Then nothing else in the pair synced, and local edits were not uploaded either. Now `Error::Io` from a download defers only that path. The path keeps its stored journal row, so the next pass plans it again from the same baseline and nothing is lost or overwritten. The rest of the pass commits, and the daemon retries the deferred path with backoff (100 ms up to 5 s). Remote errors, destination races and hash errors still fail the pass. The mass-deletion guard runs before any action, so this change does not affect it. The `local_failure` campaign is the regression test, together with the unit test `only_a_local_write_failure_defers_a_download`. DESIGN §24 now documents the rule. ### Commits - f6f62f4 fix(sync): a local write failure defers only its own download - d8987ba test(sync): make the broken campaigns able to fail; share their setup - 52cea20 docs: list the sync campaigns and their entry point in §24 - e58da8f test(sync): local_failure checks for a published partial file on every poll - 2c23e93 test(sync): the campaign runner leaves no __pycache__ in the tree ### Gates ``` integration: PASS 30 two-client convergence rounds; latency p50=1.411s p95=3.118s p99=3.474s kill_rounds: PASS 30 process-kill rounds; seed=17 precondition_race: PASS stale listing race; fresh stat preserved both revisions without a stale upload precondition_race: PASS precondition race; stale upload rejected with HTTP 412; conflict resolved before another tree scan workload: PASS edit, rename, folder move, concurrent conflict; two clients and server have identical bytes random: FAIL (exit 1, 85.7s) latency p50=1.049s p95=2.221s p99=3.781s target-p95<1.000s AssertionError: edit-to-visible p95 2.221s was not below 1.000s mid_transfer: PASS 64 MiB tus transfer resumed after server SIGKILL mid-upload; exact bytes on both clients mid_transfer: PASS local edit during tus upload kept the first revision in server versions and converged to the edit pending_rename: PASS pending uploads followed case-only rename chain and folder move type_flip: PASS reused inode identity: type flip and refused rename converge download_race: PASS unrelated edit during download; committed in 0.116s; other edits sync while a file keeps changing download_race: PASS same-file edit during download; kept as a conflict copy local_failure: PASS local write failure left no partial file, did not stop the pair, and a restart recovered exact remote bytes mass_deletion: PASS mass-deletion guard paused; --confirm allowed exactly one pass feed_rescan: PASS feed 410 caused full remote rescan and advanced journal cursor restart_collisions: PASS 200 daemon-restart collisions; lost_updates=0; seed=17 ``` - `cargo fmt --check`: exit 0, no output - `cargo clippy --workspace --all-targets -- -D warnings`: `Finished \`dev\` profile [unoptimized + debuginfo] target(s) in 29.63s` - `cargo test -p calternal-sync`: `test result: ok. 37 passed; 0 failed` / `ok. 2 passed` / `ok. 0 passed` **`random` fails only its latency target (p95 < 1 s). All 30 steps converge.** The host load average was 20 to 50 on 8 cores because other jobs were running. The base daemon misses the target in the same way. Interleaved A/B runs gave p95 of base 2.696 s / 3.655 s against fix 2.358 s / 3.389 s. So this is host load, not this change. To check the target, run `run_campaigns.py random` on an idle host.
kayg closed this issue 2026-09-25 22:39:13 +00:00
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#121
No description provided.