PERF: skip full pre-migration Index snapshots on unchanged restarts (#663) #748

Open
opened 2026-10-02 13:08:48 +00:00 by kayg · 4 comments
Owner

Context: disk IO and startup audit for #663, rule 8. Source is origin/dev c4a61e8cf090170f35b1bed3350d9de20c83ecd5; the same behavior remains in merge-round-7a 2f4482ded066d9c5d9c59130377907f7fd2916c9.

Every start awaits Db::snapshot before it checks migrations (crates/calternal-server/src/wire.rs:1135-1142; queued source :1143). main.rs:371-372 cannot bind HTTP until setup returns. crates/calternal-db/src/snapshot.rs:64-79 runs VACUUM INTO on the one writer connection, syncs the full output, renames it, syncs the directory and prunes retained backups. A no-change restart therefore copies the complete Index and consumes retention even when no migration will run.

Reasoned impact: output writes and the final sync grow with the Index, even on a host-cached data disk. This adds avoidable deploy downtime. This audit did not measure its duration. #549's hddsql comment reports 737 ms for VACUUM on one seeded restart and a 300 s readiness timeout that this statement does NOT explain. Do not attribute that timeout to this finding.

Concrete fix: check all migration versions and checksums before copying the Index. Preserve the pre-upgrade snapshot required by #23 and DESIGN §27 whenever a migration will run. Skip only no-change restart snapshots. Keep scheduled backup retention independent if an upgrade snapshot must outlive ordinary retention. Fail safely on snapshot failure before applying any migration.

Tests: first boot; unchanged restart produces no second startup snapshot; new migration snapshots old rows before applying; changed checksum fails without migration; failed snapshot leaves schema unchanged. Extend startup timing coverage with a large Index on the locked HDD emulator: at least five cold and warm starts, median/p95/max, startup bytes written and readiness phases. Keep successful readiness separate from background completion.

Duplicate check: searched all open and closed issue titles and read #23, #549 including the hddsql comments, and #698. #23 specifies pre-upgrade safety; #549 owns the current startup phase investigation; #698 concerns Admin reads. This focused issue owns the unnecessary snapshot on an unchanged restart. Coordinate the fix with job/hddsql-549; do not duplicate its startup instrumentation.

No security or correctness blocker is asserted. No product change was made by this audit.

Context: disk IO and startup audit for #663, rule 8. Source is origin/dev `c4a61e8cf090170f35b1bed3350d9de20c83ecd5`; the same behavior remains in merge-round-7a `2f4482ded066d9c5d9c59130377907f7fd2916c9`. Every start awaits `Db::snapshot` before it checks migrations (`crates/calternal-server/src/wire.rs:1135-1142`; queued source :1143). `main.rs:371-372` cannot bind HTTP until setup returns. `crates/calternal-db/src/snapshot.rs:64-79` runs `VACUUM INTO` on the one writer connection, syncs the full output, renames it, syncs the directory and prunes retained backups. A no-change restart therefore copies the complete Index and consumes retention even when no migration will run. Reasoned impact: output writes and the final sync grow with the Index, even on a host-cached data disk. This adds avoidable deploy downtime. This audit did not measure its duration. #549's hddsql comment reports 737 ms for VACUUM on one seeded restart and a 300 s readiness timeout that this statement does NOT explain. Do not attribute that timeout to this finding. Concrete fix: check all migration versions and checksums before copying the Index. Preserve the pre-upgrade snapshot required by #23 and DESIGN §27 whenever a migration will run. Skip only no-change restart snapshots. Keep scheduled backup retention independent if an upgrade snapshot must outlive ordinary retention. Fail safely on snapshot failure before applying any migration. Tests: first boot; unchanged restart produces no second startup snapshot; new migration snapshots old rows before applying; changed checksum fails without migration; failed snapshot leaves schema unchanged. Extend startup timing coverage with a large Index on the locked HDD emulator: at least five cold and warm starts, median/p95/max, startup bytes written and readiness phases. Keep successful readiness separate from background completion. Duplicate check: searched all open and closed issue titles and read #23, #549 including the hddsql comments, and #698. #23 specifies pre-upgrade safety; #549 owns the current startup phase investigation; #698 concerns Admin reads. This focused issue owns the unnecessary snapshot on an unchanged restart. Coordinate the fix with job/hddsql-549; do not duplicate its startup instrumentation. No security or correctness blocker is asserted. No product change was made by this audit.
Author
Owner

SQLite audit #663 on job/perf-arch-db. Base c4a61e8cf0; queued merge-round-7a 2f4482ded0 checked separately.

Related runtime snapshot evidence for the same snapshot owner: db/src/snapshot.rs:64 executes VACUUM INTO on writer_pool. wire.rs:4132 admin trigger and :4824 system.backup use it after startup too. The complete copy holds the only application writer connection; WAL readers can continue, but every queued write waits for the full copy. The existing comments acknowledge this. In addition to skipping unchanged-start snapshots, evaluate a separate bounded backup/maintenance connection and a consistent snapshot API with pause/yield and cancellation; preserve fsync, atomic installation and retention. Test an unrelated writer during a large runtime snapshot and pinning of checkpoint frames. No runtime timing was measured.

Local query-work figures use Python SQLite 3.53.3 and synthetic data; VM callbacks count 1,000-instruction units. They are not production latency, CPU/RSS or HDD budget samples. The locked Rust bundled SQLite is 3.51.3. Repeat with audit-sqlite.py --plans and confirm the production-version plan before implementing. No existing assertion or product behavior was changed. Reuse this issue instead of filing a duplicate.

SQLite audit #663 on job/perf-arch-db. Base c4a61e8cf090170f35b1bed3350d9de20c83ecd5; queued merge-round-7a 2f4482ded066d9c5d9c59130377907f7fd2916c9 checked separately. Related runtime snapshot evidence for the same snapshot owner: db/src/snapshot.rs:64 executes VACUUM INTO on writer_pool. wire.rs:4132 admin trigger and :4824 system.backup use it after startup too. The complete copy holds the only application writer connection; WAL readers can continue, but every queued write waits for the full copy. The existing comments acknowledge this. In addition to skipping unchanged-start snapshots, evaluate a separate bounded backup/maintenance connection and a consistent snapshot API with pause/yield and cancellation; preserve fsync, atomic installation and retention. Test an unrelated writer during a large runtime snapshot and pinning of checkpoint frames. No runtime timing was measured. Local query-work figures use Python SQLite 3.53.3 and synthetic data; VM callbacks count 1,000-instruction units. They are not production latency, CPU/RSS or HDD budget samples. The locked Rust bundled SQLite is 3.51.3. Repeat with audit-sqlite.py --plans and confirm the production-version plan before implementing. No existing assertion or product behavior was changed. Reuse this issue instead of filing a duplicate.
Author
Owner

Starting implementation for this issue on branch job/ioperf, based at 2f4482ded066d9c5d9c59130377907f7fd2916c9 (job/merge-round-7a). I have read the issue and the #663 audit evidence. I am inspecting the job/hddsql-549 changes before editing shared startup, filesystem, and database paths. I will add a regression test and a focused benchmark for each performance path, then report measured results and crate gates here.

Starting implementation for this issue on branch `job/ioperf`, based at `2f4482ded066d9c5d9c59130377907f7fd2916c9` (`job/merge-round-7a`). I have read the issue and the #663 audit evidence. I am inspecting the `job/hddsql-549` changes before editing shared startup, filesystem, and database paths. I will add a regression test and a focused benchmark for each performance path, then report measured results and crate gates here.
Author
Owner

Confirmed in crates/calternal-server/src/wire.rs: startup calls Db::snapshot before apply_migration_sets, so an unchanged restart runs VACUUM INTO and prunes backup retention before it knows whether any migration will run. I am adding a registered-version/checksum preflight and keeping snapshot failure ahead of migration application. The regression coverage is in crates/calternal-db/src/migrations.rs.

Confirmed in `crates/calternal-server/src/wire.rs`: startup calls `Db::snapshot` before `apply_migration_sets`, so an unchanged restart runs `VACUUM INTO` and prunes backup retention before it knows whether any migration will run. I am adding a registered-version/checksum preflight and keeping snapshot failure ahead of migration application. The regression coverage is in `crates/calternal-db/src/migrations.rs`.
Author
Owner

ioperf final report

Built changes for #748, #750, #764, #800 and #806. Head: d43e7829a3e0c8a483cc45bdbb965af5f5d1c713.

Changes

  • #748: validate pending migrations before taking a startup snapshot; unchanged restarts skip it.
  • #750: list and hash one folder outside the Root mutation lock, then validate generations and publish the prepared Index changes.
  • #764: idle workers subscribe to queue changes and wake for due-job and lease deadlines; the server fallback is 30 seconds.
  • #800: DAV stat/open reads use indexed hashes without holding the Root mutation lock and verify full file fingerprints.
  • #806: the server owns one bounded recursive Home watcher and forwards external hints through the shared Root change stream. Overflow triggers a coalesced rebuild, including open collaboration rooms.

Files

crates/calternal-db/src/migrations.rs, crates/calternal-db/src/worker.rs, crates/calternal-server/src/wire.rs, crates/plugins/files/src/dav.rs, crates/plugins/files/src/index.rs, crates/plugins/files/src/lib.rs, crates/calternal-fs/src/lib.rs, crates/calternal-fs/src/quota.rs, crates/calternal-fs/src/root.rs, crates/calternal-fs/tests/storage.rs, crates/calternal-search/src/indexer.rs, crates/calternal-collab/src/session.rs.

Gate output

cargo fmt --all --check: exit 0; no output.

cargo clippy --offline -p calternal-db --all-targets -- -D warnings:

Finished `dev` profile [unoptimized + debuginfo] target(s) in 1m 55s

RUST_TEST_THREADS=1 cargo test --offline -p calternal-db:

running 27 tests
test result: ok. 27 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 68.51s

running 17 tests
test result: ok. 16 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 25.17s

running 1 test
test sqlite_pools_bound_readers_and_page_cache ... ok
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 1.80s

Doc-tests calternal_db
running 0 tests
test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s

Focused Files regressions:

running 4 tests
test dav::tests::open_read_does_not_wait_for_the_home_mutation_lock ... ok
test dav::tests::stat_does_not_wait_for_the_home_mutation_lock ... ok
test tests::reconcile_hash_does_not_hold_mutation_lock_across_homes ... ok
test tests::public_password_rejection_does_not_wait_for_data_mutation_lock ... ok

test result: ok. 4 passed; 0 failed; 0 ignored; 0 measured; 156 filtered out; finished in 4.16s

bun run check:

$ node scripts/check-user-storage.mjs && node scripts/check-type-tokens.mjs && node scripts/check-motion-tokens.mjs && svelte-kit sync && svelte-check --tsconfig ./tsconfig.json
User browser caches use userStorage; only documented device/public-link exceptions remain.
Text sizes and UI shape values use shared role tokens.
UI transitions and animation options use shared motion tokens or documented exceptions.
Loading svelte-check in workspace: /home/kayg/Developer/calternal-wt/ioperf/apps/web
Getting Svelte diagnostics...

svelte-check found 0 errors and 0 warnings

Known gaps

  • The full Files test command was interrupted with exit 130 when the shared target stayed blocked on I/O. The focused four tests passed.
  • Final clippy/test gates for Files, server, fs, search and collaboration were not completed. The #806 server/filesystem integration is therefore not compile-verified in this run.
  • bun run test, adversarial probing, and the requested per-feature performance profiles and measurements were not completed.
  • cargo clean and removal of the generated web build output remain undone.
  • No UI changed, so UX gaps and screenshots are N/A.

Decisions

  • The #764 cross-process fallback is 30 seconds; local queue hints and durable deadlines provide prompt wakeups.
  • Watcher overflow recovery waits for a 100 ms quiet window before repeating a full repair.
  • Folder reconciliation reuses the existing eight-attempt policy for changed snapshots.
  • External watcher events carry an explicit source on the existing Root change stream, so only external changes publish the Home feed.

No UX gaps were introduced. Commits are on job/ioperf; no push or deploy was made.

## ioperf final report Built changes for #748, #750, #764, #800 and #806. Head: `d43e7829a3e0c8a483cc45bdbb965af5f5d1c713`. ### Changes - #748: validate pending migrations before taking a startup snapshot; unchanged restarts skip it. - #750: list and hash one folder outside the Root mutation lock, then validate generations and publish the prepared Index changes. - #764: idle workers subscribe to queue changes and wake for due-job and lease deadlines; the server fallback is 30 seconds. - #800: DAV stat/open reads use indexed hashes without holding the Root mutation lock and verify full file fingerprints. - #806: the server owns one bounded recursive Home watcher and forwards external hints through the shared Root change stream. Overflow triggers a coalesced rebuild, including open collaboration rooms. ### Files `crates/calternal-db/src/migrations.rs`, `crates/calternal-db/src/worker.rs`, `crates/calternal-server/src/wire.rs`, `crates/plugins/files/src/dav.rs`, `crates/plugins/files/src/index.rs`, `crates/plugins/files/src/lib.rs`, `crates/calternal-fs/src/lib.rs`, `crates/calternal-fs/src/quota.rs`, `crates/calternal-fs/src/root.rs`, `crates/calternal-fs/tests/storage.rs`, `crates/calternal-search/src/indexer.rs`, `crates/calternal-collab/src/session.rs`. ### Gate output `cargo fmt --all --check`: exit 0; no output. `cargo clippy --offline -p calternal-db --all-targets -- -D warnings`: ``` Finished `dev` profile [unoptimized + debuginfo] target(s) in 1m 55s ``` `RUST_TEST_THREADS=1 cargo test --offline -p calternal-db`: ``` running 27 tests test result: ok. 27 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 68.51s running 17 tests test result: ok. 16 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 25.17s running 1 test test sqlite_pools_bound_readers_and_page_cache ... ok test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 1.80s Doc-tests calternal_db running 0 tests test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s ``` Focused Files regressions: ``` running 4 tests test dav::tests::open_read_does_not_wait_for_the_home_mutation_lock ... ok test dav::tests::stat_does_not_wait_for_the_home_mutation_lock ... ok test tests::reconcile_hash_does_not_hold_mutation_lock_across_homes ... ok test tests::public_password_rejection_does_not_wait_for_data_mutation_lock ... ok test result: ok. 4 passed; 0 failed; 0 ignored; 0 measured; 156 filtered out; finished in 4.16s ``` `bun run check`: ``` $ node scripts/check-user-storage.mjs && node scripts/check-type-tokens.mjs && node scripts/check-motion-tokens.mjs && svelte-kit sync && svelte-check --tsconfig ./tsconfig.json User browser caches use userStorage; only documented device/public-link exceptions remain. Text sizes and UI shape values use shared role tokens. UI transitions and animation options use shared motion tokens or documented exceptions. Loading svelte-check in workspace: /home/kayg/Developer/calternal-wt/ioperf/apps/web Getting Svelte diagnostics... svelte-check found 0 errors and 0 warnings ``` ### Known gaps - The full Files test command was interrupted with exit 130 when the shared target stayed blocked on I/O. The focused four tests passed. - Final clippy/test gates for Files, server, fs, search and collaboration were not completed. The #806 server/filesystem integration is therefore not compile-verified in this run. - `bun run test`, adversarial probing, and the requested per-feature performance profiles and measurements were not completed. - `cargo clean` and removal of the generated web build output remain undone. - No UI changed, so UX gaps and screenshots are N/A. ### Decisions - The #764 cross-process fallback is 30 seconds; local queue hints and durable deadlines provide prompt wakeups. - Watcher overflow recovery waits for a 100 ms quiet window before repeating a full repair. - Folder reconciliation reuses the existing eight-attempt policy for changed snapshots. - External watcher events carry an explicit source on the existing Root change stream, so only external changes publish the Home feed. No UX gaps were introduced. Commits are on `job/ioperf`; no push or deploy was made.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#748
No description provided.