Bulk import: semantic index queue full and job worker crash on SQLITE_BUSY #1157

Open
opened 2026-10-05 15:31:26 +00:00 by kayg · 18 comments
Owner

Found while seeding the website showcase instance (#1143) on 2026-10-05. Build: dev 04afb00, server image from the root Containerfile (with the #1143 fixes), arm64, OrbStack.

The seed uploads about 300 photos and about 480 files through WebDAV and the API in a few minutes, as a new User would on a first import.

Server log during the burst:

WARN calternal_server::wire: watched semantic search change notification failed error=semantic index queue is full   (many times)
WARN calternal_plugin_notes: Notes Index operation failed sqlite_code="5" pool_timeout=false   (twice)
ERROR calternal_server::wire: job worker failed; restarting error=SQLite operation failed: error returned from database: (code: 5) database is locked

Expected:

  1. A full semantic index queue does not drop changes. It applies back-pressure or schedules a catch-up scan, so search finds every file after the burst.
  2. A busy Index (SQLITE_BUSY) does not crash the job worker. The worker retries with a busy timeout and does not restart.
  3. A Notes Index operation that gets SQLITE_BUSY is retried, not dropped.

Not checked yet: whether search and Notes recovered after the burst. Repro: apps/website/shots on branch site/showcase-seed (bun run seed --reset).

Found while seeding the website showcase instance (#1143) on 2026-10-05. Build: `dev` 04afb00, server image from the root Containerfile (with the #1143 fixes), arm64, OrbStack. The seed uploads about 300 photos and about 480 files through WebDAV and the API in a few minutes, as a new User would on a first import. Server log during the burst: ``` WARN calternal_server::wire: watched semantic search change notification failed error=semantic index queue is full (many times) WARN calternal_plugin_notes: Notes Index operation failed sqlite_code="5" pool_timeout=false (twice) ERROR calternal_server::wire: job worker failed; restarting error=SQLite operation failed: error returned from database: (code: 5) database is locked ``` Expected: 1. A full semantic index queue does not drop changes. It applies back-pressure or schedules a catch-up scan, so search finds every file after the burst. 2. A busy Index (SQLITE_BUSY) does not crash the job worker. The worker retries with a busy timeout and does not restart. 3. A Notes Index operation that gets SQLITE_BUSY is retried, not dropped. Not checked yet: whether search and Notes recovered after the burst. Repro: `apps/website/shots` on branch `site/showcase-seed` (`bun run seed --reset`).
Author
Owner

Starting #1157 on job/bulkimport-1157 at a493c314edfd40027a72e0d7fc1cb97fa69b1eb9. The worktree tracks origin/dev and is six commits behind at start; I will merge the updated origin/dev before final gates as required. I am tracing semantic notification overflow, Notes Index transactions, and the worker's SQLite error handling before making changes.

Starting #1157 on `job/bulkimport-1157` at `a493c314edfd40027a72e0d7fc1cb97fa69b1eb9`. The worktree tracks `origin/dev` and is six commits behind at start; I will merge the updated `origin/dev` before final gates as required. I am tracing semantic notification overflow, Notes Index transactions, and the worker's SQLite error handling before making changes.
Author
Owner

Trace finding: Db::connect opens separate ordinary and authority writer connections to the same WAL Index, each with a 5 s busy timeout. crates/plugins/files/src/index.rs calls Notes::refresh_target_path after Files Index changes; crates/plugins/notes/src/store.rs starts that operation with deferred begin(), reads within the transaction, then writes. A commit from the other writer pool can stale the WAL snapshot, so the read-to-write upgrade returns SQLITE_BUSY_SNAPSHOT; a longer busy timeout cannot wait out that case. JobQueue::fail has the same deferred read-then-write shape, while Worker::wait_for_writer retries only pool checkout timeouts. Separately, SemanticIndexer::notify_changed/notify_removed return QueueFull on try_send, and the caller logs it; the five-minute reconcile is the only recovery. I will reserve the writer before these read-modify-write transactions, retry complete transient operations, keep worker queue calls alive through BUSY, and turn semantic overflow into one coalesced catch-up scan.

Trace finding: `Db::connect` opens separate ordinary and authority writer connections to the same WAL Index, each with a 5 s busy timeout. `crates/plugins/files/src/index.rs` calls `Notes::refresh_target_path` after Files Index changes; `crates/plugins/notes/src/store.rs` starts that operation with deferred `begin()`, reads within the transaction, then writes. A commit from the other writer pool can stale the WAL snapshot, so the read-to-write upgrade returns `SQLITE_BUSY_SNAPSHOT`; a longer busy timeout cannot wait out that case. `JobQueue::fail` has the same deferred read-then-write shape, while `Worker::wait_for_writer` retries only pool checkout timeouts. Separately, `SemanticIndexer::notify_changed/notify_removed` return `QueueFull` on `try_send`, and the caller logs it; the five-minute reconcile is the only recovery. I will reserve the writer before these read-modify-write transactions, retry complete transient operations, keep worker queue calls alive through BUSY, and turn semantic overflow into one coalesced catch-up scan.
Author
Owner

The first focused adversarial run built and started the real server, then stopped before uploads. The runner uses runpy.run_path("bulk_import_1157.py"), so __file__ was relative; pathlib.Path(__file__).parents[2] raised IndexError while opening the JPEG fixture. I am updating the probe to resolve its own file path before locating the fixture, then I will rerun the focused burst. This attempt produced no product finding because the upload probe did not start.

The first focused adversarial run built and started the real server, then stopped before uploads. The runner uses `runpy.run_path("bulk_import_1157.py")`, so `__file__` was relative; `pathlib.Path(__file__).parents[2]` raised `IndexError` while opening the JPEG fixture. I am updating the probe to resolve its own file path before locating the fixture, then I will rerun the focused burst. This attempt produced no product finding because the upload probe did not start.
Author
Owner

The full probe ran against the real server. With 24 upload workers, the 780 requests finished their first attempts in 90.96 s, but the API returned 429 too_many_requests for active upload admission and several requests timed out. After the 300 s index wait, only 119 import files were on disk and in Files Index; Notes Index, semantic Index and Search had 0 bulk-import paths. The worker restart count was 0 and dropped semantic-change warnings were 0. Host load average was 18.50 before and 24.23 after. The probe exceeded the Files upload admission capacity, so this run does not establish the full import result. I am lowering concurrent uploads and adding bounded retry for explicit 429 admission responses before the next focused run.

The full probe ran against the real server. With 24 upload workers, the 780 requests finished their first attempts in 90.96 s, but the API returned `429 too_many_requests` for active upload admission and several requests timed out. After the 300 s index wait, only 119 import files were on disk and in Files Index; Notes Index, semantic Index and Search had 0 bulk-import paths. The worker restart count was 0 and dropped semantic-change warnings were 0. Host load average was 18.50 before and 24.23 after. The probe exceeded the Files upload admission capacity, so this run does not establish the full import result. I am lowering concurrent uploads and adding bounded retry for explicit 429 admission responses before the next focused run.
Author
Owner

The backpressured real-server probe uploaded all 300 Photos and 480 Markdown files. Disk, Files Index, Notes Index and semantic Index matched the expected paths. The import took 487.596 s and all Indexes were ready at 514.526 s; upload p50/p95 were 3366.38/16240.13 ms, sampled server CPU was 205.01 s, mean/peak RSS was 588398135/700784640 bytes, and host load average was 19.06 before and 13.43 after. The run had 0 upload 429 retries, 0 job-worker restarts and 0 dropped semantic-change warnings. The Files Search assertion returned 0 expected hits out of 780. I am checking whether the probe uses the current Search hit identity/query contract or found a product indexing defect.

The backpressured real-server probe uploaded all 300 Photos and 480 Markdown files. Disk, Files Index, Notes Index and semantic Index matched the expected paths. The import took 487.596 s and all Indexes were ready at 514.526 s; upload p50/p95 were 3366.38/16240.13 ms, sampled server CPU was 205.01 s, mean/peak RSS was 588398135/700784640 bytes, and host load average was 19.06 before and 13.43 after. The run had 0 upload 429 retries, 0 job-worker restarts and 0 dropped semantic-change warnings. The Files Search assertion returned 0 expected hits out of 780. I am checking whether the probe uses the current Search hit identity/query contract or found a product indexing defect.
Author
Owner

Finding: the prior real-server burst had 0/780 probe Search hits, but its request set semantic=false. In crates/calternal-server/src/main.rs, that mode filters providers to plugin_id == "search"; it excludes the Files provider. The expected Files hit was therefore not part of the request contract. The burst had already passed the exact disk, Files Index, Notes Index, and semantic Index checks, with 0 worker restarts and 0 dropped-change warnings. I updated the probe to use the normal hybrid Search route (the default), which queries the Files provider, and am rerunning the burst against the real server.

Finding: the prior real-server burst had 0/780 probe Search hits, but its request set `semantic=false`. In `crates/calternal-server/src/main.rs`, that mode filters providers to `plugin_id == "search"`; it excludes the Files provider. The expected Files hit was therefore not part of the request contract. The burst had already passed the exact disk, Files Index, Notes Index, and semantic Index checks, with 0 worker restarts and 0 dropped-change warnings. I updated the probe to use the normal hybrid Search route (the default), which queries the Files provider, and am rerunning the burst against the real server.
Author
Owner

Finding from the corrected 300 Photo + 480 Markdown live-server run: disk, Files Index, Notes Index and semantic Index reached the expected paths; worker restarts and dropped-change warnings were both 0. On this shared host, load average was 42.9 before the run and 30.3 after it. The probe issued 780 simultaneous Search requests against the route's 200 ms provider deadline; all 780 returned timed_out=true with no hits, and one upload client timed out after the server persisted the item. This is a load-sensitive verification failure, not an index mismatch. I am changing the probe to search the exact filename set in small stable-prefix groups through the normal Search route, so it checks every path without creating a second 780-request storm.

Finding from the corrected 300 Photo + 480 Markdown live-server run: disk, Files Index, Notes Index and semantic Index reached the expected paths; worker restarts and dropped-change warnings were both 0. On this shared host, load average was 42.9 before the run and 30.3 after it. The probe issued 780 simultaneous Search requests against the route's 200 ms provider deadline; all 780 returned `timed_out=true` with no hits, and one upload client timed out after the server persisted the item. This is a load-sensitive verification failure, not an index mismatch. I am changing the probe to search the exact filename set in small stable-prefix groups through the normal Search route, so it checks every path without creating a second 780-request storm.
Author
Owner

Finding in the third real-server burst: after the 300 Photo uploads completed, three expected paths were absent from both disk and Files Index: Photos/BulkImport1157/bulk1157-photos-045.jpg, ...-053.jpg, and ...-265.jpg. Notes and semantic Index counts reached 480. The server log still has zero database is locked, worker-restart, semantic-drop, and upload-admission warnings. I am waiting for the probe's captured HTTP outcomes before attributing these three missing items.

Finding in the third real-server burst: after the 300 Photo uploads completed, three expected paths were absent from both disk and Files Index: `Photos/BulkImport1157/bulk1157-photos-045.jpg`, `...-053.jpg`, and `...-265.jpg`. Notes and semantic Index counts reached 480. The server log still has zero `database is locked`, worker-restart, semantic-drop, and upload-admission warnings. I am waiting for the probe's captured HTTP outcomes before attributing these three missing items.
Author
Owner

The third live-server run returned three HTTP 503 responses on upload creation. Two had empty bodies; one said Authentication database is busy; retry shortly. Those three requests correspond exactly to the three Photo paths absent from disk and Files Index. The probe profile completed the other indexes, with 0 worker restarts and 0 dropped semantic-change warnings. This narrows the remaining contention to transient SQLite BUSY during request authentication, which currently becomes a client-visible 503. I am tracing the AuthStore request path and will add a bounded retry plus a regression test before the next verification pass.

The third live-server run returned three HTTP 503 responses on upload creation. Two had empty bodies; one said `Authentication database is busy; retry shortly`. Those three requests correspond exactly to the three Photo paths absent from disk and Files Index. The probe profile completed the other indexes, with 0 worker restarts and 0 dropped semantic-change warnings. This narrows the remaining contention to transient SQLite BUSY during request authentication, which currently becomes a client-visible 503. I am tracing the AuthStore request path and will add a bounded retry plus a regression test before the next verification pass.
Author
Owner

Completed: #1157 bulk-import contention and index recovery

Head: b0265619929917f9f9e195e4f2f900f403ce7749

What changed

  • Set the SQLite busy timeout on all application pools. The root cause was that ordinary and authority connections share one SQLite writer, while a deferred WAL read transaction can become a stale snapshot before its first write. SQLite returns BUSY for that snapshot upgrade without waiting for busy_timeout. Read-modify-write paths now reserve the writer with BEGIN IMMEDIATE before reading.
  • Kept the job Worker alive through transient BUSY and pool checkout timeouts. It retries complete durable queue operations in bounded jittered batches, so a leased Job is not lost to a Worker restart.
  • Retained Notes Index event work through transient BUSY. Each retry re-reads disk and runs a fresh transaction; the event consumer keeps retrying after the per-pass retry budget.
  • Made semantic queue overflow set a coalesced full-scan marker. Failed incremental batches also request a scan, which repairs the Index from current disk state without growing the queue.
  • Added the real-server #1157 burst probe and its profile reporter. Updated exact performance pins for the added retry paths.

Real-server burst

The probe uploaded 300 Photos and 480 Markdown files through API Tus and File WebDAV, then compared disk, Files Search, Notes Index, and semantic Index paths. Result:

bulk import profile: {"all_indexes_seconds": 612.584, "dropped_change_warnings": 0, "load_average": {"after": [27.625, 20.92041015625, 22.99267578125], "before": [7.37353515625, 15.48779296875, 26.84228515625], "samples": 6171}, "markdown_files": 480, "photos": 300, "search_hits": 780, "search_latency_ms": {"p50": 138.19, "p95": 183.93}, "search_queries": 78, "search_timed_out_requests": 20, "server_cpu_seconds": 228.14, "server_rss_bytes": {"after": 675581952, "mean_sampled": 597091365, "peak_sampled": 675581952}, "upload_429_retries": 0, "upload_latency_ms": {"p50": 4088.79, "p95": 13514.28}, "upload_seconds": 560.399, "upload_workers": 8, "uploads": 780, "worker_restarts": 0}
Bulk import #1157 passed: disk, Search paths, Notes Index and semantic Index match.

The host was shared and load rose during the run, so these are local measurements, not a controlled comparison with the perf-test baseline. Twenty grouped Search requests reported provider timeouts; the probe retried timed-out or incomplete groups and found all 780 expected paths.

Gates

cargo fmt --check passed with no output.

$ cargo clippy -p calternal-db --all-targets -- -D warnings
    Finished `dev` profile [unoptimized + debuginfo] target(s) in 11.11s
$ cargo test -p calternal-db
test result: ok. 36 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 5.91s
test result: ok. 7 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 4.30s
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.20s
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.23s
test result: ok. 6 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.53s
test result: ok. 21 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 1.40s
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.07s

$ cargo clippy -p calternal-plugin-notes --all-targets -- -D warnings
    Finished `dev` profile [unoptimized + debuginfo] target(s) in 1m 29s
$ cargo test -p calternal-plugin-notes
test result: ok. 288 passed; 0 failed; 3 ignored; 0 measured; 0 filtered out; finished in 238.70s
test result: ok. 2 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 1.38s

$ cargo clippy -p calternal-embed --all-targets -- -D warnings
    Finished `dev` profile [unoptimized + debuginfo] target(s) in 34.35s
$ cargo test -p calternal-embed
test result: ok. 39 passed; 0 failed; 4 ignored; 0 measured; 0 filtered out; finished in 24.49s

$ cargo clippy -p calternal-auth --all-targets -- -D warnings
    Finished `dev` profile [unoptimized + debuginfo] target(s) in 11.83s
$ cargo test -p calternal-auth
test result: ok. 124 passed; 0 failed; 3 ignored; 0 measured; 0 filtered out; finished in 55.57s

$ cd apps/web && bun run check
perf-lint: PASS; 0 violations; 22104 scoped exceptions
svelte-check found 0 errors and 4 warnings in 3 files
$ cd apps/web && bun run test
 Test Files  266 passed (266)
      Tests  1827 passed (1827)

The Notes, DB, Embed, and Auth test commands also reported passing doc tests (zero doc tests in each applicable binary). The four Svelte warnings are existing empty/unused CSS selectors.

Files

  • crates/calternal-db/src/{db,jobs,sqlite,worker}.rs
  • crates/plugins/notes/src/{lib,store,tasks_store}.rs
  • crates/calternal-embed/src/store.rs
  • crates/calternal-auth/src/api.rs
  • tests/adversarial/{bulk_import_1157.py,run.sh,setup.mjs,webdav.py}
  • bench/bulk-import-1157.py
  • contracts/perf/{exceptions,adoption-1058}.json

Decisions and gaps

  • Used a one-bit watch marker for semantic overflow. A marker coalesces any burst size while keeping server writes non-blocking and bounded; the Worker scans current disk after queued events drain.
  • Kept transient Notes event retries alive with a 500 ms pause after each three-attempt transaction retry batch. Used 50 ms and 250 ms jitter caps for Worker BUSY batches and authority lookup retries.
  • Added a bounded retry around session authority lookup because the live import exposed transient authentication 503s under pool contention; no API response contract changed.
  • The only observed verification limitation was the 20 timed-out grouped Search requests noted above; retries returned every expected path. No worker restart or dropped-change warning occurred.

cargo clean removed 20,422 files (16.2 GiB); web build output was removed. No push, deploy, or merge was performed.

## Completed: #1157 bulk-import contention and index recovery Head: `b0265619929917f9f9e195e4f2f900f403ce7749` ### What changed - Set the SQLite busy timeout on all application pools. The root cause was that ordinary and authority connections share one SQLite writer, while a deferred WAL read transaction can become a stale snapshot before its first write. SQLite returns `BUSY` for that snapshot upgrade without waiting for `busy_timeout`. Read-modify-write paths now reserve the writer with `BEGIN IMMEDIATE` before reading. - Kept the job Worker alive through transient BUSY and pool checkout timeouts. It retries complete durable queue operations in bounded jittered batches, so a leased Job is not lost to a Worker restart. - Retained Notes Index event work through transient BUSY. Each retry re-reads disk and runs a fresh transaction; the event consumer keeps retrying after the per-pass retry budget. - Made semantic queue overflow set a coalesced full-scan marker. Failed incremental batches also request a scan, which repairs the Index from current disk state without growing the queue. - Added the real-server #1157 burst probe and its profile reporter. Updated exact performance pins for the added retry paths. ### Real-server burst The probe uploaded 300 Photos and 480 Markdown files through API Tus and File WebDAV, then compared disk, Files Search, Notes Index, and semantic Index paths. Result: ```text bulk import profile: {"all_indexes_seconds": 612.584, "dropped_change_warnings": 0, "load_average": {"after": [27.625, 20.92041015625, 22.99267578125], "before": [7.37353515625, 15.48779296875, 26.84228515625], "samples": 6171}, "markdown_files": 480, "photos": 300, "search_hits": 780, "search_latency_ms": {"p50": 138.19, "p95": 183.93}, "search_queries": 78, "search_timed_out_requests": 20, "server_cpu_seconds": 228.14, "server_rss_bytes": {"after": 675581952, "mean_sampled": 597091365, "peak_sampled": 675581952}, "upload_429_retries": 0, "upload_latency_ms": {"p50": 4088.79, "p95": 13514.28}, "upload_seconds": 560.399, "upload_workers": 8, "uploads": 780, "worker_restarts": 0} Bulk import #1157 passed: disk, Search paths, Notes Index and semantic Index match. ``` The host was shared and load rose during the run, so these are local measurements, not a controlled comparison with the perf-test baseline. Twenty grouped Search requests reported provider timeouts; the probe retried timed-out or incomplete groups and found all 780 expected paths. ### Gates `cargo fmt --check` passed with no output. ```text $ cargo clippy -p calternal-db --all-targets -- -D warnings Finished `dev` profile [unoptimized + debuginfo] target(s) in 11.11s $ cargo test -p calternal-db test result: ok. 36 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 5.91s test result: ok. 7 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 4.30s test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.20s test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.23s test result: ok. 6 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.53s test result: ok. 21 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 1.40s test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.07s $ cargo clippy -p calternal-plugin-notes --all-targets -- -D warnings Finished `dev` profile [unoptimized + debuginfo] target(s) in 1m 29s $ cargo test -p calternal-plugin-notes test result: ok. 288 passed; 0 failed; 3 ignored; 0 measured; 0 filtered out; finished in 238.70s test result: ok. 2 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 1.38s $ cargo clippy -p calternal-embed --all-targets -- -D warnings Finished `dev` profile [unoptimized + debuginfo] target(s) in 34.35s $ cargo test -p calternal-embed test result: ok. 39 passed; 0 failed; 4 ignored; 0 measured; 0 filtered out; finished in 24.49s $ cargo clippy -p calternal-auth --all-targets -- -D warnings Finished `dev` profile [unoptimized + debuginfo] target(s) in 11.83s $ cargo test -p calternal-auth test result: ok. 124 passed; 0 failed; 3 ignored; 0 measured; 0 filtered out; finished in 55.57s $ cd apps/web && bun run check perf-lint: PASS; 0 violations; 22104 scoped exceptions svelte-check found 0 errors and 4 warnings in 3 files $ cd apps/web && bun run test Test Files 266 passed (266) Tests 1827 passed (1827) ``` The Notes, DB, Embed, and Auth test commands also reported passing doc tests (zero doc tests in each applicable binary). The four Svelte warnings are existing empty/unused CSS selectors. ### Files - `crates/calternal-db/src/{db,jobs,sqlite,worker}.rs` - `crates/plugins/notes/src/{lib,store,tasks_store}.rs` - `crates/calternal-embed/src/store.rs` - `crates/calternal-auth/src/api.rs` - `tests/adversarial/{bulk_import_1157.py,run.sh,setup.mjs,webdav.py}` - `bench/bulk-import-1157.py` - `contracts/perf/{exceptions,adoption-1058}.json` ### Decisions and gaps - Used a one-bit watch marker for semantic overflow. A marker coalesces any burst size while keeping server writes non-blocking and bounded; the Worker scans current disk after queued events drain. - Kept transient Notes event retries alive with a 500 ms pause after each three-attempt transaction retry batch. Used 50 ms and 250 ms jitter caps for Worker BUSY batches and authority lookup retries. - Added a bounded retry around session authority lookup because the live import exposed transient authentication 503s under pool contention; no API response contract changed. - The only observed verification limitation was the 20 timed-out grouped Search requests noted above; retries returned every expected path. No worker restart or dropped-change warning occurred. `cargo clean` removed 20,422 files (16.2 GiB); web build output was removed. No push, deploy, or merge was performed.
Author
Owner

Independent review started on job/bulkimport-1157, head b026561992; origin/dev d0061ec3df. Reviewing recovery, transaction boundaries and retry idempotency; adding regression tests and running the real-server burst and 700-Note reconcile. Finding: semantic recovery uses only an in-memory watch bit and waits for an empty queue. Restart loses the bit and sustained arrivals delay recovery. No push or deploy.

Independent review started on job/bulkimport-1157, head b0265619929917f9f9e195e4f2f900f403ce7749; origin/dev d0061ec3df127d81c86729d07d899b0bf2b6de91. Reviewing recovery, transaction boundaries and retry idempotency; adding regression tests and running the real-server burst and 700-Note reconcile. Finding: semantic recovery uses only an in-memory watch bit and waits for an empty queue. Restart loses the bit and sustained arrivals delay recovery. No push or deploy.
Author
Owner

Review findings at b02656199: runtime-only semantic marker, queue-empty recovery starvation, stale queued deletion after a missed recreate, and swallowed file preparation/second-fingerprint-race failures. Fixes add one permanent synced recovery witness through minimal calternal-fs Root hooks, alternate scan/bounded queue work, reconcile deletes from current disk, and return incomplete-batch errors. First accepted event persists the witness once; subsequent events do no filesystem write. Rust gates are compiling in the fresh job target. The 700-Note test source uses direct reconcile calls with no test-only retry. Live probe now samples ordinary authenticated Files reads during uploads.

Review findings at b02656199: runtime-only semantic marker, queue-empty recovery starvation, stale queued deletion after a missed recreate, and swallowed file preparation/second-fingerprint-race failures. Fixes add one permanent synced recovery witness through minimal calternal-fs Root hooks, alternate scan/bounded queue work, reconcile deletes from current disk, and return incomplete-batch errors. First accepted event persists the witness once; subsequent events do no filesystem write. Rust gates are compiling in the fresh job target. The 700-Note test source uses direct reconcile calls with no test-only retry. Live probe now samples ordinary authenticated Files reads during uploads.
Author
Owner

Safety fix committed as fc6e3c7c8 (Root hook afeabc3b7). The real pinned-model regression passed in 33.53 s: continuous overflow reaches both current files; an old delete hint preserves a recreated file; preparation SQL errors return failure; cold restart repairs changed bytes with no notification; one failed private Index does not starve another Home. Ordinary Embed tests: 42 passed, 5 ignored (the added real-model test was then run explicitly and passed). calternal-fs clippy passed; its test binaries passed 94, 1, 1, 48 and 2 tests. calternal-db clippy passed; its test binaries passed 38, 7, 1, 1, 6, 21 and 1 tests. A notification-count assertion is now in focused verification. Full web attempt: 265 files passed, one unchanged InfoPanel test hit its existing 5000 ms timeout; no assertion failure. Preserved logs; one focused confirmation and one full confirmation use maxWorkers=2. Real server build and 700-Note check remain in progress.

Safety fix committed as fc6e3c7c8 (Root hook afeabc3b7). The real pinned-model regression passed in 33.53 s: continuous overflow reaches both current files; an old delete hint preserves a recreated file; preparation SQL errors return failure; cold restart repairs changed bytes with no notification; one failed private Index does not starve another Home. Ordinary Embed tests: 42 passed, 5 ignored (the added real-model test was then run explicitly and passed). calternal-fs clippy passed; its test binaries passed 94, 1, 1, 48 and 2 tests. calternal-db clippy passed; its test binaries passed 38, 7, 1, 1, 6, 21 and 1 tests. A notification-count assertion is now in focused verification. Full web attempt: 265 files passed, one unchanged InfoPanel test hit its existing 5000 ms timeout; no assertion failure. Preserved logs; one focused confirmation and one full confirmation use maxWorkers=2. Real server build and 700-Note check remain in progress.
Author
Owner

The #1163 700-Note regression passed once with its unchanged direct reconcile calls and no test-only retry:

$ cargo test -p calternal-plugin-notes seven_hundred_notes_reconcile_without_feedback -- --test-threads=1
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 290 filtered out; finished in 116.46s

It retains 700 initial per-source notices and silence on unchanged repair. Full web confirmation also passed after the initial 5000 ms InfoPanel timeout (the focused file passed unchanged):

$ cd apps/web && bun run check
perf-lint: PASS; 0 violations; 22104 scoped exceptions
svelte-check found 0 errors and 4 warnings in 3 files
$ cd apps/web && bun run test --maxWorkers=2
 Test Files  266 passed (266)
      Tests  1830 passed (1830)

Current head 24a90faf03. The patched real-server build and one time-boxed import burst remain.

The #1163 700-Note regression passed once with its unchanged direct reconcile calls and no test-only retry: ```text $ cargo test -p calternal-plugin-notes seven_hundred_notes_reconcile_without_feedback -- --test-threads=1 test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 290 filtered out; finished in 116.46s ``` It retains 700 initial per-source notices and silence on unchanged repair. Full web confirmation also passed after the initial 5000 ms InfoPanel timeout (the focused file passed unchanged): ```text $ cd apps/web && bun run check perf-lint: PASS; 0 violations; 22104 scoped exceptions svelte-check found 0 errors and 4 warnings in 3 files $ cd apps/web && bun run test --maxWorkers=2 Test Files 266 passed (266) Tests 1830 passed (1830) ``` Current head 24a90faf03e84a5a60a66ea62601af68d292bbee. The patched real-server build and one time-boxed import burst remain.
Author
Owner

Additional review finding: configure in crates/plugins/notes/src/lib.rs acquired lock_user(&user) outside retry_notes_index_until_ready. That helper can retain BUSY work across arbitrarily many batches and sleeps 500 ms between exhausted batches. The Notes User writer guard therefore stayed held through the outer sleeps. This is separate from the SQLite writer pool, which the earlier settlement test proved is released. Fix in progress: scope the User guard to each adoption/target-refresh pass and re-read disk on the next pass; retain the existing bounded inner projection retries. A new regression holds the authority transaction, waits for a failed pass, checks a foreground Notes writer can take its User lock before SQLite is released, then checks one repaired Note row. Full Notes clippy/test are queued; the live profile now alternates normal Notes creation and Files folder creation alongside ordinary Files reads.

Additional review finding: `configure` in crates/plugins/notes/src/lib.rs acquired `lock_user(&user)` outside `retry_notes_index_until_ready`. That helper can retain BUSY work across arbitrarily many batches and sleeps 500 ms between exhausted batches. The Notes User writer guard therefore stayed held through the outer sleeps. This is separate from the SQLite writer pool, which the earlier settlement test proved is released. Fix in progress: scope the User guard to each adoption/target-refresh pass and re-read disk on the next pass; retain the existing bounded inner projection retries. A new regression holds the authority transaction, waits for a failed pass, checks a foreground Notes writer can take its User lock before SQLite is released, then checks one repaired Note row. Full Notes clippy/test are queued; the live profile now alternates normal Notes creation and Files folder creation alongside ordinary Files reads.
Author
Owner

Independent review progress at c766ce589 (base author head b0265619929917f9f9e195e4f2f900f403ce7749; merged origin/dev d0061ec3df127d81c86729d07d899b0bf2b6de91 once).

Fixed recovery defects: lost restart marker, queue-empty starvation, stale deletion after recreation, swallowed partial-scan errors, and failed first witness write without runtime repair. Added a confined synced restart witness and real-model overflow/recreation/restart regression. Also moved the Notes User writer guard inside each retained retry pass so outer backoff releases it.

Evidence: FS, DB, Embed and Notes clippy/tests pass. Notes: test result: ok. 289 passed; 0 failed; 3 ignored; 0 measured; 0 filtered out; finished in 387.81s, including the 700-Note reconcile without a test-only retry. Latest Embed: test result: ok. 43 passed; 0 failed; 5 ignored; 0 measured; 0 filtered out; finished in 60.88s. Explicit real-model regression passed in 33.53s. DB regressions check one settled outcome/notification, writer-pool release during BUSY, and immediate missing-schema failure.

Full web suite passed: 266 files / 1830 tests, after one unchanged InfoPanel timeout on the shared host; focused InfoPanel and full confirmation both passed without changing expectations. Web check: zero errors, four existing warnings.

The real-server 300-Photo + 480-Markdown burst is running now, with ordinary Files reads, Files writes and Notes writes sampled during upload. Final verdict waits for its disk/Index/Search comparisons and restart/drop counters. Decisions and findings are in review-1157.md.

Independent review progress at `c766ce589` (base author head `b0265619929917f9f9e195e4f2f900f403ce7749`; merged origin/dev `d0061ec3df127d81c86729d07d899b0bf2b6de91` once). Fixed recovery defects: lost restart marker, queue-empty starvation, stale deletion after recreation, swallowed partial-scan errors, and failed first witness write without runtime repair. Added a confined synced restart witness and real-model overflow/recreation/restart regression. Also moved the Notes User writer guard inside each retained retry pass so outer backoff releases it. Evidence: FS, DB, Embed and Notes clippy/tests pass. Notes: `test result: ok. 289 passed; 0 failed; 3 ignored; 0 measured; 0 filtered out; finished in 387.81s`, including the 700-Note reconcile without a test-only retry. Latest Embed: `test result: ok. 43 passed; 0 failed; 5 ignored; 0 measured; 0 filtered out; finished in 60.88s`. Explicit real-model regression passed in 33.53s. DB regressions check one settled outcome/notification, writer-pool release during BUSY, and immediate missing-schema failure. Full web suite passed: 266 files / 1830 tests, after one unchanged InfoPanel timeout on the shared host; focused InfoPanel and full confirmation both passed without changing expectations. Web check: zero errors, four existing warnings. The real-server 300-Photo + 480-Markdown burst is running now, with ordinary Files reads, Files writes and Notes writes sampled during upload. Final verdict waits for its disk/Index/Search comparisons and restart/drop counters. Decisions and findings are in `review-1157.md`.
Author
Owner

The independent live burst passed on code head c766ce589: 300 Photos + 480 Markdown files through API Tus and WebDAV, eight upload workers. Search returned all 780 paths. Disk and Index sets matched; a separate read-only row check found 480 Note rows, zero duplicate Note paths, 780 Files rows, and 480 semantic document rows. Note identity/title mismatches: 0. Worker restarts: 0. Dropped-change warnings: 0.

Ordinary requests during upload all succeeded. Files reads: 458 samples, p50 145.33 ms / p95 1860.73 ms. Files folder writes: 25 samples, p50 1028.16 ms / p95 5236.82 ms. Notes writes: 25 samples, p50 223.96 ms / p95 842.89 ms. Upload elapsed: 398.815 s; all Indexes current: 448.079 s. Upload p50/p95: 2620.76 / 14082.41 ms. Server CPU: 260.81 s; mean/peak RSS: 587157014 / 655114240 bytes. Shared-host load average rose from 10.09 to 27.76; gate compilation also ran during the burst. These are local debug-server measurements, not a controlled baseline comparison.

Reference in docs/perf/baseline.json: search_100k.freshness_storm, perf-test, 1000 writes, upload elapsed 373.45 s, p50/p95 13643.83 / 15972.65 ms, mean/peak RSS 952356266 / 983887872 bytes. Fixture sizes and hosts differ.

One reporting defect found: Search latency samples were not sorted before percentile lookup. Fixed separately at 378ea4358. The completed run's Search percentile pair is invalid and is excluded from evidence. Its 780-path assertions and all ordinary-request/upload percentiles are valid. Search logged 32 timed-out first requests but bounded Search retries reached every file. No test expectation changed; no second burst was run.

Auth clippy/tests passed (124 tests). Server crate gates are still running; the final verdict follows them. The review and regression evidence are retained locally in review-1157.md and ignored artifacts/.

The independent live burst passed on code head `c766ce589`: 300 Photos + 480 Markdown files through API Tus and WebDAV, eight upload workers. Search returned all 780 paths. Disk and Index sets matched; a separate read-only row check found 480 Note rows, zero duplicate Note paths, 780 Files rows, and 480 semantic document rows. Note identity/title mismatches: 0. Worker restarts: 0. Dropped-change warnings: 0. Ordinary requests during upload all succeeded. Files reads: 458 samples, p50 145.33 ms / p95 1860.73 ms. Files folder writes: 25 samples, p50 1028.16 ms / p95 5236.82 ms. Notes writes: 25 samples, p50 223.96 ms / p95 842.89 ms. Upload elapsed: 398.815 s; all Indexes current: 448.079 s. Upload p50/p95: 2620.76 / 14082.41 ms. Server CPU: 260.81 s; mean/peak RSS: 587157014 / 655114240 bytes. Shared-host load average rose from 10.09 to 27.76; gate compilation also ran during the burst. These are local debug-server measurements, not a controlled baseline comparison. Reference in `docs/perf/baseline.json`: `search_100k.freshness_storm`, perf-test, 1000 writes, upload elapsed 373.45 s, p50/p95 13643.83 / 15972.65 ms, mean/peak RSS 952356266 / 983887872 bytes. Fixture sizes and hosts differ. One reporting defect found: Search latency samples were not sorted before percentile lookup. Fixed separately at `378ea4358`. The completed run's Search percentile pair is invalid and is excluded from evidence. Its 780-path assertions and all ordinary-request/upload percentiles are valid. Search logged 32 timed-out first requests but bounded Search retries reached every file. No test expectation changed; no second burst was run. Auth clippy/tests passed (124 tests). Server crate gates are still running; the final verdict follows them. The review and regression evidence are retained locally in `review-1157.md` and ignored `artifacts/`.
Author
Owner

Independent review verdict: SAFE TO MERGE — yes.

Branch: job/bulkimport-1157
HEAD: a82f2e5679d6795b287924fe465f027dbfe4b655
Merged origin/dev once: d0061ec3df127d81c86729d07d899b0bf2b6de91

Independent data-safety review of #1157

Reviewed input head: b0265619929917f9f9e195e4f2f900f403ce7749.
The branch merged current origin/dev once before review gates.

Findings

  1. The semantic recovery marker was only in memory. A restart with no new
    upsert did not request model loading or an immediate scan. Fixed: a fixed,
    private, synced .system/semantic-recovery witness arms recovery on restart.
    The first accepted event persists it; later events do no filesystem write.
    The witness stays after successful repair. This also covers queued events
    interrupted by a crash, and avoids a disk clear race with a new request.

  2. Recovery waited for an empty queue. Continuous arrivals could defer it.
    Fixed: alternate a recovery pass and one bounded incremental batch when
    both are pending. A scan clears only the runtime bit before it starts.

  3. An old queued deletion could erase a recreated file after a full scan,
    when its new upsert had overflowed. Fixed: loaded workers reconcile both
    event types from current disk. Before model readiness, an existing target
    requests recovery instead of deleting its vectors.

  4. File preparation errors were logged inside index_files but returned
    success. A second fingerprint race also cleared vectors and returned success.
    Fixed: report an incomplete batch after applying its other files. Catch-up
    retains failed work without needing a second watcher event.

  5. Worker and Notes event retries have bounded batches, but no total attempt
    limit for BUSY. This retains durable work through contention. Other Index
    errors return. A missing-schema test checks that worker errors surface.

  6. The Files event consumer held the Notes User writer lock across outer
    retries and sleeps. Fixed: take that lock inside each retry pass. A failed
    pass releases it before the next outer pause. A focused test checks that a
    foreground editor can take it while SQLite remains busy, then repair stores
    one Note projection after authority releases SQLite.

  7. A failed first recovery witness write returned before runtime repair was
    armed. Fixed: surface the filesystem error, set the coalesced runtime bit,
    and request the model. A confined I/O failure regression checks all three.

  8. The Search latency reporter used unsorted samples. Fixed: sort them before
    percentile lookup. The completed burst has valid ordinary-request and upload
    percentiles, but its Search percentile pair is not used as evidence.

Decisions

Keep one permanent restart witness after the first accepted semantic event.
Restart scans are conservative; successful live scans clear only the runtime
bit. No disk write is needed per event after the witness has been synced.
This adds a small confined Root hook; no other filesystem behavior changes.

Alternate a full scan with one bounded event batch when both are pending.
Continue other batches and Homes after a partial error, but do not prune an
incomplete scan. Surface failed witness writes and arm runtime recovery.
Release the Notes User guard between retained retry passes. Keep existing
bounded projection retries inside a pass so its disk view stays consistent.
Retain durable background work across bounded BUSY batches; do not apply that
open-ended policy to HTTP request retries.

Verification

The local real-server burst passed with 300 Photos and 480 Markdown files,
through API Tus and WebDAV, with eight upload workers. Search found all 780
paths. Disk, Notes, Files and semantic Index path sets matched. Note identities
and titles matched the final disk bytes. Worker restarts and dropped-change
warnings were zero. A separate read-only check found 480 Note rows, zero
duplicate Note paths, 780 Files rows and 480 semantic document rows.

The 700-Note reconcile test passed in isolation and in the full Notes suite.
Its direct reconcile calls remain unchanged. No test-only retry was added.
The real-model regression passed. It covered continuous overflow, stale
removal after recreation, duplicate vector checks, cold restart without a new
notification, and recovery of another Home after a partial failure.

Worker tests check one stored outcome and one committed notification, writer
pool release during BUSY, and immediate error return for a missing schema.
Only BUSY and pool checkout timeouts are retained by the Worker. Each retry
batch is bounded, with capped jittered delay. Durable background work has no
total BUSY attempt limit. Other errors return. This policy prevents lost work
and Worker restarts during temporary contention.

The new BEGIN IMMEDIATE sites contain Index reads and writes. File reads,
model inference and Note preparation happen outside these transactions. No
network or heavy filesystem work was added inside them. Notes source-change
notifications check the committed revision. Link-health notifications publish
changed resolutions after commit.

Local performance evidence

The server used the current Rust debug build and a production web build.
The host was shared. Gate compilation ran during the burst. Load average rose
from 10.09 to 27.76. These numbers are not a controlled baseline comparison.

Ordinary request Samples Errors p50 ms p95 ms
Files entries read 458 0 145.33 1860.73
Files folder write 25 0 1028.16 5236.82
Notes write 25 0 223.96 842.89

Upload elapsed: 398.815 s. All Indexes were current at 448.079 s. Upload p50
and p95 were 2620.76 and 14082.41 ms. Server CPU was 260.81 s. Mean and peak
sampled RSS were 587157014 and 655114240 bytes. No upload retry was needed.
Search had 32 timed-out first requests. Bounded Search retries found all files.
The Search percentile pair from this run is invalid because its samples were
not sorted; the separate reporter fix sorts future samples. The ordinary and
upload percentile samples were sorted and are valid. One burst was run.

The closest reference in docs/perf/baseline.json is
search_100k.freshness_storm: perf-test, 1000 writes, upload elapsed 373.45 s,
upload p50/p95 13643.83/15972.65 ms, mean/peak RSS 952356266/983887872 bytes.
The reference has a different data set, host and build profile. It does not
establish a regression for the ordinary request figures above.

Files changed by this review

  • crates/calternal-fs/src/root.rs: fixed, private, synced recovery witness.
  • crates/calternal-embed/src/store.rs: restart recovery, scan fairness,
    current-disk deletion handling, partial-error propagation and regressions.
  • crates/calternal-db/src/worker.rs: fatal-error, writer-release and single
    outcome/notification regression tests.
  • crates/plugins/notes/src/lib.rs: User lock scope and regression.
  • tests/adversarial/bulk_import_1157.py: ordinary request profile, Note
    metadata comparison and ordered Search percentile reporting.
  • review-1157.md: findings, decisions and verification evidence.

No dependency, migration, route or UI screen was added by this review.
All changed doc comments were read again before this report.

Known gaps

No unresolved data-safety finding remains in the completed scenarios.
The Search latency reporter fix was syntax checked after the single burst;
that burst's invalid Search percentile pair was not reused. Existing ignored
model tests were not all run; the new real-model regression was run explicitly.
Web check has four warnings in three files. An unchanged InfoPanel test hit
its existing five-second timeout on the first full web run. Its focused test
and the full confirmation passed without changes to expectations.
No UI screen changed, so UX walkthroughs and screenshots do not apply.
The first server suite invocation passed units and the performance guard, but
cleanup removed its last integration executable before it could run. The
missing target passed after rebuild. The full per-crate command was then run
again to obtain one complete zero-exit gate. This was a cleanup error, not a
changed test expectation or a product failure.

Gate output (verbatim)

cargo fmt --check exited 0 with no output.

calternal-fs

    Finished `dev` profile [unoptimized + debuginfo] target(s) in 12m 10s
    Finished `test` profile [unoptimized + debuginfo] target(s) in 1m 12s
test result: ok. 94 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 38.73s
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.10s
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.62s
test result: ok. 48 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 9.76s
test result: ok. 2 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.17s
test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s

calternal-db

    Finished `dev` profile [unoptimized + debuginfo] target(s) in 1m 12s
    Finished `test` profile [unoptimized + debuginfo] target(s) in 54.14s
test result: ok. 38 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 6.67s
test result: ok. 7 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 4.52s
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.26s
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.37s
test result: ok. 6 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.88s
test result: ok. 21 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 4.43s
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.19s
test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s

DB single-notice regression

    Finished `dev` profile [unoptimized + debuginfo] target(s) in 3m 57s
    Finished `test` profile [unoptimized + debuginfo] target(s) in 1m 05s
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 37 filtered out; finished in 1.01s
test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 7 filtered out; finished in 0.00s
test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 1 filtered out; finished in 0.00s
test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 1 filtered out; finished in 0.00s
test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 6 filtered out; finished in 0.00s
test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 22 filtered out; finished in 0.00s
test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 1 filtered out; finished in 0.00s

calternal-embed

    Finished `dev` profile [unoptimized + debuginfo] target(s) in 25.05s
    Finished `test` profile [unoptimized + debuginfo] target(s) in 7.88s
test result: ok. 43 passed; 0 failed; 5 ignored; 0 measured; 0 filtered out; finished in 60.88s
test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.01s

real-model regression

    Finished `test` profile [unoptimized + debuginfo] target(s) in 4m 08s
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 46 filtered out; finished in 33.53s

calternal-plugin-notes

    Finished `dev` profile [unoptimized + debuginfo] target(s) in 12m 15s
    Finished `test` profile [unoptimized + debuginfo] target(s) in 1m 01s
test result: ok. 289 passed; 0 failed; 3 ignored; 0 measured; 0 filtered out; finished in 387.81s
test result: ok. 2 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 2.37s
test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s

700-Note reconcile

    Finished `test` profile [unoptimized + debuginfo] target(s) in 10m 05s
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 290 filtered out; finished in 116.46s
test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 2 filtered out; finished in 0.00s

calternal-auth

    Finished `dev` profile [unoptimized + debuginfo] target(s) in 1m 31s
    Finished `test` profile [unoptimized + debuginfo] target(s) in 1m 49s
test result: ok. 124 passed; 0 failed; 3 ignored; 0 measured; 0 filtered out; finished in 113.77s
test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s

cd apps/web && bun run check

perf-lint: PASS; 0 violations; 22104 scoped exceptions
svelte-check found 0 errors and 4 warnings in 3 files

cd apps/web && bun run test --maxWorkers=2

 Test Files  266 passed (266)
      Tests  1830 passed (1830)

ADVERSARIAL_BULK_IMPORT_1157_ONLY=1 ADVERSARIAL_SERVER_BIN="$CARGO_TARGET_DIR/debug/calternal-server" ADVERSARIAL_SKIP_WEB_BUILD=1 ADVERSARIAL_KEEP_WORK_DIR=1 timeout --signal=TERM --kill-after=75s 1800s bash tests/adversarial/run.sh

Bulk import #1157 passed: disk, Search paths, Notes Index and semantic Index match.

Probe and benchmark JSON, row checks and gate logs are in ignored artifacts/.
cargo clippy -p calternal-server --all-targets -- -D warnings

    Finished `dev` profile [unoptimized + debuginfo] target(s) in 5m 48s

cargo test -p calternal-server -- --test-threads=4 (complete invocation, exit 0)

    Finished `test` profile [unoptimized + debuginfo] target(s) in 1m 43s
test result: ok. 256 passed; 0 failed; 10 ignored; 0 measured; 0 filtered out; finished in 151.49s
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 36.96s
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.16s

Final cargo fmt --check exited 0 with no output. Final cargo clean exited 0:

     Removed 8698 files, 7.9GiB total

Web build and .svelte-kit/output were removed after all gates completed.
The disposable live-server data was removed after retaining its profile,
row checks and redacted log. The final worktree has only committed changes.

Verdict

SAFE TO MERGE: yes. The input branch had clear recovery defects. The review
fixed them with separate commits and regression tests. The real-model recovery
round, forced-BUSY worker tests, 700-Note reconcile and live mixed burst passed.
All six reviewed Rust crates passed clippy and tests. Web check and the full
web suite passed. No data loss, duplicate projection, Worker restart or dropped
change warning was found in the completed final scenarios.

Behavior tested at c766ce589; the reporting-only Search percentile fix is
378ea4358. Later commits only record review evidence. No push or deploy was
made. The required merge from origin/dev was done once before gates.

**Independent review verdict: SAFE TO MERGE — yes.** Branch: `job/bulkimport-1157` HEAD: `a82f2e5679d6795b287924fe465f027dbfe4b655` Merged origin/dev once: `d0061ec3df127d81c86729d07d899b0bf2b6de91` # Independent data-safety review of #1157 Reviewed input head: `b0265619929917f9f9e195e4f2f900f403ce7749`. The branch merged current `origin/dev` once before review gates. ## Findings 1. The semantic recovery marker was only in memory. A restart with no new upsert did not request model loading or an immediate scan. Fixed: a fixed, private, synced `.system/semantic-recovery` witness arms recovery on restart. The first accepted event persists it; later events do no filesystem write. The witness stays after successful repair. This also covers queued events interrupted by a crash, and avoids a disk clear race with a new request. 2. Recovery waited for an empty queue. Continuous arrivals could defer it. Fixed: alternate a recovery pass and one bounded incremental batch when both are pending. A scan clears only the runtime bit before it starts. 3. An old queued deletion could erase a recreated file after a full scan, when its new upsert had overflowed. Fixed: loaded workers reconcile both event types from current disk. Before model readiness, an existing target requests recovery instead of deleting its vectors. 4. File preparation errors were logged inside `index_files` but returned success. A second fingerprint race also cleared vectors and returned success. Fixed: report an incomplete batch after applying its other files. Catch-up retains failed work without needing a second watcher event. 5. Worker and Notes event retries have bounded batches, but no total attempt limit for BUSY. This retains durable work through contention. Other Index errors return. A missing-schema test checks that worker errors surface. 6. The Files event consumer held the Notes User writer lock across outer retries and sleeps. Fixed: take that lock inside each retry pass. A failed pass releases it before the next outer pause. A focused test checks that a foreground editor can take it while SQLite remains busy, then repair stores one Note projection after authority releases SQLite. 7. A failed first recovery witness write returned before runtime repair was armed. Fixed: surface the filesystem error, set the coalesced runtime bit, and request the model. A confined I/O failure regression checks all three. 8. The Search latency reporter used unsorted samples. Fixed: sort them before percentile lookup. The completed burst has valid ordinary-request and upload percentiles, but its Search percentile pair is not used as evidence. ## Decisions Keep one permanent restart witness after the first accepted semantic event. Restart scans are conservative; successful live scans clear only the runtime bit. No disk write is needed per event after the witness has been synced. This adds a small confined Root hook; no other filesystem behavior changes. Alternate a full scan with one bounded event batch when both are pending. Continue other batches and Homes after a partial error, but do not prune an incomplete scan. Surface failed witness writes and arm runtime recovery. Release the Notes User guard between retained retry passes. Keep existing bounded projection retries inside a pass so its disk view stays consistent. Retain durable background work across bounded BUSY batches; do not apply that open-ended policy to HTTP request retries. ## Verification The local real-server burst passed with 300 Photos and 480 Markdown files, through API Tus and WebDAV, with eight upload workers. Search found all 780 paths. Disk, Notes, Files and semantic Index path sets matched. Note identities and titles matched the final disk bytes. Worker restarts and dropped-change warnings were zero. A separate read-only check found 480 Note rows, zero duplicate Note paths, 780 Files rows and 480 semantic document rows. The 700-Note reconcile test passed in isolation and in the full Notes suite. Its direct reconcile calls remain unchanged. No test-only retry was added. The real-model regression passed. It covered continuous overflow, stale removal after recreation, duplicate vector checks, cold restart without a new notification, and recovery of another Home after a partial failure. Worker tests check one stored outcome and one committed notification, writer pool release during BUSY, and immediate error return for a missing schema. Only BUSY and pool checkout timeouts are retained by the Worker. Each retry batch is bounded, with capped jittered delay. Durable background work has no total BUSY attempt limit. Other errors return. This policy prevents lost work and Worker restarts during temporary contention. The new `BEGIN IMMEDIATE` sites contain Index reads and writes. File reads, model inference and Note preparation happen outside these transactions. No network or heavy filesystem work was added inside them. Notes source-change notifications check the committed revision. Link-health notifications publish changed resolutions after commit. ### Local performance evidence The server used the current Rust debug build and a production web build. The host was shared. Gate compilation ran during the burst. Load average rose from 10.09 to 27.76. These numbers are not a controlled baseline comparison. | Ordinary request | Samples | Errors | p50 ms | p95 ms | | --- | ---: | ---: | ---: | ---: | | Files entries read | 458 | 0 | 145.33 | 1860.73 | | Files folder write | 25 | 0 | 1028.16 | 5236.82 | | Notes write | 25 | 0 | 223.96 | 842.89 | Upload elapsed: 398.815 s. All Indexes were current at 448.079 s. Upload p50 and p95 were 2620.76 and 14082.41 ms. Server CPU was 260.81 s. Mean and peak sampled RSS were 587157014 and 655114240 bytes. No upload retry was needed. Search had 32 timed-out first requests. Bounded Search retries found all files. The Search percentile pair from this run is invalid because its samples were not sorted; the separate reporter fix sorts future samples. The ordinary and upload percentile samples were sorted and are valid. One burst was run. The closest reference in `docs/perf/baseline.json` is `search_100k.freshness_storm`: perf-test, 1000 writes, upload elapsed 373.45 s, upload p50/p95 13643.83/15972.65 ms, mean/peak RSS 952356266/983887872 bytes. The reference has a different data set, host and build profile. It does not establish a regression for the ordinary request figures above. ### Files changed by this review - `crates/calternal-fs/src/root.rs`: fixed, private, synced recovery witness. - `crates/calternal-embed/src/store.rs`: restart recovery, scan fairness, current-disk deletion handling, partial-error propagation and regressions. - `crates/calternal-db/src/worker.rs`: fatal-error, writer-release and single outcome/notification regression tests. - `crates/plugins/notes/src/lib.rs`: User lock scope and regression. - `tests/adversarial/bulk_import_1157.py`: ordinary request profile, Note metadata comparison and ordered Search percentile reporting. - `review-1157.md`: findings, decisions and verification evidence. No dependency, migration, route or UI screen was added by this review. All changed doc comments were read again before this report. ### Known gaps No unresolved data-safety finding remains in the completed scenarios. The Search latency reporter fix was syntax checked after the single burst; that burst's invalid Search percentile pair was not reused. Existing ignored model tests were not all run; the new real-model regression was run explicitly. Web check has four warnings in three files. An unchanged InfoPanel test hit its existing five-second timeout on the first full web run. Its focused test and the full confirmation passed without changes to expectations. No UI screen changed, so UX walkthroughs and screenshots do not apply. The first server suite invocation passed units and the performance guard, but cleanup removed its last integration executable before it could run. The missing target passed after rebuild. The full per-crate command was then run again to obtain one complete zero-exit gate. This was a cleanup error, not a changed test expectation or a product failure. ### Gate output (verbatim) `cargo fmt --check` exited 0 with no output. calternal-fs ```text Finished `dev` profile [unoptimized + debuginfo] target(s) in 12m 10s Finished `test` profile [unoptimized + debuginfo] target(s) in 1m 12s test result: ok. 94 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 38.73s test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.10s test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.62s test result: ok. 48 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 9.76s test result: ok. 2 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.17s test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s ``` calternal-db ```text Finished `dev` profile [unoptimized + debuginfo] target(s) in 1m 12s Finished `test` profile [unoptimized + debuginfo] target(s) in 54.14s test result: ok. 38 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 6.67s test result: ok. 7 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 4.52s test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.26s test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.37s test result: ok. 6 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.88s test result: ok. 21 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 4.43s test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.19s test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s ``` DB single-notice regression ```text Finished `dev` profile [unoptimized + debuginfo] target(s) in 3m 57s Finished `test` profile [unoptimized + debuginfo] target(s) in 1m 05s test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 37 filtered out; finished in 1.01s test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 7 filtered out; finished in 0.00s test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 1 filtered out; finished in 0.00s test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 1 filtered out; finished in 0.00s test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 6 filtered out; finished in 0.00s test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 22 filtered out; finished in 0.00s test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 1 filtered out; finished in 0.00s ``` calternal-embed ```text Finished `dev` profile [unoptimized + debuginfo] target(s) in 25.05s Finished `test` profile [unoptimized + debuginfo] target(s) in 7.88s test result: ok. 43 passed; 0 failed; 5 ignored; 0 measured; 0 filtered out; finished in 60.88s test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.01s ``` real-model regression ```text Finished `test` profile [unoptimized + debuginfo] target(s) in 4m 08s test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 46 filtered out; finished in 33.53s ``` calternal-plugin-notes ```text Finished `dev` profile [unoptimized + debuginfo] target(s) in 12m 15s Finished `test` profile [unoptimized + debuginfo] target(s) in 1m 01s test result: ok. 289 passed; 0 failed; 3 ignored; 0 measured; 0 filtered out; finished in 387.81s test result: ok. 2 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 2.37s test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s ``` 700-Note reconcile ```text Finished `test` profile [unoptimized + debuginfo] target(s) in 10m 05s test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 290 filtered out; finished in 116.46s test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 2 filtered out; finished in 0.00s ``` calternal-auth ```text Finished `dev` profile [unoptimized + debuginfo] target(s) in 1m 31s Finished `test` profile [unoptimized + debuginfo] target(s) in 1m 49s test result: ok. 124 passed; 0 failed; 3 ignored; 0 measured; 0 filtered out; finished in 113.77s test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s ``` `cd apps/web && bun run check` ```text perf-lint: PASS; 0 violations; 22104 scoped exceptions svelte-check found 0 errors and 4 warnings in 3 files ``` `cd apps/web && bun run test --maxWorkers=2` ```text Test Files 266 passed (266) Tests 1830 passed (1830) ``` `ADVERSARIAL_BULK_IMPORT_1157_ONLY=1 ADVERSARIAL_SERVER_BIN="$CARGO_TARGET_DIR/debug/calternal-server" ADVERSARIAL_SKIP_WEB_BUILD=1 ADVERSARIAL_KEEP_WORK_DIR=1 timeout --signal=TERM --kill-after=75s 1800s bash tests/adversarial/run.sh` ```text Bulk import #1157 passed: disk, Search paths, Notes Index and semantic Index match. ``` Probe and benchmark JSON, row checks and gate logs are in ignored `artifacts/`. `cargo clippy -p calternal-server --all-targets -- -D warnings` ```text Finished `dev` profile [unoptimized + debuginfo] target(s) in 5m 48s ``` `cargo test -p calternal-server -- --test-threads=4` (complete invocation, exit 0) ```text Finished `test` profile [unoptimized + debuginfo] target(s) in 1m 43s test result: ok. 256 passed; 0 failed; 10 ignored; 0 measured; 0 filtered out; finished in 151.49s test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 36.96s test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.16s ``` Final `cargo fmt --check` exited 0 with no output. Final `cargo clean` exited 0: ```text Removed 8698 files, 7.9GiB total ``` Web build and `.svelte-kit/output` were removed after all gates completed. The disposable live-server data was removed after retaining its profile, row checks and redacted log. The final worktree has only committed changes. ## Verdict **SAFE TO MERGE: yes.** The input branch had clear recovery defects. The review fixed them with separate commits and regression tests. The real-model recovery round, forced-BUSY worker tests, 700-Note reconcile and live mixed burst passed. All six reviewed Rust crates passed clippy and tests. Web check and the full web suite passed. No data loss, duplicate projection, Worker restart or dropped change warning was found in the completed final scenarios. Behavior tested at `c766ce589`; the reporting-only Search percentile fix is `378ea4358`. Later commits only record review evidence. No push or deploy was made. The required merge from `origin/dev` was done once before gates.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#1157
No description provided.