SECURITY: Preserve cross-User search availability during staged rebuild #345

Closed
opened 2026-09-28 13:41:15 +00:00 by kayg · 11 comments
Owner

Evidence

The one local real-server adversarial run created a 20,000-file watcher-overflow fixture, then rebuilt the shared search index while querying a live indexed Note. tests/adversarial/search_chaos.py::rebuild_during_queries reported old Index stopped returning a live hit during staged rebuild 64 times. Each query returned HTTP 200 without the known result. The run did not expose another User's content; it showed that live search results disappear during a rebuild.

The Indexer and Tantivy Index are Instance-wide, while query roots filter results by User. A normal User's file-event storm can trigger overflow recovery, so the missing-result window can affect other Users who search at the same time. Existing coverage in crates/calternal-search/tests/indexer.rs checks that the old Index remains available while a PDF extractor blocks, but it does not catch the live-server behavior under a large rebuild.

Required work

Add a regression test that keeps querying a second User's existing result while a large staged rebuild and watcher overflow recovery run. Keep the old query generation available until the staged Index is complete and atomically published. Preserve User and Share filtering throughout publication. Link this issue to the cross-User audit in #331.

This is an availability finding. The probe found no cross-User result data.

## Evidence The one local real-server adversarial run created a 20,000-file watcher-overflow fixture, then rebuilt the shared search index while querying a live indexed Note. `tests/adversarial/search_chaos.py::rebuild_during_queries` reported `old Index stopped returning a live hit during staged rebuild` 64 times. Each query returned HTTP 200 without the known result. The run did not expose another User's content; it showed that live search results disappear during a rebuild. The Indexer and Tantivy Index are Instance-wide, while query roots filter results by User. A normal User's file-event storm can trigger overflow recovery, so the missing-result window can affect other Users who search at the same time. Existing coverage in `crates/calternal-search/tests/indexer.rs` checks that the old Index remains available while a PDF extractor blocks, but it does not catch the live-server behavior under a large rebuild. ## Required work Add a regression test that keeps querying a second User's existing result while a large staged rebuild and watcher overflow recovery run. Keep the old query generation available until the staged Index is complete and atomically published. Preserve User and Share filtering throughout publication. Link this issue to the cross-User audit in #331. This is an availability finding. The probe found no cross-User result data.
Author
Owner

Starting #345 on branch job/search-rebuild at ba7104f289 (dev base). I am tracing staged search publication and will add the required live cross-User regression test. Related to #331.

Starting #345 on branch job/search-rebuild at ba7104f289603a58d99b81500f4ffe12e0010666 (dev base). I am tracing staged search publication and will add the required live cross-User regression test. Related to #331.
Author
Owner

Code finding: search_grouped clones the active Arc<QueryIndex> and drops the query-slot read guard before the Tantivy read runs in spawn_blocking. After a successful directory exchange, the retired generation is at the fixed staging path; the next open_staging clears that path. An in-flight Search can therefore outlive the directory entries it reads. I will keep the generation read guard through Tantivy execution and exercise a second User during the overflow/rebuild sequence. This finding is related to #331’s Search isolation audit.

Code finding: `search_grouped` clones the active `Arc<QueryIndex>` and drops the query-slot read guard before the Tantivy read runs in `spawn_blocking`. After a successful directory exchange, the retired generation is at the fixed staging path; the next `open_staging` clears that path. An in-flight Search can therefore outlive the directory entries it reads. I will keep the generation read guard through Tantivy execution and exercise a second User during the overflow/rebuild sequence. This finding is related to #331’s Search isolation audit.
Author
Owner

Confirmed with the focused regression test: after the staged rebuild, Alice’s incremental result is searchable through the live Indexer but is absent when a new Indexer opens .system/index/tantivy. Tantivy reports the published actor Index still bound to .system/index/tantivy-rebuild; a later write updates that retired path. The test fails before the fix with the post-publication update was not written to the active Index. This is the concrete availability cause behind the cross-User impact in #331.

Confirmed with the focused regression test: after the staged rebuild, Alice’s incremental result is searchable through the live Indexer but is absent when a new Indexer opens `.system/index/tantivy`. Tantivy reports the published actor Index still bound to `.system/index/tantivy-rebuild`; a later write updates that retired path. The test fails before the fix with `the post-publication update was not written to the active Index`. This is the concrete availability cause behind the cross-User impact in #331.
Author
Owner

The one real-server adversarial run completed with exit 1. Its only finding was the final integrity check during 20,000-file watcher overflow recovery:

{"checked_items":5736,"missing_items":4673,"stale_items":0,"manifest_mismatches":4674,"repaired_items":4674,"running":true,"healthy":true,"last_error":null}

The second User's committed Search result remained visible, all three overflow markers were found, and the staged rebuild returned HTTP 200. The response showed a follow-up integrity scan still running, so its counters were partial; healthy was true and last_error was null. I am treating this as an early status sample under load and will make the probe wait, with its existing timeout, for the scan to become idle before asserting health. No production behavior change is indicated by this SLOW-only observation.

The one real-server adversarial run completed with exit 1. Its only finding was the final integrity check during 20,000-file watcher overflow recovery: `{"checked_items":5736,"missing_items":4673,"stale_items":0,"manifest_mismatches":4674,"repaired_items":4674,"running":true,"healthy":true,"last_error":null}` The second User's committed Search result remained visible, all three overflow markers were found, and the staged rebuild returned HTTP 200. The response showed a follow-up integrity scan still running, so its counters were partial; `healthy` was true and `last_error` was null. I am treating this as an early status sample under load and will make the probe wait, with its existing timeout, for the scan to become idle before asserting health. No production behavior change is indicated by this SLOW-only observation.
Author
Owner

The 2026-09-28 real-server adversarial round reproduced Search loss during concurrent full rebuild. During a 20,000-write watcher-overflow fixture, concurrent Search returned HTTP 200 without the committed sentinel in samples 4 and 5. The server log then showed the indexer queue full, Tantivy failing to open a term file in the staged rebuild, and the background Search indexer stopping. Subsequent two-user/authz fixture requests got transport failures while the server process remained alive. This is outside the Files write-race change; the detailed observations are also recorded on #343.

The 2026-09-28 real-server adversarial round reproduced Search loss during concurrent full rebuild. During a 20,000-write watcher-overflow fixture, concurrent Search returned HTTP 200 without the committed sentinel in samples 4 and 5. The server log then showed the indexer queue full, Tantivy failing to open a term file in the staged rebuild, and the background Search indexer stopping. Subsequent two-user/authz fixture requests got transport failures while the server process remained alive. This is outside the Files write-race change; the detailed observations are also recorded on #343.
Author
Owner

Resuming #345 after the prior run was stopped by build disk exhaustion (exit 101). Current branch: job/search-rebuild; current head: 43a7aa2835a36e74138344ac485eb3a24cc12896; dev merge base: fcba3cb1092eaae0c5c3162b8b9b0a532b40f43a. The worktree is clean. I am continuing the investigation and regression coverage already committed on this branch.

Resuming #345 after the prior run was stopped by build disk exhaustion (exit 101). Current branch: `job/search-rebuild`; current head: `43a7aa2835a36e74138344ac485eb3a24cc12896`; `dev` merge base: `fcba3cb1092eaae0c5c3162b8b9b0a532b40f43a`. The worktree is clean. I am continuing the investigation and regression coverage already committed on this branch.
Author
Owner

Root cause confirmed by a deterministic regression in calternal-search: the test holds a Search generation read lease, starts a rebuild, then waits until the publication writer is queued. At that point the active-directory marker was already gone, so the directory exchange happened before existing Search readers drained. This matches the live-server Tantivy missing-term-file errors recorded from the overflow run. The test fails before the production fix with: staged publication exchanged the active directory before Search readers drained. I will acquire the query-generation write lock before the directory exchange and hold it through reopen and pointer replacement. The second User's result remains filtered by the existing request context; this is an availability race, not a result leak. Related to #331.

Root cause confirmed by a deterministic regression in `calternal-search`: the test holds a Search generation read lease, starts a rebuild, then waits until the publication writer is queued. At that point the active-directory marker was already gone, so the directory exchange happened before existing Search readers drained. This matches the live-server Tantivy missing-term-file errors recorded from the overflow run. The test fails before the production fix with: `staged publication exchanged the active directory before Search readers drained`. I will acquire the query-generation write lock before the directory exchange and hold it through reopen and pointer replacement. The second User's result remains filtered by the existing request context; this is an availability race, not a result leak. Related to #331.
Author
Owner

Follow-up evidence from the same single adversarial round run for #291: search_chaos.py again reported a successful HTTP 200 query that omitted the committed unicodenfcsentinel result during full/staged rebuild. The local server log showed repeated full indexer-queue warnings and a Tantivy segment merge cancelled because a term file was missing under the staged rebuild directory. This matches the live-result availability finding already tracked here; it does not establish the exact race or cross-User data exposure.

Follow-up evidence from the same single adversarial round run for #291: `search_chaos.py` again reported a successful HTTP 200 query that omitted the committed `unicodenfcsentinel` result during full/staged rebuild. The local server log showed repeated full indexer-queue warnings and a Tantivy segment merge cancelled because a term file was missing under the staged rebuild directory. This matches the live-result availability finding already tracked here; it does not establish the exact race or cross-User data exposure.
Author
Owner

Finished report

Branch: job/search-rebuild
Base: fcba3cb1092eaae0c5c3162b8b9b0a532b40f43a
Merged dev at: fba83527f2cccf2334934bb1fd0932be7c0e209b
HEAD: 768ac3ba9d9a69cf9116112c61776839732e786e
Pushed: yes (git push returned Everything up-to-date).

Built

  • Hold the query-generation write lock before exchanging active and staging directory names. Keep it until the completed Index has reopened and the query pointer has changed. Search remains available during the staged scan; the lock covers only publication.
  • Add a deterministic regression. It held an existing Search read lease and proved that publication does not exchange directory names until that lease drains. Bob's existing result remains searchable before and after publication.
  • Keep the real-server Search chaos probe that queries a second User during the 20,000-file watcher overflow and rebuild. The existing owner-vs-member visibility check still rejects private-result exposure.
  • Related to the cross-User audit in #331.

Files

crates/calternal-search/src/index.rs, crates/calternal-search/src/indexer.rs, crates/calternal-search/tests/indexer.rs, tests/adversarial/search_chaos.py, tests/adversarial/test_search_result_matching.py.

Gates and adversarial result

cargo fmt --check: exited 0 with no output.

cargo clippy --all-targets -- -D warnings final output:

    Finished `dev` profile [unoptimized + debuginfo] target(s) in 82m 38s

cargo test -p calternal-search passed. Exact result lines:

test result: ok. 31 passed; 0 failed; 1 ignored; 0 measured; finished in 1.95s
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; finished in 0.91s
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; finished in 0.10s
test result: ok. 18 passed; 0 failed; 0 ignored; 0 measured; finished in 8.36s
test result: ok. 5 passed; 0 failed; 0 ignored; 0 measured; finished in 0.03s
test result: ok. 2 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.03s
test result: ok. 1 passed; 0 failed; 2 ignored; 0 measured; finished in 3.03s
test result: ok. 4 passed; 0 failed; 0 ignored; 0 measured; finished in 0.01s
test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s

The one real-server Search adversarial round passed:

search chaos: passed (01a0e9c9, 32 concurrent renames, concurrent queries and rebuild)
Ran 3 tests in 0.006s
OK

Full workspace cargo test was interrupted at the four-hour limit with exit code 130 while compiling test binaries. The log had no running or test result: lines, so workspace tests did not start. bun run check and bun run test were not run. These are the remaining gates.

cargo clean output:

     Removed 19620 files, 13.8GiB total

Web build output and local temporary artifacts were deleted.

Decisions not specified in the design

  • Search reads are blocked only for the directory exchange, Index reopen, and query-generation pointer update. The staged scan does not hold the query write lock.
  • Local verification used system OpenSSL 3.5.7 with OPENSSL_NO_VENDOR=1 after the vendored OpenSSL build exceeded the timebox. This changed no source or dependency files.
## Finished report Branch: `job/search-rebuild` Base: `fcba3cb1092eaae0c5c3162b8b9b0a532b40f43a` Merged `dev` at: `fba83527f2cccf2334934bb1fd0932be7c0e209b` HEAD: `768ac3ba9d9a69cf9116112c61776839732e786e` Pushed: yes (`git push` returned `Everything up-to-date`). ### Built - Hold the query-generation write lock before exchanging active and staging directory names. Keep it until the completed Index has reopened and the query pointer has changed. Search remains available during the staged scan; the lock covers only publication. - Add a deterministic regression. It held an existing Search read lease and proved that publication does not exchange directory names until that lease drains. Bob's existing result remains searchable before and after publication. - Keep the real-server Search chaos probe that queries a second User during the 20,000-file watcher overflow and rebuild. The existing owner-vs-member visibility check still rejects private-result exposure. - Related to the cross-User audit in #331. ### Files `crates/calternal-search/src/index.rs`, `crates/calternal-search/src/indexer.rs`, `crates/calternal-search/tests/indexer.rs`, `tests/adversarial/search_chaos.py`, `tests/adversarial/test_search_result_matching.py`. ### Gates and adversarial result `cargo fmt --check`: exited 0 with no output. `cargo clippy --all-targets -- -D warnings` final output: ```text Finished `dev` profile [unoptimized + debuginfo] target(s) in 82m 38s ``` `cargo test -p calternal-search` passed. Exact result lines: ```text test result: ok. 31 passed; 0 failed; 1 ignored; 0 measured; finished in 1.95s test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; finished in 0.91s test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; finished in 0.10s test result: ok. 18 passed; 0 failed; 0 ignored; 0 measured; finished in 8.36s test result: ok. 5 passed; 0 failed; 0 ignored; 0 measured; finished in 0.03s test result: ok. 2 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.03s test result: ok. 1 passed; 0 failed; 2 ignored; 0 measured; finished in 3.03s test result: ok. 4 passed; 0 failed; 0 ignored; 0 measured; finished in 0.01s test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s ``` The one real-server Search adversarial round passed: ```text search chaos: passed (01a0e9c9, 32 concurrent renames, concurrent queries and rebuild) Ran 3 tests in 0.006s OK ``` Full workspace `cargo test` was interrupted at the four-hour limit with exit code 130 while compiling test binaries. The log had no `running` or `test result:` lines, so workspace tests did not start. `bun run check` and `bun run test` were not run. These are the remaining gates. `cargo clean` output: ```text Removed 19620 files, 13.8GiB total ``` Web build output and local temporary artifacts were deleted. ### Decisions not specified in the design - Search reads are blocked only for the directory exchange, Index reopen, and query-generation pointer update. The staged scan does not hold the query write lock. - Local verification used system OpenSSL 3.5.7 with `OPENSSL_NO_VENDOR=1` after the vendored OpenSSL build exceeded the timebox. This changed no source or dependency files.
Author
Owner

Fixed by job/search-rebuild (publication drains readers before the directory exchange), merged into dev by Claude. Closing.

Fixed by job/search-rebuild (publication drains readers before the directory exchange), merged into dev by Claude. Closing.
kayg closed this issue 2026-09-28 22:48:28 +00:00
Author
Owner

Merged into dev by Claude after review (413ccaa7), deploying to calternal.cloud. Closing.

Merged into dev by Claude after review (413ccaa7), deploying to calternal.cloud. Closing.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#345
No description provided.