Server graceful shutdown waits without a bound for a running Search rebuild or repair #963

Open
opened 2026-10-03 00:32:02 +00:00 by kayg · 6 comments
Owner

Found by the merge-round-7a adversarial run (#427, round 5, head after the #959 runner fix). Related to #959 but not a test-infrastructure problem.

Problem

calternal-server does not finish a graceful shutdown in bounded time while the Search actor runs a long staged rebuild or integrity repair. SIGTERM is received, HTTP drains (bounded by serve::DRAIN_LIMIT, 3 s) and live Note rooms flush (bounded by session::SHUTDOWN_FLUSH_WAIT, 10 s). Then main awaits live.indexer.shutdown() without a deadline (crates/calternal-server/src/main.rs, after serve::serve).

Indexer::shutdown (crates/calternal-search/src/indexer.rs) only sets the watch flag and waits for stopped. The actor loop checks that flag only between commands in its tokio::select!. A running rebuild_all, check_and_repair_inner or overflow recovery does not check it, so the process stays alive until that whole operation ends.

Evidence

The full runner now stops each child with a 60 s grace and prints a RUNNER line. Verbatim, after Search chaos (24,000 fixture files in one Home):

02:04:18 RUNNER calternal-server pid 959591 still running 60s after SIGTERM; sending SIGKILL
02:04:19 RUNNER restart_server: graceful shutdown overran; server was killed

Server log for that process: shutdown signal received at 00:03:15.86Z, then 47 more Tantivy commits (Prepared commit 2050 up to Prepared commit 48276 at 00:04:17.11Z) and no Search Index writer released after shutdown line before SIGKILL. In merge-round-7a round 4 the same wait held the server for more than ten minutes (#959).

Impact

  • Not data loss: SIGKILL leaves the startup integrity repair to heal the Index, and the run healed. Not a merge blocker under the adversarial rule.
  • A container or systemd stop always reaches its kill timeout during a rebuild, so every deploy or restart during maintenance becomes a crash restart, followed by a full startup repair (which then also contends for the SQLite pool, see #960).

Expected

Shutdown completes within a stated bound (for example the container stop timeout minus the drain and flush budgets). Either the long Search operations check the shutdown flag between batches and stop at a committed, resumable point, or main bounds indexer.shutdown() with a timeout and logs that the next start repairs the Index.

Regression test

A server test that starts a rebuild over a large Home, sends the shutdown, and asserts that Indexer::shutdown returns within the bound, and that the next start opens the Index and reports a healthy integrity check.

Reproduce

tests/adversarial/run.sh on the merge-round-7a branch; the RUNNER line above appears at the restart after search_chaos.py.

Found by the merge-round-7a adversarial run (#427, round 5, head after the #959 runner fix). Related to #959 but not a test-infrastructure problem. ## Problem `calternal-server` does not finish a graceful shutdown in bounded time while the Search actor runs a long staged rebuild or integrity repair. SIGTERM is received, HTTP drains (bounded by `serve::DRAIN_LIMIT`, 3 s) and live Note rooms flush (bounded by `session::SHUTDOWN_FLUSH_WAIT`, 10 s). Then `main` awaits `live.indexer.shutdown()` without a deadline (`crates/calternal-server/src/main.rs`, after `serve::serve`). `Indexer::shutdown` (`crates/calternal-search/src/indexer.rs`) only sets the watch flag and waits for `stopped`. The actor loop checks that flag only between commands in its `tokio::select!`. A running `rebuild_all`, `check_and_repair_inner` or overflow recovery does not check it, so the process stays alive until that whole operation ends. ## Evidence The full runner now stops each child with a 60 s grace and prints a RUNNER line. Verbatim, after Search chaos (24,000 fixture files in one Home): ``` 02:04:18 RUNNER calternal-server pid 959591 still running 60s after SIGTERM; sending SIGKILL 02:04:19 RUNNER restart_server: graceful shutdown overran; server was killed ``` Server log for that process: `shutdown signal received` at 00:03:15.86Z, then 47 more Tantivy commits (`Prepared commit 2050` up to `Prepared commit 48276` at 00:04:17.11Z) and no `Search Index writer released after shutdown` line before SIGKILL. In merge-round-7a round 4 the same wait held the server for more than ten minutes (#959). ## Impact - Not data loss: SIGKILL leaves the startup integrity repair to heal the Index, and the run healed. Not a merge blocker under the adversarial rule. - A container or systemd stop always reaches its kill timeout during a rebuild, so every deploy or restart during maintenance becomes a crash restart, followed by a full startup repair (which then also contends for the SQLite pool, see #960). ## Expected Shutdown completes within a stated bound (for example the container stop timeout minus the drain and flush budgets). Either the long Search operations check the shutdown flag between batches and stop at a committed, resumable point, or `main` bounds `indexer.shutdown()` with a timeout and logs that the next start repairs the Index. ## Regression test A server test that starts a rebuild over a large Home, sends the shutdown, and asserts that `Indexer::shutdown` returns within the bound, and that the next start opens the Index and reports a healthy integrity check. ## Reproduce `tests/adversarial/run.sh` on the merge-round-7a branch; the RUNNER line above appears at the restart after `search_chaos.py`.
Author
Owner

Merge-round 7a full runner evidence: during settle_server authz-matrix, the release server did not exit within the runner's 60-second SIGTERM grace and the harness sent SIGKILL (RUNNER calternal-server pid ... still running 60s after SIGTERM). This followed the dedup scrub/GC probe, whose status had remained in scanning; the server restarted afterward for the authorization matrix. The run was under high host load. Full output is in #427's worktree log.

Merge-round 7a full runner evidence: during `settle_server authz-matrix`, the release server did not exit within the runner's 60-second SIGTERM grace and the harness sent SIGKILL (`RUNNER calternal-server pid ... still running 60s after SIGTERM`). This followed the dedup scrub/GC probe, whose status had remained in `scanning`; the server restarted afterward for the authorization matrix. The run was under high host load. Full output is in #427's worktree log.
Author
Owner

Starting #963 on branch job/shutdown-963, based on origin/dev at d4e7188810a89fb0e8e6b162279917f7e23989f9. I am tracing the actor shutdown and rebuild/repair boundaries before selecting the smallest resumable fix.

Starting #963 on branch `job/shutdown-963`, based on `origin/dev` at `d4e7188810a89fb0e8e6b162279917f7e23989f9`. I am tracing the actor shutdown and rebuild/repair boundaries before selecting the smallest resumable fix.
Author
Owner

Evidence from job/shutdown-963 at d4e7188: Indexer::run selects the shutdown watch only while idle; scan_tree has no shutdown check, and reconcile_and_report counts then scans the full Home. main.rs in this base also has no Indexer stop hook, and LiveApp does not retain one. I will add safe scan-boundary cancellation plus a bounded hook after the Note room flush; interrupted staged rebuilds retain the active Index, and startup reconciliation repairs it.

Evidence from `job/shutdown-963` at `d4e7188`: `Indexer::run` selects the shutdown watch only while idle; `scan_tree` has no shutdown check, and `reconcile_and_report` counts then scans the full Home. `main.rs` in this base also has no Indexer stop hook, and `LiveApp` does not retain one. I will add safe scan-boundary cancellation plus a bounded hook after the Note room flush; interrupted staged rebuilds retain the active Index, and startup reconciliation repairs it.
Author
Owner

Rechecked origin/dev d4e7188810. main flushes live Note rooms but has no Search Indexer stop hook; the Indexer actor runs full repair or rebuild work before it can select shutdown. No bounded stop path or matching regression test is present. The work note remains on job/shutdown-963; keep open until merged.

Rechecked origin/dev d4e7188810a89fb0e8e6b162279917f7e23989f9. main flushes live Note rooms but has no Search Indexer stop hook; the Indexer actor runs full repair or rebuild work before it can select shutdown. No bounded stop path or matching regression test is present. The work note remains on job/shutdown-963; keep open until merged.
Author
Owner

Server clippy found an ownership integration error at wire.rs: build_app_with_files(..., Some(indexer)) consumed the handle before LiveApp could retain it for shutdown. I changed that call to pass a clone; both fields now share the same Indexer actor.

Server clippy found an ownership integration error at `wire.rs`: `build_app_with_files(..., Some(indexer))` consumed the handle before `LiveApp` could retain it for shutdown. I changed that call to pass a clone; both fields now share the same Indexer actor.
Author
Owner

Finished #963.

Built: Search shutdown now sets a shared cancellation flag and stops rebuild, reconcile, and repair work at safe item or batch boundaries. If a staged rebuild is interrupted before manifest persistence, it restores the previous active Index. Once persistence starts, publication completes as one commit point. Indexer::shutdown() has a 10-second bound, and the server calls it after HTTP drain and Note-room flush. A later startup integrity check repairs interrupted work.

Files

  • crates/calternal-search/src/indexer.rs
  • crates/calternal-search/src/index.rs
  • crates/calternal-search/tests/indexer.rs
  • crates/calternal-server/src/main.rs
  • crates/calternal-server/src/wire.rs

Commit: a048bad23c57f08724890cb3822de7dd0255dc32

Gate output (verbatim excerpts)

cargo fmt --check
[no output; exit status 0]

cargo clippy -p calternal-search --all-targets -- -D warnings
    Finished `dev` profile [unoptimized + debuginfo] target(s) in 2.70s

cargo test -p calternal-search
test result: ok. 36 passed; 0 failed; 1 ignored
test result: ok. 1 passed
test result: ok. 1 passed
test result: ok. 22 passed; 0 failed
test result: ok. 5 passed
test result: ok. 2 passed
test result: ok. 1 passed; 2 ignored
test result: ok. 4 passed
test result: ok. 0 passed

cargo clippy -p calternal-server --all-targets -- -D warnings
    Finished `dev` profile [unoptimized + debuginfo] target(s) in 1m 55s

cargo test -p calternal-server
test result: ok. 107 passed; 0 failed; 3 ignored; 0 measured; 0 filtered out; finished in 31.24s

Known gap: The regression test creates 256 Notes and uses a gated PDF extractor to hold the rebuild at a deterministic cancellation point, then verifies a healthy integrity report after restart. It does not recreate the adversarial runner's 24,000-file Home; creating that fixture through Root.write took over six minutes for about 1,200 files in this worktree.

Decision not specified by DESIGN: The Search actor gets a 10-second shutdown deadline, leaving time inside the reported 60-second stop grace after the 3-second HTTP drain and 10-second Note flush. The regression test uses a deterministic extractor gate so it can exercise shutdown reliably without building the 24,000-file fixture.

For the merge round: Run tests/adversarial/run.sh and verify that Search chaos restart exits within the runner's 60-second SIGTERM grace and the next start reports a healthy Index.

Finished #963. **Built:** Search shutdown now sets a shared cancellation flag and stops rebuild, reconcile, and repair work at safe item or batch boundaries. If a staged rebuild is interrupted before manifest persistence, it restores the previous active Index. Once persistence starts, publication completes as one commit point. `Indexer::shutdown()` has a 10-second bound, and the server calls it after HTTP drain and Note-room flush. A later startup integrity check repairs interrupted work. **Files** - `crates/calternal-search/src/indexer.rs` - `crates/calternal-search/src/index.rs` - `crates/calternal-search/tests/indexer.rs` - `crates/calternal-server/src/main.rs` - `crates/calternal-server/src/wire.rs` **Commit:** `a048bad23c57f08724890cb3822de7dd0255dc32` **Gate output (verbatim excerpts)** ``` cargo fmt --check [no output; exit status 0] cargo clippy -p calternal-search --all-targets -- -D warnings Finished `dev` profile [unoptimized + debuginfo] target(s) in 2.70s cargo test -p calternal-search test result: ok. 36 passed; 0 failed; 1 ignored test result: ok. 1 passed test result: ok. 1 passed test result: ok. 22 passed; 0 failed test result: ok. 5 passed test result: ok. 2 passed test result: ok. 1 passed; 2 ignored test result: ok. 4 passed test result: ok. 0 passed cargo clippy -p calternal-server --all-targets -- -D warnings Finished `dev` profile [unoptimized + debuginfo] target(s) in 1m 55s cargo test -p calternal-server test result: ok. 107 passed; 0 failed; 3 ignored; 0 measured; 0 filtered out; finished in 31.24s ``` **Known gap:** The regression test creates 256 Notes and uses a gated PDF extractor to hold the rebuild at a deterministic cancellation point, then verifies a healthy integrity report after restart. It does not recreate the adversarial runner's 24,000-file Home; creating that fixture through `Root.write` took over six minutes for about 1,200 files in this worktree. **Decision not specified by DESIGN:** The Search actor gets a 10-second shutdown deadline, leaving time inside the reported 60-second stop grace after the 3-second HTTP drain and 10-second Note flush. The regression test uses a deterministic extractor gate so it can exercise shutdown reliably without building the 24,000-file fixture. **For the merge round:** Run `tests/adversarial/run.sh` and verify that Search chaos restart exits within the runner's 60-second SIGTERM grace and the next start reports a healthy Index.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#963
No description provided.