Adversarial Search rebuild probe loses local server during watcher overflow #325

Open
opened 2026-09-28 09:51:24 +00:00 by kayg · 3 comments
Owner

Evidence

On 2026-09-28, tests/adversarial/run.sh started the local server and completed owner and second-user setup. During search_chaos.py, the 20,000-write watcher-overflow fixture reached its 5,000, 10,000 and 15,000 checkpoints. The probe then reported that the old Index stopped returning a live hit during a staged rebuild. Search requests returned HTTP 502 with local adversarial server is unavailable.

The runner waited 1,800 times at 100 ms for restart-ready-4, then timed out. It reported that the Owner passkey refresh failed before a Search admin operation. The server process was no longer present when the probe exited. The runner's cleanup removed its temporary server logs.

Context and impact

This is a non-SLOW availability failure under the staged Search rebuild probe. The shared host was running other large Search and adversarial workloads in parallel. The probe therefore does not show whether the server failure came from the product or host resource contention. Kernel log queries did not reveal an OOM record. The runner stopped at this failure; its later probe groups did not run.

Please reproduce this Search rebuild and watcher-overflow scenario on an otherwise idle server and retain the server log. If the server exits again, identify and fix the cause. The issue was observed after merging dev on job/glass, at head bb174bcedc4b779424a4cb27e3e0bafa809c20db.

## Evidence On 2026-09-28, `tests/adversarial/run.sh` started the local server and completed owner and second-user setup. During `search_chaos.py`, the 20,000-write watcher-overflow fixture reached its 5,000, 10,000 and 15,000 checkpoints. The probe then reported that the old Index stopped returning a live hit during a staged rebuild. Search requests returned HTTP 502 with `local adversarial server is unavailable`. The runner waited 1,800 times at 100 ms for `restart-ready-4`, then timed out. It reported that the Owner passkey refresh failed before a Search admin operation. The server process was no longer present when the probe exited. The runner's cleanup removed its temporary server logs. ## Context and impact This is a non-SLOW availability failure under the staged Search rebuild probe. The shared host was running other large Search and adversarial workloads in parallel. The probe therefore does not show whether the server failure came from the product or host resource contention. Kernel log queries did not reveal an OOM record. The runner stopped at this failure; its later probe groups did not run. Please reproduce this Search rebuild and watcher-overflow scenario on an otherwise idle server and retain the server log. If the server exits again, identify and fix the cause. The issue was observed after merging `dev` on `job/glass`, at head `bb174bcedc4b779424a4cb27e3e0bafa809c20db`.
1.1 MiB
Author
Owner

Additional evidence from the single post-merge adversarial round for #319:

  • Branch: job/single-pills, head 46497b6a7c03eb02a7cbbbfac3f36b563c906af7.
  • tests/adversarial/run.sh reached search_chaos.py against a local server built from this worktree.
  • During the staged rebuild, 16 queries returned HTTP 200 without the live unicodenfcsentinel result. The probe recorded each as “old Index stopped returning a live hit during staged rebuild.”
  • Other cargo and adversarial jobs were active on the shared host, so this run does not establish the cause. The run used the default cleanup and did not retain its temporary server log.

This matches the live-hit inconsistency in this issue. Please reproduce it on an idle server as requested here.

Additional evidence from the single post-merge adversarial round for #319: - Branch: `job/single-pills`, head `46497b6a7c03eb02a7cbbbfac3f36b563c906af7`. - `tests/adversarial/run.sh` reached `search_chaos.py` against a local server built from this worktree. - During the staged rebuild, 16 queries returned HTTP 200 without the live `unicodenfcsentinel` result. The probe recorded each as “old Index stopped returning a live hit during staged rebuild.” - Other cargo and adversarial jobs were active on the shared host, so this run does not establish the cause. The run used the default cleanup and did not retain its temporary server log. This matches the live-hit inconsistency in this issue. Please reproduce it on an idle server as requested here.
Author
Owner

Additional evidence from job/motion-477 at b62297fa0d5e22dd70b1263027d3a6527c5dc476, on 2026-09-30: the local server log records a Tokio worker stack overflow and process abort at 13:53:19 UTC during the broad hostile-input round. The log shows WebDAV upload stages, Tantivy commits and ONNX Runtime allocation shortly before the crash, but this run did not isolate the triggering request. Client probes after the crash reported 502 local adversarial server is unavailable; a later boot appears in the log, but the 502 cascade continued. The captured log is redacted and will be attached to this issue. This may be the same class of failure as #325; I cannot confirm the exact trigger from this run.

Additional evidence from `job/motion-477` at `b62297fa0d5e22dd70b1263027d3a6527c5dc476`, on 2026-09-30: the local server log records a Tokio worker stack overflow and process abort at 13:53:19 UTC during the broad hostile-input round. The log shows WebDAV upload stages, Tantivy commits and ONNX Runtime allocation shortly before the crash, but this run did not isolate the triggering request. Client probes after the crash reported `502 local adversarial server is unavailable`; a later boot appears in the log, but the 502 cascade continued. The captured log is redacted and will be attached to this issue. This may be the same class of failure as #325; I cannot confirm the exact trigger from this run.
Author
Owner

The redacted log captured for the 2026-09-30 run is attached: server.log. It contains the Tokio worker stack-overflow abort at 13:53:19 UTC and the following server boot attempt.

The redacted log captured for the 2026-09-30 run is attached: [server.log](https://git.kayg.org/attachments/50c826cc-1315-4e81-bdbf-3a93628dc9aa). It contains the Tokio worker stack-overflow abort at 13:53:19 UTC and the following server boot attempt.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#325
No description provided.