Adversarial API requests time out during concurrent Notes and sync probes #166

Open
opened 2026-09-26 10:50:17 +00:00 by kayg · 8 comments
Owner

Finding

The real-server adversarial run on job/agenda (head ae48cd11c78686e9c01ce2649a1791f529ef3dde) found repeated HTTP request timeouts and missing results during its concurrent API probes. This is separate from the Agenda UI work.

In tests/adversarial/attack.py, the runner reported:

  • all 5 template-create storm requests timed out at the 30-second client deadline; the expected 16 unique note IDs reached only 11
  • 6 of 16 bookmark capture storm requests timed out; 10 unique IDs were returned
  • the stale Journal line fix timed out instead of returning its expected 412
  • Reminders VTODO discovery and incremental sync timed out

In tests/adversarial/attack2.py, it reported:

  • the sync upload probe did not see its 11 initial files in the remote folder within 25 seconds
  • a Journal seed write timed out; the log-rewrite checks then saw 4 of 5 expected entries
  • the Journal race storm had timed-out create and move requests
  • one Unicode VTODO read and two Unicode task title requests timed out

The server process stayed alive, and /readyz returned 200 in 0.651 seconds during the first attack. The same run also logged many responses as SLOW. At the time, the shared host had several concurrent builds and probes; one-minute load averages reached about 85. The timeout and consistency observations need a controlled-load replay before assigning a product cause. The runner did not report a crash, 5xx, or accepted hostile input in these cases.

Requested follow-up

Replay the affected checks with bounded, controlled concurrency and record whether each request completes with an HTTP status or times out. If reproduced, bound or queue the expensive work so overload returns a defined response and journal/sync writes keep their expected identities and contents.

Related: #98, #149, #154.

## Finding The real-server adversarial run on `job/agenda` (head `ae48cd11c78686e9c01ce2649a1791f529ef3dde`) found repeated HTTP request timeouts and missing results during its concurrent API probes. This is separate from the Agenda UI work. In `tests/adversarial/attack.py`, the runner reported: - all 5 template-create storm requests timed out at the 30-second client deadline; the expected 16 unique note IDs reached only 11 - 6 of 16 bookmark capture storm requests timed out; 10 unique IDs were returned - the stale Journal line fix timed out instead of returning its expected 412 - Reminders VTODO discovery and incremental sync timed out In `tests/adversarial/attack2.py`, it reported: - the sync upload probe did not see its 11 initial files in the remote folder within 25 seconds - a Journal seed write timed out; the log-rewrite checks then saw 4 of 5 expected entries - the Journal race storm had timed-out create and move requests - one Unicode VTODO read and two Unicode task title requests timed out The server process stayed alive, and `/readyz` returned 200 in 0.651 seconds during the first attack. The same run also logged many responses as `SLOW`. At the time, the shared host had several concurrent builds and probes; one-minute load averages reached about 85. The timeout and consistency observations need a controlled-load replay before assigning a product cause. The runner did not report a crash, 5xx, or accepted hostile input in these cases. ## Requested follow-up Replay the affected checks with bounded, controlled concurrency and record whether each request completes with an HTTP status or times out. If reproduced, bound or queue the expensive work so overload returns a defined response and journal/sync writes keep their expected identities and contents. Related: #98, #149, #154.
Author
Owner

Additional evidence from the ask-page branch's one post-merge adversarial run on 2026-09-26 (head 32f9d1a9):

  • tests/adversarial/run.sh ran attack2.py against a real local server.
  • The sync upload probe did not find these initial files in the remote folder: a.txt, keep/k0.txt, keep/k1.txt, keep/k2.txt, and keep/k3.txt.
  • The sync log reported the mass-deletion guard paused a pair. The server process remained alive.
  • Several independent builds and adversarial probes were active on the shared host, so this is not a controlled-load replay. It adds evidence to the timeout/missing-result follow-up here.
Additional evidence from the `ask-page` branch's one post-merge adversarial run on 2026-09-26 (head `32f9d1a9`): - `tests/adversarial/run.sh` ran `attack2.py` against a real local server. - The sync upload probe did not find these initial files in the remote folder: `a.txt`, `keep/k0.txt`, `keep/k1.txt`, `keep/k2.txt`, and `keep/k3.txt`. - The sync log reported the mass-deletion guard paused a pair. The server process remained alive. - Several independent builds and adversarial probes were active on the shared host, so this is not a controlled-load replay. It adds evidence to the timeout/missing-result follow-up here.
Author
Owner

Additional evidence from the full post-merge real-server adversarial pass on 2026-09-27:

  • The change-feed request during the load storm returned HTTP 503 with service_unavailable and Authentication database is busy; retry shortly.
  • Saved-search rename and create storms also returned this 503 response. The create probe expected 40 distinct saved searches and observed 20; the remaining create requests were reported as 503.
  • The server stayed alive. These requests received a defined overload response rather than timing out or returning HTTP 500. This shared-host run does not establish normal capacity, but the client-visible concurrency failure remains reproducible under the combined probe load.
Additional evidence from the full post-merge real-server adversarial pass on 2026-09-27: - The change-feed request during the load storm returned HTTP 503 with `service_unavailable` and `Authentication database is busy; retry shortly`. - Saved-search rename and create storms also returned this 503 response. The create probe expected 40 distinct saved searches and observed 20; the remaining create requests were reported as 503. - The server stayed alive. These requests received a defined overload response rather than timing out or returning HTTP 500. This shared-host run does not establish normal capacity, but the client-visible concurrency failure remains reproducible under the combined probe load.
Author
Owner

The latest post-merge real-server adversarial round returned retryable Authentication database is busy; retry shortly 503 responses during several concurrent write storms: one task create, three requests in the Calendar duplicate-account storm, multiple saved-search renames, four Analytics writes, and one Journal sync PATCH. The server stayed alive. The Journal race consistency check found 46 entries for 40 successful creates, 6 seeds, and 8 uploads, with no lost or duplicate lines.

The host had multiple concurrent Cargo builds and adversarial runs, so this does not establish a quiet-host defect. These are status failures rather than SLOW-only results; this run records them for follow-up alongside the existing Index-busy responses on #148.

The latest post-merge real-server adversarial round returned retryable `Authentication database is busy; retry shortly` 503 responses during several concurrent write storms: one task create, three requests in the Calendar duplicate-account storm, multiple saved-search renames, four Analytics writes, and one Journal sync PATCH. The server stayed alive. The Journal race consistency check found 46 entries for 40 successful creates, 6 seeds, and 8 uploads, with no lost or duplicate lines. The host had multiple concurrent Cargo builds and adversarial runs, so this does not establish a quiet-host defect. These are status failures rather than SLOW-only results; this run records them for follow-up alongside the existing Index-busy responses on #148.
Author
Owner

Post-merge adversarial evidence from CSP issue #118 (job/csp merge commit c2ff7b40, seed 25608414): the editor's real WebSocket collaboration probe for a Note with 10,000 top-level blocks timed out during sync at its 20-second deadline. The next one-megabyte paragraph probe passed in 4.281 seconds, and the server remained alive. This is a non-SLOW timeout; the run continues to check for other findings.

Post-merge adversarial evidence from CSP issue #118 (`job/csp` merge commit `c2ff7b40`, seed 25608414): the editor's real WebSocket collaboration probe for a Note with 10,000 top-level blocks timed out during sync at its 20-second deadline. The next one-megabyte paragraph probe passed in 4.281 seconds, and the server remained alive. This is a non-SLOW timeout; the run continues to check for other findings.
Author
Owner

The current one-round adversarial run also reported analytics 70 KB query: NO RESPONSE (b'[Errno 104] Connection reset by peer'). This is the probe's oversized request-line case (GET /api/v1/analytics?..., 70 KB tz value); it is sent through the Node proxy and the server process stayed alive. The reset may be the proxy's request-line size refusal, so this observation does not establish a Rust crash or handler acceptance. It needs the same controlled replay/response classification noted here. Full log: target/tmp/phone-chrome-adversarial-retry.log in the phone-chrome worktree.

The current one-round adversarial run also reported `analytics 70 KB query: NO RESPONSE (b'[Errno 104] Connection reset by peer')`. This is the probe's oversized request-line case (`GET /api/v1/analytics?...`, 70 KB `tz` value); it is sent through the Node proxy and the server process stayed alive. The reset may be the proxy's request-line size refusal, so this observation does not establish a Rust crash or handler acceptance. It needs the same controlled replay/response classification noted here. Full log: `target/tmp/phone-chrome-adversarial-retry.log` in the phone-chrome worktree.
Author
Owner

Post-merge adversarial run at c2ff7b40: one 70 KiB Analytics request ended with a peer reset. During the Analytics storm, requests that did receive 201 took 20.5–29.9 seconds and ten requests timed out at the 30-second client deadline. The server was alive at the end. Several worktrees were running probes/builds on the shared host, so this needs a controlled-load replay; I have not attributed the timeouts to a product defect.

Post-merge adversarial run at c2ff7b40: one 70 KiB Analytics request ended with a peer reset. During the Analytics storm, requests that did receive 201 took 20.5–29.9 seconds and ten requests timed out at the 30-second client deadline. The server was alive at the end. Several worktrees were running probes/builds on the shared host, so this needs a controlled-load replay; I have not attributed the timeouts to a product defect.
Author
Owner

Additional gate evidence from the #148 robustness branch (7655d18ad79b4635ba5c754859b69b9c68a3598c): the full workspace retry completed most test targets, but six targets ended nonzero. Several failures were Sqlx(PoolTimedOut) while test fixtures called Db::connect (Collab shared_notes/two_clients/untouched_bytes, DB queue, and two Notes tests). Photos also had one such fixture timeout on the first workspace pass; an isolated serial Photos run passed 39 tests with 2 ignored, and the retry passed Photos. Server (43 passed, 2 ignored) and Sync (47 library + 2 CLI tests passed) also passed in isolation/serial runs.

This points to connection-pool pressure in the shared test host; these are fixture setup errors, not a confirmed API regression. The same host load caused latency and transient retryable 503 responses during the adversarial API round. I did not change those unrelated crates in this job.

Additional gate evidence from the #148 robustness branch (`7655d18ad79b4635ba5c754859b69b9c68a3598c`): the full workspace retry completed most test targets, but six targets ended nonzero. Several failures were `Sqlx(PoolTimedOut)` while test fixtures called `Db::connect` (Collab shared_notes/two_clients/untouched_bytes, DB queue, and two Notes tests). Photos also had one such fixture timeout on the first workspace pass; an isolated serial Photos run passed 39 tests with 2 ignored, and the retry passed Photos. Server (43 passed, 2 ignored) and Sync (47 library + 2 CLI tests passed) also passed in isolation/serial runs. This points to connection-pool pressure in the shared test host; these are fixture setup errors, not a confirmed API regression. The same host load caused latency and transient retryable `503` responses during the adversarial API round. I did not change those unrelated crates in this job.
Author
Owner

The first fonts-branch adversarial run saw GET for saved-search ID .. return 502 local adversarial server is unavailable; subsequent saved-search creates completed, and the probe reported the server alive at the end. The post-merge full run completed the saved-search probe without a 502 and kept the server alive. Please keep the first response in the controlled-load replay evidence.

The first fonts-branch adversarial run saw `GET` for saved-search ID `..` return `502 local adversarial server is unavailable`; subsequent saved-search creates completed, and the probe reported the server alive at the end. The post-merge full run completed the saved-search probe without a 502 and kept the server alive. Please keep the first response in the controlled-load replay evidence.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#166
No description provided.