PERF audit: diagnose socket hang-up during read-only Admin burst (#663) #705

Open
opened 2026-10-02 05:40:13 +00:00 by kayg · 3 comments
Owner

Context: #663 production/HDD audit. This issue owns diagnosis of a non-SLOW transport failure, not a claimed server crash or authorization bug. #664 was checked; its API fixture/tag/DAV findings are different.

Evidence:

  • Perf VM root@10.69.69.63, shared release server source cc25c441b7; binary SHA-256 2f3567d91c34839851247bc0acbc25a56aaacd14dca269b8f0342ddf83447ed9. Embedded production SPA, browser/request client on build host through SSH and the shared HTTPS front.
  • Every phase held flock -w 14400 /root/perf.lock. bench/hdd-emu.sh: direct-I/O loop/ext4, 8 ms read/write dm-delay, 200 IOPS and 150 MiB/s caps. Qualification QD1 125.008329 IOPS, p50 8.028160 ms, p99 8.159232 ms; QD16 200.892560 IOPS, p50 100.139008 ms, p99 104.333312 ms. Load at qualification 0.74/1.66/1.40 and 0.73/1.60/1.38.
  • Synthetic fixture: 366 Daily notes, 10,980 Logs, 100 Files/Photos, 20 Notes/Tasks, three Budgets and 100 transactions; one User, empty Mail. No production User data.
  • After eight endpoint profiles completed (each five serial and five concurrent GETs), the five-request burst of GET /api/v1/auth/users ended with get: socket hang up. Call-log credentials were redacted. The script did not save an Admin burst latency and did not reach its later Note-detail probe. No successful sample was invented for the interrupted burst.
  • An earlier independent phase did complete the same endpoint: five serial HTTP 200 reads, median/p95/max 140.4/929.4/929.4 ms; five-request burst HTTP 200, median/p95/max 81.6/82.2/82.2 ms. This shows variability, not a reliably reproduced server defect.
  • Post-cleanup inspection found no kernel OOM/killed-process entry for the phase. Cleanup stops the temporary server, so a missing process after cleanup is not crash evidence. Health and the process exit status were not captured at the instant of the hang-up.
  • Source boundaries to inspect: bench/tab-switch.mjs:306 temporary server lifecycle, apps/web/e2e/harness.mjs HTTPS front/tunnel helpers, crates/calternal-auth/src/api.rs:1384 list_users and store.rs:1611 query.

Expected:
Keep successful read bursts connected or return a complete bounded API error. Determine whether the failure was in the server, SSH tunnel, HTTPS front or request client before changing product behavior. Instrument content-free server exit/health and transport state at failure, while still holding the perf lock. Never log cookies, credentials, raw Admin configuration or User content.

Tests:
One bounded production/HDD retest with the same fixture and lock, five serial reads plus one five-request burst. Record every status/error, health at failure, server exit and SSH/proxy exit. If a product defect is established, add a regression at that boundary and preserve existing expectations. Do not loop until the host is quiet or discard failed samples. Reuse #549 harness and #697 Auth paging/read-priority follow-ups; do not duplicate #664.

Context: #663 production/HDD audit. This issue owns diagnosis of a non-SLOW transport failure, not a claimed server crash or authorization bug. #664 was checked; its API fixture/tag/DAV findings are different. Evidence: - Perf VM root@10.69.69.63, shared release server source cc25c441b7a974185622a1dee853cf38686d2b67; binary SHA-256 2f3567d91c34839851247bc0acbc25a56aaacd14dca269b8f0342ddf83447ed9. Embedded production SPA, browser/request client on build host through SSH and the shared HTTPS front. - Every phase held flock -w 14400 /root/perf.lock. bench/hdd-emu.sh: direct-I/O loop/ext4, 8 ms read/write dm-delay, 200 IOPS and 150 MiB/s caps. Qualification QD1 125.008329 IOPS, p50 8.028160 ms, p99 8.159232 ms; QD16 200.892560 IOPS, p50 100.139008 ms, p99 104.333312 ms. Load at qualification 0.74/1.66/1.40 and 0.73/1.60/1.38. - Synthetic fixture: 366 Daily notes, 10,980 Logs, 100 Files/Photos, 20 Notes/Tasks, three Budgets and 100 transactions; one User, empty Mail. No production User data. - After eight endpoint profiles completed (each five serial and five concurrent GETs), the five-request burst of GET /api/v1/auth/users ended with `get: socket hang up`. Call-log credentials were redacted. The script did not save an Admin burst latency and did not reach its later Note-detail probe. No successful sample was invented for the interrupted burst. - An earlier independent phase did complete the same endpoint: five serial HTTP 200 reads, median/p95/max 140.4/929.4/929.4 ms; five-request burst HTTP 200, median/p95/max 81.6/82.2/82.2 ms. This shows variability, not a reliably reproduced server defect. - Post-cleanup inspection found no kernel OOM/killed-process entry for the phase. Cleanup stops the temporary server, so a missing process after cleanup is not crash evidence. Health and the process exit status were not captured at the instant of the hang-up. - Source boundaries to inspect: bench/tab-switch.mjs:306 temporary server lifecycle, apps/web/e2e/harness.mjs HTTPS front/tunnel helpers, crates/calternal-auth/src/api.rs:1384 list_users and store.rs:1611 query. Expected: Keep successful read bursts connected or return a complete bounded API error. Determine whether the failure was in the server, SSH tunnel, HTTPS front or request client before changing product behavior. Instrument content-free server exit/health and transport state at failure, while still holding the perf lock. Never log cookies, credentials, raw Admin configuration or User content. Tests: One bounded production/HDD retest with the same fixture and lock, five serial reads plus one five-request burst. Record every status/error, health at failure, server exit and SSH/proxy exit. If a product defect is established, add a regression at that boundary and preserve existing expectations. Do not loop until the host is quiet or discard failed samples. Reuse #549 harness and #697 Auth paging/read-priority follow-ups; do not duplicate #664.
Author
Owner

Started #705 on job/admin-burst-705, based on origin/dev at c4a61e8cf090170f35b1bed3350d9de20c83ecd5. I have read the issue and parent #663 matrix. origin/dev does not include DESIGN §58, so I will cherry-pick its two documentation commits from job/instant-663 before the bounded transport retest. Scope remains diagnosis of the Admin burst socket hang-up; I will only change product behavior if evidence establishes a server defect.

Started #705 on `job/admin-burst-705`, based on `origin/dev` at `c4a61e8cf090170f35b1bed3350d9de20c83ecd5`. I have read the issue and parent #663 matrix. `origin/dev` does not include DESIGN §58, so I will cherry-pick its two documentation commits from `job/instant-663` before the bounded transport retest. Scope remains diagnosis of the Admin burst socket hang-up; I will only change product behavior if evidence establishes a server defect.
Author
Owner

#705 bounded retest completed on the pinned production server (cc25c441b7a974185622a1dee853cf38686d2b67, SHA-256 2f3567d91c34839851247bc0acbc25a56aaacd14dca269b8f0342ddf83447ed9) with the #663 fixture and HDD profile. All five serial and five concurrent GET /api/v1/auth/users requests returned HTTP 200 (198 bytes, one User). Serial median/p95/max: 208.4/1165.2/1165.2 ms; burst: 98.8/99.5/99.5 ms. No socket hang-up was reproduced.

At the post-burst snapshot, the remote server PID was alive in state S; its SSH process and the API tunnel process had no exit code or signal. The HTTPS front was listening, reported zero upstream errors, and gained no early closes during the Admin burst. /healthz and /readyz both returned HTTP 200 through both the HTTPS front and direct SSH tunnel. Load at the snapshot, inside the lock, was 3.33/2.18/1.02. These observations do not identify the earlier interruption's transport hop, so I made no server/API behavior change. The shared E2E front now exposes bounded, content-free transport counters for future diagnosis.

#705 bounded retest completed on the pinned production server (`cc25c441b7a974185622a1dee853cf38686d2b67`, SHA-256 `2f3567d91c34839851247bc0acbc25a56aaacd14dca269b8f0342ddf83447ed9`) with the #663 fixture and HDD profile. All five serial and five concurrent `GET /api/v1/auth/users` requests returned HTTP 200 (198 bytes, one User). Serial median/p95/max: 208.4/1165.2/1165.2 ms; burst: 98.8/99.5/99.5 ms. No socket hang-up was reproduced. At the post-burst snapshot, the remote server PID was alive in state `S`; its SSH process and the API tunnel process had no exit code or signal. The HTTPS front was listening, reported zero upstream errors, and gained no early closes during the Admin burst. `/healthz` and `/readyz` both returned HTTP 200 through both the HTTPS front and direct SSH tunnel. Load at the snapshot, inside the lock, was 3.33/2.18/1.02. These observations do not identify the earlier interruption's transport hop, so I made no server/API behavior change. The shared E2E front now exposes bounded, content-free transport counters for future diagnosis.
Author
Owner

Final report — Forgejo #705

Branch: job/admin-burst-705
Base: origin/dev at c4a61e8cf090170f35b1bed3350d9de20c83ecd5
Head: 23a6fe0e0e326f789c886f366880f5b86683b287

Built

  • Cherry-picked DESIGN §58 from job/instant-663 in two atomic documentation commits. This adds the #663 interaction rules and the cache access/evidence invariants.
  • Added bounded, content-free transport diagnostics to the shared E2E HTTPS front. It retains at most 32 recent request states and cumulative counts. It records no path, query, headers or body. Added a focused test for a normal response and an upstream reset.
  • No server route, API contract, Rust crate, migration or dependency changed. bun.lock and Cargo.lock are unchanged.

Commits:

  • 02679cb9e docs: define instant interaction rules and shared ownership for #663
  • 8ced71a02 docs: bind instant caches to access and require complete audit evidence
  • 23a6fe0e0 test: expose content-free HTTPS transport diagnostics

Bounded production/HDD retest

One run used bench/hdd-emu.sh, the embedded production SPA and the exact #705 server binary: source cc25c441b7a974185622a1dee853cf38686d2b67, SHA-256 2f3567d91c34839851247bc0acbc25a56aaacd14dca269b8f0342ddf83447ed9. Each measurement phase held flock -w 14400 /root/perf.lock. The 8 ms delay qualification passed: QD1 125.008329 IOPS, p50/p99 8.028160/8.159232 ms; QD16 200.932712 IOPS, p50/p99 100.139008/104.333312 ms.

The fixture matched #663: 366 Daily notes, 10,980 Logs, 100 Files and Photos, 20 Notes and Tasks, three Budgets and 100 transactions. The eight preceding endpoint profiles completed before Admin. For GET /api/v1/auth/users, all five serial reads and all five concurrent burst reads returned HTTP 200 with 198 response bytes and one User. No socket hang-up or non-200 status occurred.

Admin profile p50 / p95 / max CPU delta RSS after profile Load average in lock
5 serial reads 208.4 / 1165.2 / 1165.2 ms 310 ms 535,474,176 bytes 3.18 / 2.13 / 1.00
5-request burst 98.8 / 99.5 / 99.5 ms 90 ms 535,474,176 bytes 3.18 / 2.13 / 1.00

With n=5, nearest-rank p95 equals the maximum. The post-burst observation was also inside the lock at load 3.33 / 2.18 / 1.02. The remote server process was alive in state S; its SSH process and API tunnel had no exit code or signal. The HTTPS front was listening with zero upstream errors and no new early closes during the burst. /healthz and /readyz returned HTTP 200 through both the HTTPS front and the SSH tunnel.

The SSH tunnel had 203,205 stderr bytes classified as other; the raw text was not retained. This does not establish a tunnel failure: its process remained alive and all five Admin requests and health checks succeeded. The single retry therefore does not identify the hop responsible for the earlier hang-up. The Admin Users endpoint has no matching auth/users profile in docs/perf/baseline.json; these numbers are diagnostic and do not claim a budget pass or regression. The earlier independent #663 successful burst (81.6 / 82.2 / 82.2 ms) remains a separate sample, not a baseline.

The ignored, content-free run artifacts remain local at artifacts/admin-burst-705/ (diagnostic script SHA-256 bff4d6f62d05a893f5f60fe02cd654be86ca20d619915bc0ac8ca7188fee905e; results SHA-256 811835387bd050fe33a00c222d3922693183163cb25cc082e79dff299c714351). They are not committed.

Gates

cargo fmt --check: exit 0, no output. No Rust crate changed, so per-crate clippy/test gates were not applicable. No API route changed, so no cross-User route matrix or separate adversarial round was applicable; the requested live burst probe was completed.

Focused transport test (exit 0):

# Subtest: HTTPS front reports bounded, content-free transport state
ok 1 - HTTPS front reports bounded, content-free transport state
1..1
# tests 1
# suites 0
# pass 1
# fail 0
# cancelled 0
# skipped 0
# todo 0

bun run check (exit 0):

$ node scripts/check-user-storage.mjs && node scripts/check-type-tokens.mjs && node scripts/check-motion-tokens.mjs && svelte-kit sync && svelte-check --tsconfig ./tsconfig.json
User browser caches use userStorage; only documented device/public-link exceptions remain.
Text sizes and UI shape values use shared role tokens.
UI transitions and animation options use shared motion tokens or documented exceptions.
Loading svelte-check in workspace: /home/kayg/Developer/calternal-wt/admin-burst-705/apps/web
Getting Svelte diagnostics...

svelte-check found 0 errors and 0 warnings

bun run test (exit 0; Vitest also emitted the existing jsdom scrollTo() and CSS parser notices):

 Test Files  153 passed (153)
      Tests  1053 passed (1053)
   Start at  09:05:48
   Duration  81.21s (transform 52%, import 18%, environment 16%, tests 11%, setup 3%)

Environment  |component| jsdom was created 48 times · 63.00s total, 29% of tracked time
             create it once per worker with pool: 'vmThreads' (keeps per-file isolation) or isolate: false (shares it across files)
             learn more: https://vitest.dev/guide/improving-performance#test-environments

git diff --check: exit 0, no output. cargo clean:

     Removed 1 file, 356B total

The initial web check ran before dependency installation and stopped at missing typescript. bun install --frozen-lockfile installed the locked workspace packages; the lockfile remained unchanged, and the final web check above passed. apps/web/build and apps/web/.svelte-kit were removed.

UX gaps closed / left

  • Closed: no product UI gap changed. The shared test front can now distinguish a completed response from an upstream reset without retaining request content.
  • Left: the earlier intermittent Admin socket hang-up was not reproduced. No user-facing failure or recovery behavior was observed, so this job made no product behavior change.

Decisions and known gaps

  • DESIGN §58 does not set a retention size for test transport evidence. I capped the recent state list at 32 and kept cumulative counts; only method, status and transport state are retained.
  • I kept the change in the shared E2E front and did not alter the server because the bounded retest found no server/API failure. The original transport source remains unknown; the tunnel's unclassified stderr is an evidence gap.
  • A bare node --test of the whole existing harness.test.mjs also reported seven legacy theme cases with window is not defined. The focused #705 test passed, and the required Vitest gate passed; no existing expectation was changed.
  • No UI was changed, so visual screenshots were not applicable.

origin/dev was fetched and merged before the final gates; it was already up to date. No push, deploy or issue close was performed.

## Final report — Forgejo #705 Branch: `job/admin-burst-705` Base: `origin/dev` at `c4a61e8cf090170f35b1bed3350d9de20c83ecd5` Head: `23a6fe0e0e326f789c886f366880f5b86683b287` ### Built - Cherry-picked DESIGN §58 from `job/instant-663` in two atomic documentation commits. This adds the #663 interaction rules and the cache access/evidence invariants. - Added bounded, content-free transport diagnostics to the shared E2E HTTPS front. It retains at most 32 recent request states and cumulative counts. It records no path, query, headers or body. Added a focused test for a normal response and an upstream reset. - No server route, API contract, Rust crate, migration or dependency changed. `bun.lock` and `Cargo.lock` are unchanged. Commits: - `02679cb9e` docs: define instant interaction rules and shared ownership for #663 - `8ced71a02` docs: bind instant caches to access and require complete audit evidence - `23a6fe0e0` test: expose content-free HTTPS transport diagnostics ### Bounded production/HDD retest One run used `bench/hdd-emu.sh`, the embedded production SPA and the exact #705 server binary: source `cc25c441b7a974185622a1dee853cf38686d2b67`, SHA-256 `2f3567d91c34839851247bc0acbc25a56aaacd14dca269b8f0342ddf83447ed9`. Each measurement phase held `flock -w 14400 /root/perf.lock`. The 8 ms delay qualification passed: QD1 125.008329 IOPS, p50/p99 8.028160/8.159232 ms; QD16 200.932712 IOPS, p50/p99 100.139008/104.333312 ms. The fixture matched #663: 366 Daily notes, 10,980 Logs, 100 Files and Photos, 20 Notes and Tasks, three Budgets and 100 transactions. The eight preceding endpoint profiles completed before Admin. For `GET /api/v1/auth/users`, all five serial reads and all five concurrent burst reads returned HTTP 200 with 198 response bytes and one User. No socket hang-up or non-200 status occurred. | Admin profile | p50 / p95 / max | CPU delta | RSS after profile | Load average in lock | | --- | --- | ---: | ---: | --- | | 5 serial reads | 208.4 / 1165.2 / 1165.2 ms | 310 ms | 535,474,176 bytes | 3.18 / 2.13 / 1.00 | | 5-request burst | 98.8 / 99.5 / 99.5 ms | 90 ms | 535,474,176 bytes | 3.18 / 2.13 / 1.00 | With n=5, nearest-rank p95 equals the maximum. The post-burst observation was also inside the lock at load 3.33 / 2.18 / 1.02. The remote server process was alive in state `S`; its SSH process and API tunnel had no exit code or signal. The HTTPS front was listening with zero upstream errors and no new early closes during the burst. `/healthz` and `/readyz` returned HTTP 200 through both the HTTPS front and the SSH tunnel. The SSH tunnel had 203,205 stderr bytes classified as `other`; the raw text was not retained. This does not establish a tunnel failure: its process remained alive and all five Admin requests and health checks succeeded. The single retry therefore does not identify the hop responsible for the earlier hang-up. The Admin Users endpoint has no matching `auth/users` profile in `docs/perf/baseline.json`; these numbers are diagnostic and do not claim a budget pass or regression. The earlier independent #663 successful burst (81.6 / 82.2 / 82.2 ms) remains a separate sample, not a baseline. The ignored, content-free run artifacts remain local at `artifacts/admin-burst-705/` (diagnostic script SHA-256 `bff4d6f62d05a893f5f60fe02cd654be86ca20d619915bc0ac8ca7188fee905e`; results SHA-256 `811835387bd050fe33a00c222d3922693183163cb25cc082e79dff299c714351`). They are not committed. ### Gates `cargo fmt --check`: exit 0, no output. No Rust crate changed, so per-crate clippy/test gates were not applicable. No API route changed, so no cross-User route matrix or separate adversarial round was applicable; the requested live burst probe was completed. Focused transport test (exit 0): ```text # Subtest: HTTPS front reports bounded, content-free transport state ok 1 - HTTPS front reports bounded, content-free transport state 1..1 # tests 1 # suites 0 # pass 1 # fail 0 # cancelled 0 # skipped 0 # todo 0 ``` `bun run check` (exit 0): ```text $ node scripts/check-user-storage.mjs && node scripts/check-type-tokens.mjs && node scripts/check-motion-tokens.mjs && svelte-kit sync && svelte-check --tsconfig ./tsconfig.json User browser caches use userStorage; only documented device/public-link exceptions remain. Text sizes and UI shape values use shared role tokens. UI transitions and animation options use shared motion tokens or documented exceptions. Loading svelte-check in workspace: /home/kayg/Developer/calternal-wt/admin-burst-705/apps/web Getting Svelte diagnostics... svelte-check found 0 errors and 0 warnings ``` `bun run test` (exit 0; Vitest also emitted the existing jsdom `scrollTo()` and CSS parser notices): ```text Test Files 153 passed (153) Tests 1053 passed (1053) Start at 09:05:48 Duration 81.21s (transform 52%, import 18%, environment 16%, tests 11%, setup 3%) Environment |component| jsdom was created 48 times · 63.00s total, 29% of tracked time create it once per worker with pool: 'vmThreads' (keeps per-file isolation) or isolate: false (shares it across files) learn more: https://vitest.dev/guide/improving-performance#test-environments ``` `git diff --check`: exit 0, no output. `cargo clean`: ```text Removed 1 file, 356B total ``` The initial web check ran before dependency installation and stopped at missing `typescript`. `bun install --frozen-lockfile` installed the locked workspace packages; the lockfile remained unchanged, and the final web check above passed. `apps/web/build` and `apps/web/.svelte-kit` were removed. ### UX gaps closed / left - Closed: no product UI gap changed. The shared test front can now distinguish a completed response from an upstream reset without retaining request content. - Left: the earlier intermittent Admin socket hang-up was not reproduced. No user-facing failure or recovery behavior was observed, so this job made no product behavior change. ### Decisions and known gaps - DESIGN §58 does not set a retention size for test transport evidence. I capped the recent state list at 32 and kept cumulative counts; only method, status and transport state are retained. - I kept the change in the shared E2E front and did not alter the server because the bounded retest found no server/API failure. The original transport source remains unknown; the tunnel's unclassified stderr is an evidence gap. - A bare `node --test` of the whole existing `harness.test.mjs` also reported seven legacy theme cases with `window is not defined`. The focused #705 test passed, and the required Vitest gate passed; no existing expectation was changed. - No UI was changed, so visual screenshots were not applicable. `origin/dev` was fetched and merged before the final gates; it was already up to date. No push, deploy or issue close was performed.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#705
No description provided.