PERF: Admin rule 8 — keep counts and background work off the first usable path (#663) #698

Open
opened 2026-10-02 05:36:53 +00:00 by kayg · 3 comments
Owner

Context: #663 matrix and DESIGN §58 rule 8. Scope: Admin must keep counts and background work off the first usable path. Shared client machinery belongs only to #665–#668.

Evidence at c4a61e8cf0:

  • Endpoint: GET /api/v1/auth/users.
  • Source: crates/calternal-auth/src/store.rs:1611. Auth queries use same pool as writes and return all users before paint. Core Db reader_pool at db.rs:69 is existing shared infrastructure to use.
  • Representative endpoint numbers, server/client revisions, fixture size, lock/HDD qualification and limits are in the #663 production/HDD table. Those numbers do not prove this rule passes; the structural gap above is separate evidence.

Expected and regression tests:
Compare first-row/first-card paint with counts/stats deliberately delayed. Interactive input and usable rows must not await counts, thumbnail jobs or non-selected sidebars. Under a held writer and indexing/backfill batches, reads use a separate read-only pool and return a complete committed answer. Bound batches, yield between them and record CPU/RSS plus per-phase latency. Keep all visible totals correct; do not remove counts or replace them with fake values. Record accepted and durable timings via #667, without content or credentials.

Performance test:
Extend the existing bench profile for this hot path and the #549/#641 harness. Use production builds on root@10.69.69.63, bench/hdd-emu.sh and flock -w 14400 /root/perf.lock. Record load inside the lock; ≥5 samples, median/p95/max, average CPU/RSS and one realistic large-data/burst case. Separate warm, cold, accepted and durable boundaries. First usable 10k view ≤1.5 s; cached open/warm return ≤100 ms and accepted action ≤150 ms where applicable. Compare only a matching baseline in docs/perf/baseline.json; missing profiles require a new recorded baseline, not a made-up comparison.

Reuse/ownership:

Reuse #555 userStorage, #549 route caches, Files signed keysets/change feed and the existing Db reader_pool. #641 owns blaze measurements; #642 owns Settings opening; #640 owns Mail layouts; #639 owns linked Note opening. Integrate their active/completed branches before changing related code.

Acceptance:
Existing tests and status expectations stay intact. Run the per-crate gates (and calternal-server for route/contract changes), web gates if changed, and one time-boxed real-server regression/adversarial round for any new API contract. Preserve calternal-fs as the only filesystem interface and the server as the only writer. UI changes need pointer/touch/keyboard/screen-reader coverage and real-production captures at 390/820/1440 in both themes. Do not change shared motion for keyboard input.

Representative measurement from the audit (not a full-view budget result):
/api/v1/auth/users, five serial requests after fixture readiness, all HTTP 200. Median/p95/max 140.4/929.4/929.4 ms; five-request burst median/p95/max 81.6/82.2/82.2 ms. Serial window server CPU 220 ms, RSS 515657728 bytes; no ETag on these sampled responses. Load inside lock 3.79/2.21/0.98.
Shared release server source cc25c441b7a974185622a1dee853cf38686d2b67, binary SHA-256 2f3567d91c34839851247bc0acbc25a56aaacd14dca269b8f0342ddf83447ed9; embedded production SPA; Chromium browser on the build host through SSH/HTTPS. Server/Home/Index on perf VM HDD emulator: direct-I/O loop, 8 ms read/write dm-delay, 200 IOPS and 150 MiB/s caps. Every measured phase held flock -w 14400 /root/perf.lock. Qualification QD1 115.3 IOPS/8.028 ms median, QD16 200.7 IOPS/96.993 ms. Fixture: 366 Daily notes, 10,980 Logs, 100 Files/Photos, 20 Notes/Tasks, three Budgets and 100 transactions; Mail empty, Admin one User.
Structural source evidence above is the newer audit base, not the measured binary revision. No claim that these revisions are equivalent. The baseline in docs/perf/baseline.json uses another fixture/build/transport; no regression ratio is valid here. See #663 for matching baseline endpoint values and coverage gaps.

Context: #663 matrix and DESIGN §58 rule 8. Scope: Admin must keep counts and background work off the first usable path. Shared client machinery belongs only to #665–#668. Evidence at c4a61e8cf090170f35b1bed3350d9de20c83ecd5: - Endpoint: `GET /api/v1/auth/users`. - Source: `crates/calternal-auth/src/store.rs:1611`. Auth queries use same pool as writes and return all users before paint. Core Db reader_pool at db.rs:69 is existing shared infrastructure to use. - Representative endpoint numbers, server/client revisions, fixture size, lock/HDD qualification and limits are in the #663 production/HDD table. Those numbers do not prove this rule passes; the structural gap above is separate evidence. Expected and regression tests: Compare first-row/first-card paint with counts/stats deliberately delayed. Interactive input and usable rows must not await counts, thumbnail jobs or non-selected sidebars. Under a held writer and indexing/backfill batches, reads use a separate read-only pool and return a complete committed answer. Bound batches, yield between them and record CPU/RSS plus per-phase latency. Keep all visible totals correct; do not remove counts or replace them with fake values. Record accepted and durable timings via #667, without content or credentials. Performance test: Extend the existing bench profile for this hot path and the #549/#641 harness. Use production builds on root@10.69.69.63, bench/hdd-emu.sh and flock -w 14400 /root/perf.lock. Record load inside the lock; ≥5 samples, median/p95/max, average CPU/RSS and one realistic large-data/burst case. Separate warm, cold, accepted and durable boundaries. First usable 10k view ≤1.5 s; cached open/warm return ≤100 ms and accepted action ≤150 ms where applicable. Compare only a matching baseline in docs/perf/baseline.json; missing profiles require a new recorded baseline, not a made-up comparison. Reuse/ownership: Reuse #555 userStorage, #549 route caches, Files signed keysets/change feed and the existing Db reader_pool. #641 owns blaze measurements; #642 owns Settings opening; #640 owns Mail layouts; #639 owns linked Note opening. Integrate their active/completed branches before changing related code. Acceptance: Existing tests and status expectations stay intact. Run the per-crate gates (and calternal-server for route/contract changes), web gates if changed, and one time-boxed real-server regression/adversarial round for any new API contract. Preserve calternal-fs as the only filesystem interface and the server as the only writer. UI changes need pointer/touch/keyboard/screen-reader coverage and real-production captures at 390/820/1440 in both themes. Do not change shared motion for keyboard input. Representative measurement from the audit (not a full-view budget result): `/api/v1/auth/users`, five serial requests after fixture readiness, all HTTP 200. Median/p95/max 140.4/929.4/929.4 ms; five-request burst median/p95/max 81.6/82.2/82.2 ms. Serial window server CPU 220 ms, RSS 515657728 bytes; no ETag on these sampled responses. Load inside lock 3.79/2.21/0.98. Shared release server source `cc25c441b7a974185622a1dee853cf38686d2b67`, binary SHA-256 `2f3567d91c34839851247bc0acbc25a56aaacd14dca269b8f0342ddf83447ed9`; embedded production SPA; Chromium browser on the build host through SSH/HTTPS. Server/Home/Index on perf VM HDD emulator: direct-I/O loop, 8 ms read/write dm-delay, 200 IOPS and 150 MiB/s caps. Every measured phase held `flock -w 14400 /root/perf.lock`. Qualification QD1 115.3 IOPS/8.028 ms median, QD16 200.7 IOPS/96.993 ms. Fixture: 366 Daily notes, 10,980 Logs, 100 Files/Photos, 20 Notes/Tasks, three Budgets and 100 transactions; Mail empty, Admin one User. Structural source evidence above is the newer audit base, not the measured binary revision. No claim that these revisions are equivalent. The baseline in docs/perf/baseline.json uses another fixture/build/transport; no regression ratio is valid here. See #663 for matching baseline endpoint values and coverage gaps.
Author
Owner

Additional #663 source evidence at c4a61e8cf: admin_jobs_view (server/src/wire.rs:3189) loads registered kinds and job rows, then awaits job_kind_control once for every registered kind. This adds K SQL round trips to every job snapshot/refresh. Batch the control rows with registrations/counters and page job details separately; verify query count stays constant as kinds increase. Several Security-state SELECT lists also use the writer pool (auth/src/store.rs:1351 users, :1180 App Passwords, :2046 sessions, :2183 passkeys), so pure reads can wait behind indexing. Keep authority current while moving those SELECTs to reader_pool. Config, integrity and scrub GETs already use in-memory snapshots; do not classify every Admin GET as a filesystem scan. No timing measurement.

Additional #663 source evidence at c4a61e8cf: admin_jobs_view (`server/src/wire.rs:3189`) loads registered kinds and job rows, then awaits job_kind_control once for every registered kind. This adds K SQL round trips to every job snapshot/refresh. Batch the control rows with registrations/counters and page job details separately; verify query count stays constant as kinds increase. Several Security-state SELECT lists also use the writer pool (`auth/src/store.rs:1351` users, :1180 App Passwords, :2046 sessions, :2183 passkeys), so pure reads can wait behind indexing. Keep authority current while moving those SELECTs to reader_pool. Config, integrity and scrub GETs already use in-memory snapshots; do not classify every Admin GET as a filesystem scan. No timing measurement.
Author
Owner

SQLite audit #663 on job/perf-arch-db. Base c4a61e8cf0; queued merge-round-7a 2f4482ded0 checked separately.

Precise Admin pool evidence: auth/src/store.rs users() (:1353), count_users() (:1360), list_users() (:1649) and list_invites() (:1613) borrow self.pool, which server wire.rs:1233 maps to the sole core writer. The read-only pool already exists and is used for normal User/session checks. Round 7a retains these writer-pool list methods. Use read_pool for independent reads and preserve current authorization/transaction invariants. Test under a held core writer, then separately bound list responses via #697. Do not attribute slow reads to WAL write locks when the actual wait is pool checkout.

Local query-work figures use Python SQLite 3.53.3 and synthetic data; VM callbacks count 1,000-instruction units. They are not production latency, CPU/RSS or HDD budget samples. The locked Rust bundled SQLite is 3.51.3. Repeat with audit-sqlite.py --plans and confirm the production-version plan before implementing. No existing assertion or product behavior was changed. Reuse this issue instead of filing a duplicate.

SQLite audit #663 on job/perf-arch-db. Base c4a61e8cf090170f35b1bed3350d9de20c83ecd5; queued merge-round-7a 2f4482ded066d9c5d9c59130377907f7fd2916c9 checked separately. Precise Admin pool evidence: auth/src/store.rs users() (:1353), count_users() (:1360), list_users() (:1649) and list_invites() (:1613) borrow self.pool, which server wire.rs:1233 maps to the sole core writer. The read-only pool already exists and is used for normal User/session checks. Round 7a retains these writer-pool list methods. Use read_pool for independent reads and preserve current authorization/transaction invariants. Test under a held core writer, then separately bound list responses via #697. Do not attribute slow reads to WAL write locks when the actual wait is pool checkout. Local query-work figures use Python SQLite 3.53.3 and synthetic data; VM callbacks count 1,000-instruction units. They are not production latency, CPU/RSS or HDD budget samples. The locked Rust bundled SQLite is 3.51.3. Repeat with audit-sqlite.py --plans and confirm the production-version plan before implementing. No existing assertion or product behavior was changed. Reuse this issue instead of filing a duplicate.
Author
Owner

SQLite audit #663. Base c4a61e8cf0; queued round-7a 2f4482ded0.

Source-line correction: at base c4a61e8cf, list_users() starts at store.rs:1611 and fetches at :1613; list_invites() starts at :1646 and fetches at :1649. I transposed those names in the prior comment. Both use self.pool, the sole server writer. The finding is unchanged.

Figures are local Python SQLite 3.53.3 query work in 1,000-instruction callback units, not production latency. Rust bundles 3.51.3. Recheck production plans before implementation; no product edit or assertion change.

SQLite audit #663. Base c4a61e8cf090170f35b1bed3350d9de20c83ecd5; queued round-7a 2f4482ded066d9c5d9c59130377907f7fd2916c9. Source-line correction: at base c4a61e8cf, list_users() starts at store.rs:1611 and fetches at :1613; list_invites() starts at :1646 and fetches at :1649. I transposed those names in the prior comment. Both use self.pool, the sole server writer. The finding is unchanged. Figures are local Python SQLite 3.53.3 query work in 1,000-instruction callback units, not production latency. Rust bundles 3.51.3. Recheck production plans before implementation; no product edit or assertion change.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#698
No description provided.