PERF: Search rule 8 — keep counts and background work off the first usable path (#663) #695

Open
opened 2026-10-02 05:36:49 +00:00 by kayg · 2 comments
Owner

Context: #663 matrix and DESIGN §58 rule 8. Scope: Search must keep counts and background work off the first usable path. Shared client machinery belongs only to #665–#668.

Evidence at c4a61e8cf0:

  • Endpoint: GET /api/v1/search.
  • Source: apps/web/src/lib/search/server.ts:145. Keyword first then hybrid after 80 ms is reuse; base Notes metadata loader still drains corpus and providers perform request-time detail work. Background read priority not established for all providers.
  • Representative endpoint numbers, server/client revisions, fixture size, lock/HDD qualification and limits are in the #663 production/HDD table. Those numbers do not prove this rule passes; the structural gap above is separate evidence.

Expected and regression tests:
Compare first-row/first-card paint with counts/stats deliberately delayed. Interactive input and usable rows must not await counts, thumbnail jobs or non-selected sidebars. Under a held writer and indexing/backfill batches, reads use a separate read-only pool and return a complete committed answer. Bound batches, yield between them and record CPU/RSS plus per-phase latency. Keep all visible totals correct; do not remove counts or replace them with fake values. Record accepted and durable timings via #667, without content or credentials.

Performance test:
Extend the existing bench profile for this hot path and the #549/#641 harness. Use production builds on root@10.69.69.63, bench/hdd-emu.sh and flock -w 14400 /root/perf.lock. Record load inside the lock; ≥5 samples, median/p95/max, average CPU/RSS and one realistic large-data/burst case. Separate warm, cold, accepted and durable boundaries. First usable 10k view ≤1.5 s; cached open/warm return ≤100 ms and accepted action ≤150 ms where applicable. Compare only a matching baseline in docs/perf/baseline.json; missing profiles require a new recorded baseline, not a made-up comparison.

Reuse/ownership:
Notes metadata drain is owned by the Notes rule-8 issue; this issue only adapts Search to bounded metadata and gives interactive queries priority over semantic/index work. Do not fix the same loader in two jobs.
Reuse #555 userStorage, #549 route caches, Files signed keysets/change feed and the existing Db reader_pool. #641 owns blaze measurements; #642 owns Settings opening; #640 owns Mail layouts; #639 owns linked Note opening. Integrate their active/completed branches before changing related code.

Acceptance:
Existing tests and status expectations stay intact. Run the per-crate gates (and calternal-server for route/contract changes), web gates if changed, and one time-boxed real-server regression/adversarial round for any new API contract. Preserve calternal-fs as the only filesystem interface and the server as the only writer. UI changes need pointer/touch/keyboard/screen-reader coverage and real-production captures at 390/820/1440 in both themes. Do not change shared motion for keyboard input.

Representative measurement from the audit (not a full-view budget result):
/api/v1/search?q=log&limit=100&semantic=false, five serial requests after fixture readiness, all HTTP 200. Median/p95/max 49.3/55.3/55.3 ms; five-request burst median/p95/max 52.1/54.2/54.2 ms. Serial window server CPU 130 ms, RSS 515657728 bytes; no ETag on these sampled responses. Load inside lock 3.69/2.16/0.96.
Shared release server source cc25c441b7a974185622a1dee853cf38686d2b67, binary SHA-256 2f3567d91c34839851247bc0acbc25a56aaacd14dca269b8f0342ddf83447ed9; embedded production SPA; Chromium browser on the build host through SSH/HTTPS. Server/Home/Index on perf VM HDD emulator: direct-I/O loop, 8 ms read/write dm-delay, 200 IOPS and 150 MiB/s caps. Every measured phase held flock -w 14400 /root/perf.lock. Qualification QD1 115.3 IOPS/8.028 ms median, QD16 200.7 IOPS/96.993 ms. Fixture: 366 Daily notes, 10,980 Logs, 100 Files/Photos, 20 Notes/Tasks, three Budgets and 100 transactions; Mail empty, Admin one User.
Structural source evidence above is the newer audit base, not the measured binary revision. No claim that these revisions are equivalent. The baseline in docs/perf/baseline.json uses another fixture/build/transport; no regression ratio is valid here. See #663 for matching baseline endpoint values and coverage gaps.

Context: #663 matrix and DESIGN §58 rule 8. Scope: Search must keep counts and background work off the first usable path. Shared client machinery belongs only to #665–#668. Evidence at c4a61e8cf090170f35b1bed3350d9de20c83ecd5: - Endpoint: `GET /api/v1/search`. - Source: `apps/web/src/lib/search/server.ts:145`. Keyword first then hybrid after 80 ms is reuse; base Notes metadata loader still drains corpus and providers perform request-time detail work. Background read priority not established for all providers. - Representative endpoint numbers, server/client revisions, fixture size, lock/HDD qualification and limits are in the #663 production/HDD table. Those numbers do not prove this rule passes; the structural gap above is separate evidence. Expected and regression tests: Compare first-row/first-card paint with counts/stats deliberately delayed. Interactive input and usable rows must not await counts, thumbnail jobs or non-selected sidebars. Under a held writer and indexing/backfill batches, reads use a separate read-only pool and return a complete committed answer. Bound batches, yield between them and record CPU/RSS plus per-phase latency. Keep all visible totals correct; do not remove counts or replace them with fake values. Record accepted and durable timings via #667, without content or credentials. Performance test: Extend the existing bench profile for this hot path and the #549/#641 harness. Use production builds on root@10.69.69.63, bench/hdd-emu.sh and flock -w 14400 /root/perf.lock. Record load inside the lock; ≥5 samples, median/p95/max, average CPU/RSS and one realistic large-data/burst case. Separate warm, cold, accepted and durable boundaries. First usable 10k view ≤1.5 s; cached open/warm return ≤100 ms and accepted action ≤150 ms where applicable. Compare only a matching baseline in docs/perf/baseline.json; missing profiles require a new recorded baseline, not a made-up comparison. Reuse/ownership: Notes metadata drain is owned by the Notes rule-8 issue; this issue only adapts Search to bounded metadata and gives interactive queries priority over semantic/index work. Do not fix the same loader in two jobs. Reuse #555 userStorage, #549 route caches, Files signed keysets/change feed and the existing Db reader_pool. #641 owns blaze measurements; #642 owns Settings opening; #640 owns Mail layouts; #639 owns linked Note opening. Integrate their active/completed branches before changing related code. Acceptance: Existing tests and status expectations stay intact. Run the per-crate gates (and calternal-server for route/contract changes), web gates if changed, and one time-boxed real-server regression/adversarial round for any new API contract. Preserve calternal-fs as the only filesystem interface and the server as the only writer. UI changes need pointer/touch/keyboard/screen-reader coverage and real-production captures at 390/820/1440 in both themes. Do not change shared motion for keyboard input. Representative measurement from the audit (not a full-view budget result): `/api/v1/search?q=log&limit=100&semantic=false`, five serial requests after fixture readiness, all HTTP 200. Median/p95/max 49.3/55.3/55.3 ms; five-request burst median/p95/max 52.1/54.2/54.2 ms. Serial window server CPU 130 ms, RSS 515657728 bytes; no ETag on these sampled responses. Load inside lock 3.69/2.16/0.96. Shared release server source `cc25c441b7a974185622a1dee853cf38686d2b67`, binary SHA-256 `2f3567d91c34839851247bc0acbc25a56aaacd14dca269b8f0342ddf83447ed9`; embedded production SPA; Chromium browser on the build host through SSH/HTTPS. Server/Home/Index on perf VM HDD emulator: direct-I/O loop, 8 ms read/write dm-delay, 200 IOPS and 150 MiB/s caps. Every measured phase held `flock -w 14400 /root/perf.lock`. Qualification QD1 115.3 IOPS/8.028 ms median, QD16 200.7 IOPS/96.993 ms. Fixture: 366 Daily notes, 10,980 Logs, 100 Files/Photos, 20 Notes/Tasks, three Budgets and 100 transactions; Mail empty, Admin one User. Structural source evidence above is the newer audit base, not the measured binary revision. No claim that these revisions are equivalent. The baseline in docs/perf/baseline.json uses another fixture/build/transport; no regression ratio is valid here. See #663 for matching baseline endpoint values and coverage gaps.
Author
Owner

Additional disk IO evidence from perf-arch-io (#663) for #695. Source origin/dev c4a61e8cf090170f35b1bed3350d9de20c83ecd5; queued Search changes checked at merge-round-7a 2f4482ded066d9c5d9c59130377907f7fd2916c9.

Keyword reconcile runs every five minutes (crates/calternal-search/src/indexer.rs:55,1175). In scan_tree:1498, indexed_file reads/parses source bytes BEFORE record_matches:1511 checks whether they changed. In indexed_file:2704-2707, every PDF up to 16 MiB is read, hashed and passed to the extractor on each scan, even with an unchanged manifest. Each full integrity pass first counts the tree (:2054) and then walks it again. The dedicated actor has nice=10 (:529) but no idle IO priority or byte pacing. Existing 1,024-file/32 MiB commit bounds and the minimum commit interval are valuable; they do not bound the scan's disk bandwidth.

merge-round-7a improves memory bounds with SQLite audit tables and paged directory walks (#496). It STILL calls indexed_file_for_actor before record_matches (indexer.rs:1878-1892), and its PDF path still extracts before that comparison (:3447 approximately; inspect the named function at this queued revision). Do not refile the already-fixed whole-Home memory issue.

Semantic indexing has its own five-minute scan (crates/calternal-embed/src/store.rs:406). prepare_path:1124 reads, hashes and chunks each text file before comparing existing_hash at :1140. Queued #503 bounds those scans and loads models on demand, but retains that read-before-hash-comparison order. Server sync_homes also sends Home upserts to the semantic worker before HTTP binds (wire.rs:2924), so lazy loading is still requested during startup for existing Users.

Reasoned impact: unchanged corpora get recurring random opens and source reads; cached source bytes still incur parsing/PDF subprocesses, SQL lookups and CPU contention. Independent keyword, semantic and Notes scans duplicate source work. No new latency number is claimed.

Fix within #695: share verified ingest source revisions/parsed projections where suitable; compare verified hashes before expensive PDF extraction/chunking; keep a bounded rotating integrity scan to detect same-size/mtime external changes. Pace background bytes and give request work priority between pages. Low CPU priority alone is insufficient for disk IO or model mutex wait.

Regression tests: a counting PDF extractor must not run on an unchanged second ordinary reconcile; replacing bytes with equal size/mtime must be detected by the integrity scan; pause a background page and prove keyword/semantic queries can finish. Keep deleted-hit and private-generation isolation guarantees. Measure large cold PDF/Markdown corpora plus interactive reads on the locked HDD emulator, with scan throughput and request p50/p95/max separated.

Additional disk IO evidence from perf-arch-io (#663) for #695. Source origin/dev `c4a61e8cf090170f35b1bed3350d9de20c83ecd5`; queued Search changes checked at merge-round-7a `2f4482ded066d9c5d9c59130377907f7fd2916c9`. Keyword reconcile runs every five minutes (`crates/calternal-search/src/indexer.rs:55,1175`). In `scan_tree:1498`, indexed_file reads/parses source bytes BEFORE `record_matches:1511` checks whether they changed. In `indexed_file:2704-2707`, every PDF up to 16 MiB is read, hashed and passed to the extractor on each scan, even with an unchanged manifest. Each full integrity pass first counts the tree (`:2054`) and then walks it again. The dedicated actor has nice=10 (`:529`) but no idle IO priority or byte pacing. Existing 1,024-file/32 MiB commit bounds and the minimum commit interval are valuable; they do not bound the scan's disk bandwidth. merge-round-7a improves memory bounds with SQLite audit tables and paged directory walks (#496). It STILL calls indexed_file_for_actor before record_matches (`indexer.rs:1878-1892`), and its PDF path still extracts before that comparison (`:3447` approximately; inspect the named function at this queued revision). Do not refile the already-fixed whole-Home memory issue. Semantic indexing has its own five-minute scan (`crates/calternal-embed/src/store.rs:406`). `prepare_path:1124` reads, hashes and chunks each text file before comparing existing_hash at :1140. Queued #503 bounds those scans and loads models on demand, but retains that read-before-hash-comparison order. Server sync_homes also sends Home upserts to the semantic worker before HTTP binds (`wire.rs:2924`), so lazy loading is still requested during startup for existing Users. Reasoned impact: unchanged corpora get recurring random opens and source reads; cached source bytes still incur parsing/PDF subprocesses, SQL lookups and CPU contention. Independent keyword, semantic and Notes scans duplicate source work. No new latency number is claimed. Fix within #695: share verified ingest source revisions/parsed projections where suitable; compare verified hashes before expensive PDF extraction/chunking; keep a bounded rotating integrity scan to detect same-size/mtime external changes. Pace background bytes and give request work priority between pages. Low CPU priority alone is insufficient for disk IO or model mutex wait. Regression tests: a counting PDF extractor must not run on an unchanged second ordinary reconcile; replacing bytes with equal size/mtime must be detected by the integrity scan; pause a background page and prove keyword/semantic queries can finish. Keep deleted-hit and private-generation isolation guarantees. Measure large cold PDF/Markdown corpora plus interactive reads on the locked HDD emulator, with scan throughput and request p50/p95/max separated.
Author
Owner

Memory/CPU audit #663: round 7a removes the resident manifest but still calls read_indexed_file before record_matches in keyword reconciliation (crates/calternal-search/src/indexer.rs:1880–1890 at 2f4482ded). Reconcile interval is five minutes. For 100k unchanged 4 KiB Notes, a complete pass reads/parses about 390.6 MiB: 1.30 MiB/s and 333 file visits/s if it finishes within the interval. This is a work model, not measured CPU or disk throughput. Include repeated unchanged-corpus reconciles in background-work acceptance; a byte budget alone does not bound duty cycle. Reuse #496/#503 fixes and test periodic work rather than reopening the old manifest issue.

Memory/CPU audit #663: round 7a removes the resident manifest but still calls read_indexed_file before record_matches in keyword reconciliation (crates/calternal-search/src/indexer.rs:1880–1890 at 2f4482ded). Reconcile interval is five minutes. For 100k unchanged 4 KiB Notes, a complete pass reads/parses about 390.6 MiB: 1.30 MiB/s and 333 file visits/s if it finishes within the interval. This is a work model, not measured CPU or disk throughput. Include repeated unchanged-corpus reconciles in background-work acceptance; a byte budget alone does not bound duty cycle. Reuse #496/#503 fixes and test periodic work rather than reopening the old manifest issue.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#695
No description provided.