PERF: Search indexing peaks at 1.9 GB RSS for 1M files and 1.1 GB during a 100k rebuild (caps cover only the writer) #496

Open
opened 2026-09-30 08:04:40 +00:00 by kayg · 6 comments
Owner

Problem

Search indexing has a memory ceiling well above its configured caps. On a 7.7 GiB host, a 1M-file first index peaks at 1.9 GiB RSS and a 100k full rebuild at 1.1 GiB, while the configured Tantivy writer heap is 61 MiB and the pending batch cap is 32 MiB / 1,024 items.

Found by the #367 perf round (perf-test VM, 4 vCPU; release 369ab6a2f; tests/perf/search_scale.py; load < 1 at start under /root/perf.lock).

Measurements

Workload Duration Peak RSS Mean RSS Peak CPU
1M-file first index (search-stress) 283.6 s 1,987,190,784 B (1.85 GiB) — 223.7%
100k mixed first index 22.65 s 501,870,592 B — 387.5%
100k full rebuild while querying 39.8 s 1,188,155,392 B (1.11 GiB) 1,110,614,400 B 399.5%
100k semantic embedding backfill — 916,357,120 B 881,016,111 B 346.1%
Idle server, 10 Users 600 s 162.4 MB 154.7 MiB 4.0%

Thread CPU during the 1M first index: calternal-searc 56.9 s, merge_thread_0 26.3 s, notify-rs inoti 20.4 s, sqlite 9.9 s, tokio 10.2 s.

Queries stay available during the rebuild (0 failures), but at p50 / p95 / p99 = 203.7 / 265.4 / 385.0 ms versus 6.2 / 18.1 / 23.5 ms at rest.

Why (to confirm)

The caps bound the Tantivy writer and the pending update batch only. The remaining peak is outside them: candidates are the startup/rebuild path collecting the whole manifest or path list in memory, merge threads holding several segments, and the notify watcher state for 1M paths. A heap profile (heaptrack or dhat) of search_scale.py --profile stress --items 200000 would attribute it.

Suggested direction (not decided)

Stream the rebuild/manifest diff in bounded pages and make every remaining buffer a named config value with a documented default (the #367 rule). Related: #338 (1M Home 503s).

Raw: docs/perf/runs/2026-09-29T224501Z-369ab6a2/raw/search-{mixed,stress}.json.

## Problem Search indexing has a memory ceiling well above its configured caps. On a 7.7 GiB host, a 1M-file first index peaks at 1.9 GiB RSS and a 100k full rebuild at 1.1 GiB, while the configured Tantivy writer heap is 61 MiB and the pending batch cap is 32 MiB / 1,024 items. Found by the #367 perf round (perf-test VM, 4 vCPU; release `369ab6a2f`; `tests/perf/search_scale.py`; load < 1 at start under `/root/perf.lock`). ## Measurements | Workload | Duration | Peak RSS | Mean RSS | Peak CPU | |---|---:|---:|---:|---:| | 1M-file first index (`search-stress`) | 283.6 s | **1,987,190,784 B (1.85 GiB)** | — | 223.7% | | 100k mixed first index | 22.65 s | 501,870,592 B | — | 387.5% | | 100k full rebuild while querying | 39.8 s | **1,188,155,392 B (1.11 GiB)** | 1,110,614,400 B | 399.5% | | 100k semantic embedding backfill | — | 916,357,120 B | 881,016,111 B | 346.1% | | Idle server, 10 Users | 600 s | 162.4 MB | 154.7 MiB | 4.0% | Thread CPU during the 1M first index: `calternal-searc` 56.9 s, `merge_thread_0` 26.3 s, `notify-rs inoti` 20.4 s, sqlite 9.9 s, tokio 10.2 s. Queries stay available during the rebuild (0 failures), but at p50 / p95 / p99 = 203.7 / 265.4 / 385.0 ms versus 6.2 / 18.1 / 23.5 ms at rest. ## Why (to confirm) The caps bound the Tantivy writer and the pending update batch only. The remaining peak is outside them: candidates are the startup/rebuild path collecting the whole manifest or path list in memory, merge threads holding several segments, and the `notify` watcher state for 1M paths. A heap profile (heaptrack or dhat) of `search_scale.py --profile stress --items 200000` would attribute it. ## Suggested direction (not decided) Stream the rebuild/manifest diff in bounded pages and make every remaining buffer a named config value with a documented default (the #367 rule). Related: #338 (1M Home 503s). Raw: `docs/perf/runs/2026-09-29T224501Z-369ab6a2/raw/search-{mixed,stress}.json`.
Author
Owner

heaptrack attribution (30% large Home, 300k files; details in #503): the Search rebuild path holds whole-Home maps at the same time. BTreeMap<String, ManifestEntry> 127.4 MB (flush_upserts / scan_tree / reconcile_and_report), String::clone 175.4 MB (file_result_ids, scan_tree, manifest_entry), hashbrown tables 126.5 MB (file_result_ids, SearchIndex::documents_by_path) and Vec growth 50.9 MB in documents_by_path (rebuild_all). That is about 0.48 GB at 300k files and grows linearly, which accounts for most of the 1.9 GB at 1M files. None of these structures is covered by the writer-heap or pending-batch caps.

heaptrack attribution (30% large Home, 300k files; details in #503): the Search rebuild path holds whole-Home maps at the same time. `BTreeMap<String, ManifestEntry>` 127.4 MB (`flush_upserts` / `scan_tree` / `reconcile_and_report`), `String::clone` 175.4 MB (`file_result_ids`, `scan_tree`, `manifest_entry`), `hashbrown` tables 126.5 MB (`file_result_ids`, `SearchIndex::documents_by_path`) and `Vec` growth 50.9 MB in `documents_by_path` (`rebuild_all`). That is about 0.48 GB at 300k files and grows linearly, which accounts for most of the 1.9 GB at 1M files. None of these structures is covered by the writer-heap or pending-batch caps.
Author
Owner

Starting #496 on branch job/perf-496, based on 6c87f5ff9442cd658572139bc536d018fd5222a4 (merge-base with origin/dev). I am tracing the full indexing memory path, then will add bounded scan/rebuild memory with a regression test and compare the issue workload.

Starting #496 on branch `job/perf-496`, based on `6c87f5ff9442cd658572139bc536d018fd5222a4` (merge-base with `origin/dev`). I am tracing the full indexing memory path, then will add bounded scan/rebuild memory with a regression test and compare the issue workload.
Author
Owner

I traced the reported heap profile to the current implementation: load_manifest materializes all search_manifest rows in a BTreeMap; documents_by_path builds a second whole-Index map; and reconcile_and_report adds whole-Home expected and seen maps. Root::list_page already provides a directory-handle-relative bounded page API. I am using it with disk-backed reconciliation rows and bounded batches so path count no longer sets the actor's working-set size.

I traced the reported heap profile to the current implementation: `load_manifest` materializes all `search_manifest` rows in a `BTreeMap`; `documents_by_path` builds a second whole-Index map; and `reconcile_and_report` adds whole-Home `expected` and `seen` maps. `Root::list_page` already provides a directory-handle-relative bounded page API. I am using it with disk-backed reconciliation rows and bounded batches so path count no longer sets the actor's working-set size.
Author
Owner

I am handling related issue #503. I found the committed Search memory fix (0072a577c, fa52fa1d7) on job/perf-496; it is not yet present in origin/dev. To avoid duplicate Search implementation, I will integrate those existing commits into #503 after aligning with current origin/dev, then add the separate ONNX and SQLite/buffer work. Please flag if you are already coordinating a different integration.

I am handling related issue #503. I found the committed Search memory fix (`0072a577c`, `fa52fa1d7`) on `job/perf-496`; it is not yet present in `origin/dev`. To avoid duplicate Search implementation, I will integrate those existing commits into #503 after aligning with current `origin/dev`, then add the separate ONNX and SQLite/buffer work. Please flag if you are already coordinating a different integration.
Author
Owner

The first 1M-file profile completed and passed integrity and keyword-hit verification. Local results: 1,000,000 files / 1,001,003 filesystem paths; 6,629.89s initial index; peak RSS 303,697,920 B (289.5 MiB); peak CPU 151.63%; worker budget 268,435,456 B, writer cap 67,108,864 B, batch cap 33,554,432 B. The issue baseline was 1,987,190,784 B (1.85 GiB), so measured peak RSS fell about 85%. Host load was 26.24/32.01/28.84 at start and 12.80/15.96/19.36 at end; the perf VM lock was held, so this is a noisy local timing and not a controlled latency comparison.

The 6,629.89s runtime is unexpectedly high. The new scan currently performs one SQLite manifest lookup per visited path. I am replacing those lookups with bounded per-directory-page queries, then I will rerun the 1M profile to check that the memory reduction does not retain a per-item database round trip.

The first 1M-file profile completed and passed integrity and keyword-hit verification. Local results: 1,000,000 files / 1,001,003 filesystem paths; 6,629.89s initial index; peak RSS 303,697,920 B (289.5 MiB); peak CPU 151.63%; worker budget 268,435,456 B, writer cap 67,108,864 B, batch cap 33,554,432 B. The issue baseline was 1,987,190,784 B (1.85 GiB), so measured peak RSS fell about 85%. Host load was 26.24/32.01/28.84 at start and 12.80/15.96/19.36 at end; the perf VM lock was held, so this is a noisy local timing and not a controlled latency comparison. The 6,629.89s runtime is unexpectedly high. The new scan currently performs one SQLite manifest lookup per visited path. I am replacing those lookups with bounded per-directory-page queries, then I will rerun the 1M profile to check that the memory reduction does not retain a per-item database round trip.
Author
Owner

perf-496 report

Head: 50a4bbf049030e84e8397fa3864e606bdb8e5772 (job/perf-496). No push or merge was made.

Built

  • Replaced whole-Home manifest and reconcile maps with SQLite-backed staging and audit tables, and bounded scan, event, and Tantivy update buffers.
  • Batched manifest reads by directory page. Each SQL query binds at most 900 paths to remain below SQLite's older 999-parameter limit.
  • Added tests for budget scaling, disk-backed audit flushing, bounded directory pagination, and page lookup chunking across the SQL parameter boundary.
  • Added search profile reporting for worker, writer, page, batch, and watcher caps.

Files: crates/calternal-search/migrations/0004_bounded_reconcile.sql, crates/calternal-search/src/index.rs, crates/calternal-search/src/indexer.rs, crates/calternal-search/src/lib.rs, crates/calternal-search/tests/indexer.rs, bench/record.py, bench/test_record.py, tests/perf/search_scale.py.

Validation

cargo fmt --all --check: exit 0, no output.

cargo clippy -p calternal-search --all-targets -- -D warnings output:

    Checking calternal-search v0.0.1 (/home/kayg/Developer/calternal-wt/perf-496/crates/calternal-search)
    Finished `dev` profile [unoptimized + debuginfo] target(s) in 4.57s

cargo test -p calternal-search result lines:

test result: ok. 37 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 5.61s
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.67s
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.05s
test result: ok. 20 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 14.67s
test result: ok. 5 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.04s
test result: ok. 2 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.01s
test result: ok. 1 passed; 0 failed; 2 ignored; 0 measured; 0 filtered out; finished in 2.69s
test result: ok. 4 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.01s
test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s

python3 -m unittest bench.test_record: 6 tests passed. cargo build --release -p calternal-server completed. cargo clean output: Removed 14420 files, 6.2GiB total. Web build output was deleted; no web source changed.

Performance

Issue baseline, perf VM: 1M-file first index 283.6s, peak RSS 1,987,190,784 B (1.85 GiB), peak CPU 223.7%.

Completed local 1M profile before the final page-batching optimization: 1,000,000 files / 1,001,003 paths; 6,629.89s; peak RSS 303,697,920 B (289.5 MiB), peak CPU 151.63%. The measured peak was about 85% lower than baseline. Integrity coverage and the live keyword probe passed. Worker budget was 268,435,456 B; writer cap 67,108,864 B; pending batch cap 33,554,432 B; page cap 256; watcher queue 8,192. Host load average was 26.24/32.01/28.84 at start and 12.80/15.96/19.36 at end. The perf VM lock was held, so this noisy local timing is not comparable to the isolated baseline.

I then added per-page manifest reads to remove the per-file SQL round trip. The repeated 1M run reached 430,876/1,001,003 paths when the four-hour job time box required stopping it. It has no completed peak or timing result. Its start/end host loads were 18.73/17.31/16.71 and 37.91/41.06/35.25. The partial corpus was removed.

Gaps and decisions

  • The final page-batched revision still needs a complete 1M profile and a 100k mixed rebuild profile on the perf VM under /root/perf.lock. The one time-boxed adversarial probe also remains to run.
  • I chose a 256 MiB default worker budget and kept storage-dependent reconciliation state in SQLite. The design did not specify the exact budget or query page size; I capped SQL batches at 900 parameters for compatibility with SQLite builds using the older limit.
# perf-496 report Head: `50a4bbf049030e84e8397fa3864e606bdb8e5772` (`job/perf-496`). No push or merge was made. ## Built - Replaced whole-Home manifest and reconcile maps with SQLite-backed staging and audit tables, and bounded scan, event, and Tantivy update buffers. - Batched manifest reads by directory page. Each SQL query binds at most 900 paths to remain below SQLite's older 999-parameter limit. - Added tests for budget scaling, disk-backed audit flushing, bounded directory pagination, and page lookup chunking across the SQL parameter boundary. - Added search profile reporting for worker, writer, page, batch, and watcher caps. Files: `crates/calternal-search/migrations/0004_bounded_reconcile.sql`, `crates/calternal-search/src/index.rs`, `crates/calternal-search/src/indexer.rs`, `crates/calternal-search/src/lib.rs`, `crates/calternal-search/tests/indexer.rs`, `bench/record.py`, `bench/test_record.py`, `tests/perf/search_scale.py`. ## Validation `cargo fmt --all --check`: exit 0, no output. `cargo clippy -p calternal-search --all-targets -- -D warnings` output: ```text Checking calternal-search v0.0.1 (/home/kayg/Developer/calternal-wt/perf-496/crates/calternal-search) Finished `dev` profile [unoptimized + debuginfo] target(s) in 4.57s ``` `cargo test -p calternal-search` result lines: ```text test result: ok. 37 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 5.61s test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.67s test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.05s test result: ok. 20 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 14.67s test result: ok. 5 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.04s test result: ok. 2 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.01s test result: ok. 1 passed; 0 failed; 2 ignored; 0 measured; 0 filtered out; finished in 2.69s test result: ok. 4 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.01s test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s ``` `python3 -m unittest bench.test_record`: 6 tests passed. `cargo build --release -p calternal-server` completed. `cargo clean` output: `Removed 14420 files, 6.2GiB total`. Web build output was deleted; no web source changed. ## Performance Issue baseline, perf VM: 1M-file first index 283.6s, peak RSS 1,987,190,784 B (1.85 GiB), peak CPU 223.7%. Completed local 1M profile before the final page-batching optimization: 1,000,000 files / 1,001,003 paths; 6,629.89s; peak RSS 303,697,920 B (289.5 MiB), peak CPU 151.63%. The measured peak was about 85% lower than baseline. Integrity coverage and the live keyword probe passed. Worker budget was 268,435,456 B; writer cap 67,108,864 B; pending batch cap 33,554,432 B; page cap 256; watcher queue 8,192. Host load average was 26.24/32.01/28.84 at start and 12.80/15.96/19.36 at end. The perf VM lock was held, so this noisy local timing is not comparable to the isolated baseline. I then added per-page manifest reads to remove the per-file SQL round trip. The repeated 1M run reached 430,876/1,001,003 paths when the four-hour job time box required stopping it. It has no completed peak or timing result. Its start/end host loads were 18.73/17.31/16.71 and 37.91/41.06/35.25. The partial corpus was removed. ## Gaps and decisions - The final page-batched revision still needs a complete 1M profile and a 100k mixed rebuild profile on the perf VM under `/root/perf.lock`. The one time-boxed adversarial probe also remains to run. - I chose a 256 MiB default worker budget and kept storage-dependent reconciliation state in SQLite. The design did not specify the exact budget or query page size; I capped SQL batches at 900 parameters for compatibility with SQLite builds using the older limit.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#496
No description provided.