PERF: a 1M-file Home drives the server to 6.65 GB RSS on a 7.7 GiB host (Search maps, two ONNX models, SQLite) #503

Open
opened 2026-09-30 08:44:27 +00:00 by kayg · 7 comments
Owner

Problem

Opening a large Home drives the server to 6.65 GB peak RSS on a 7.7 GiB host (884 MB left for everything else). Almost all of it is anonymous heap, so it does not come from mmapped indexes and the kernel cannot reclaim it. A smaller host would be OOM-killed.

Found by the #367 perf round (perf-test VM, 4 vCPU / 7.7 GiB; release 369ab6a2f; photos-perf.mjs --items 20000 --files 1000000 --notes 100000 --analytics-entries 10000 --large-note, i.e. bench/run.sh --full large Home; lock held, load < 1 at start).

Measurements

Run 2 server memory from /proc/<pid>/smaps_rollup, sampled every 30 s:

Server age RSS Anonymous
156 s 1.39 GB 1.32 GB
517 s 2.88 GB 2.81 GB
790 s 2.50 GB 2.43 GB
1,031 s 4.99 GB 4.97 GB
1,514 s 5.16 GB 5.13 GB
2,754 s 6.28 GB 6.26 GB

VmHWM 6,654,084 kB. Run 1 (same Home, same build): RSS 3.79 GB (anonymous 3.69 GB) at 24 min, VmHWM 4.20 GB at 55 min; 2 h 08 min of server CPU in 55 min of wall time.

Index timeline (run 1): 5k-item folder at ~210 s, the other 995k files at ~1,330 s, 100,001 Notes at ~1,520 s. Photos never appeared (#495, #470).

CPU at 10 min and 25 min (perf, 30 s at 99 Hz): ONNX embedding 26–27% (MlasGemmU8U8KernelAvx2), sqlite walFindFrame 5.9% (large WAL), subjournalPageIfRequired present; threads 49% tokio workers, 46% sqlite worker. A notes.reconcile job was still leased at 47 min.

For scale: idle server with 10 Users is 162 MB; the 1M-file keyword index alone peaks at 1.9 GB (#496).

Why (heaptrack)

heaptrack on a 30% Home (300k files, 30k Notes, 6k Photos, 3k Log entries; same build; the server ran under libheaptrack_preload.so; 748 M allocation calls). Peak heap 1.82 GB (peak RSS with heaptrack 2.83 GB). Largest peak consumers:

Peak Where Notes
447.6 MB ONNX Runtime arenas text embedding EmbeddingModel::embed_batch session run 134 + 84 + 34 MB; CLIP session prepack (OrtClipInference::load) 86 + 38 MB. Both models are resident at once.
345.6 MB sqlite3MemMalloc 245 MB in one sqlx connection worker, 87 MB in another (page cache and statement memory under long transactions)
175.4 MB String::clone Search indexer::file_result_ids, IndexActor::scan_tree (4 × 25 MB), manifest_entry 19 MB
153.4 MB SqliteRow::current 4.4 M rows buffered by fetch_all-style reads
127.4 MB BTreeMap<String, ManifestEntry>::insert Search flush_upserts / scan_tree / reconcile_and_report hold the whole manifest
126.5 MB hashbrown tables in calternal_search file_result_ids 50 MB, SearchIndex::documents_by_path 2 × 38 MB
66.1 MB Files reconcile_folder_locked startup reconcile_all
50.9 MB Vec growth in SearchIndex::documents_by_path rebuild_all → reconcile_and_report

So at 30% scale about 0.6 GB is the Search rebuild's whole-Home maps (manifest, path→document, result IDs), 0.45 GB is two ONNX models with arenas, and 0.5 GB is SQLite and buffered rows. The Search part grows linearly with files, which matches the 1.9 GB 1M-file Search peak in #496. Trace: /root/perf-367/profiles/ht-lh.1288364.zst on the perf VM.

Suggested direction (not decided)

Bound each startup scanner and backfill by a named config value (the #367 rule: every cap is a documented config value), stream instead of collecting, and do not run the embedding backfill, Notes reconcile and Photos rebuild at full width at the same time on a fresh large Home.

## Problem Opening a large Home drives the server to **6.65 GB peak RSS on a 7.7 GiB host** (884 MB left for everything else). Almost all of it is anonymous heap, so it does not come from mmapped indexes and the kernel cannot reclaim it. A smaller host would be OOM-killed. Found by the #367 perf round (perf-test VM, 4 vCPU / 7.7 GiB; release `369ab6a2f`; `photos-perf.mjs --items 20000 --files 1000000 --notes 100000 --analytics-entries 10000 --large-note`, i.e. `bench/run.sh --full` large Home; lock held, load < 1 at start). ## Measurements Run 2 server memory from `/proc/<pid>/smaps_rollup`, sampled every 30 s: | Server age | RSS | Anonymous | |---:|---:|---:| | 156 s | 1.39 GB | 1.32 GB | | 517 s | 2.88 GB | 2.81 GB | | 790 s | 2.50 GB | 2.43 GB | | 1,031 s | 4.99 GB | 4.97 GB | | 1,514 s | 5.16 GB | 5.13 GB | | 2,754 s | 6.28 GB | 6.26 GB | VmHWM 6,654,084 kB. Run 1 (same Home, same build): RSS 3.79 GB (anonymous 3.69 GB) at 24 min, VmHWM 4.20 GB at 55 min; 2 h 08 min of server CPU in 55 min of wall time. Index timeline (run 1): 5k-item folder at ~210 s, the other 995k files at ~1,330 s, 100,001 Notes at ~1,520 s. Photos never appeared (#495, #470). CPU at 10 min and 25 min (`perf`, 30 s at 99 Hz): ONNX embedding 26–27% (`MlasGemmU8U8KernelAvx2`), sqlite `walFindFrame` 5.9% (large WAL), `subjournalPageIfRequired` present; threads 49% tokio workers, 46% sqlite worker. A `notes.reconcile` job was still leased at 47 min. For scale: idle server with 10 Users is 162 MB; the 1M-file keyword index alone peaks at 1.9 GB (#496). ## Why (heaptrack) heaptrack on a 30% Home (300k files, 30k Notes, 6k Photos, 3k Log entries; same build; the server ran under `libheaptrack_preload.so`; 748 M allocation calls). **Peak heap 1.82 GB** (peak RSS with heaptrack 2.83 GB). Largest peak consumers: | Peak | Where | Notes | |---:|---|---| | 447.6 MB | ONNX Runtime arenas | text embedding `EmbeddingModel::embed_batch` session run 134 + 84 + 34 MB; CLIP session prepack (`OrtClipInference::load`) 86 + 38 MB. Both models are resident at once. | | 345.6 MB | `sqlite3MemMalloc` | 245 MB in one sqlx connection worker, 87 MB in another (page cache and statement memory under long transactions) | | 175.4 MB | `String::clone` | Search `indexer::file_result_ids`, `IndexActor::scan_tree` (4 × 25 MB), `manifest_entry` 19 MB | | 153.4 MB | `SqliteRow::current` | 4.4 M rows buffered by `fetch_all`-style reads | | 127.4 MB | `BTreeMap<String, ManifestEntry>::insert` | Search `flush_upserts` / `scan_tree` / `reconcile_and_report` hold the whole manifest | | 126.5 MB | `hashbrown` tables in `calternal_search` | `file_result_ids` 50 MB, `SearchIndex::documents_by_path` 2 × 38 MB | | 66.1 MB | Files `reconcile_folder_locked` | startup `reconcile_all` | | 50.9 MB | `Vec` growth in `SearchIndex::documents_by_path` | `rebuild_all` → `reconcile_and_report` | So at 30% scale about 0.6 GB is the Search rebuild's whole-Home maps (manifest, path→document, result IDs), 0.45 GB is two ONNX models with arenas, and 0.5 GB is SQLite and buffered rows. The Search part grows linearly with files, which matches the 1.9 GB 1M-file Search peak in #496. Trace: `/root/perf-367/profiles/ht-lh.1288364.zst` on the perf VM. ## Suggested direction (not decided) Bound each startup scanner and backfill by a named config value (the #367 rule: every cap is a documented config value), stream instead of collecting, and do not run the embedding backfill, Notes reconcile and Photos rebuild at full width at the same time on a fresh large Home.
Author
Owner

Starting #503 on branch job/fix-503. Worktree base: 558457cf32e1d429da3834a05ff2720d11284302 (merge-base with current origin/dev; origin/dev is cc25c441b7a974185622a1dee853cf38686d2b67, 471 commits ahead). I will trace the current Search overlap with #496, implement and commit atomic fixes, then fetch and merge origin/dev once before final per-crate gates as instructed.

Starting #503 on branch `job/fix-503`. Worktree base: `558457cf32e1d429da3834a05ff2720d11284302` (merge-base with current `origin/dev`; `origin/dev` is `cc25c441b7a974185622a1dee853cf38686d2b67`, 471 commits ahead). I will trace the current Search overlap with #496, implement and commit atomic fixes, then fetch and merge `origin/dev` once before final per-crate gates as instructed.
Author
Owner

Finding: DbOptions.max_readers was passed directly to SQLx with no upper bound. The regression test requested 1,024 readers; before the fix, the pool accepted that value. An exploratory u32::MAX request made SQLx abort while allocating its pool semaphore (memory allocation of 309237645240 bytes failed). Commit a3410a1ba clamps readers to 8, caps the per-connection page cache at 1 MiB and the statement cache at 32, and tests the writer and reader pool settings. Search map streaming remains with #496; its current patch conflicts with newer per-User Search and Sidecar behavior, so I am not copying it here.

Finding: `DbOptions.max_readers` was passed directly to SQLx with no upper bound. The regression test requested 1,024 readers; before the fix, the pool accepted that value. An exploratory `u32::MAX` request made SQLx abort while allocating its pool semaphore (`memory allocation of 309237645240 bytes failed`). Commit `a3410a1ba` clamps readers to 8, caps the per-connection page cache at 1 MiB and the statement cache at 32, and tests the writer and reader pool settings. Search map streaming remains with #496; its current patch conflicts with newer per-User Search and Sidecar behavior, so I am not copying it here.
Author
Owner

The #503 implementation slice is committed as 2153a341b793b32f370d944e8391a02afdeb7fa0.

It pages Home scans through a bounded queue, records seen paths in bounded SQLite batches, and prunes stale rows with anti-joins. SQLite connections now share the #503 cache/reader caps from calternal-db. Semantic and CLIP ONNX sessions initialize once on demand; deleting vectors does not load an inference model. Search map streaming remains with #496 because its branch conflicts with current per-User isolation and Sidecar changes.

Regression coverage: 300-file multi-page scan, 300-row stale cleanup across multiple batches, shared concurrent model initialization, no model load on removal, and existing hidden-vector cleanup. Gates: cargo fmt --all --check passed; cargo clippy -p calternal-embed --all-targets -- -D warnings passed; cargo test -p calternal-embed passed (36 passed, 4 ignored).

I am building the release server and preparing the locked perf-VM run. I will post the before/after measurement and final head SHA when it completes.

The #503 implementation slice is committed as `2153a341b793b32f370d944e8391a02afdeb7fa0`. It pages Home scans through a bounded queue, records seen paths in bounded SQLite batches, and prunes stale rows with anti-joins. SQLite connections now share the #503 cache/reader caps from `calternal-db`. Semantic and CLIP ONNX sessions initialize once on demand; deleting vectors does not load an inference model. Search map streaming remains with #496 because its branch conflicts with current per-User isolation and Sidecar changes. Regression coverage: 300-file multi-page scan, 300-row stale cleanup across multiple batches, shared concurrent model initialization, no model load on removal, and existing hidden-vector cleanup. Gates: `cargo fmt --all --check` passed; `cargo clippy -p calternal-embed --all-targets -- -D warnings` passed; `cargo test -p calternal-embed` passed (36 passed, 4 ignored). I am building the release server and preparing the locked perf-VM run. I will post the before/after measurement and final head SHA when it completes.
Author
Owner

Local #503 profile update (provisional): the perf VM run could not start because flock -n /root/perf.lock found the lock busy, so I started one local fallback run. At about 38 minutes, the server was using 1.36 GiB RSS; the 5,000-file folder and 10,000 Analytics entries were visible, while the remaining 995,000 files and Notes still reported zero. This is a point-in-time observation, not a peak or a completed before/after result. The local host load average was 33.72, 31.87, 29.68. The one-hour run remains in progress; I will report its final result and limitation separately.

Local #503 profile update (provisional): the perf VM run could not start because `flock -n /root/perf.lock` found the lock busy, so I started one local fallback run. At about 38 minutes, the server was using 1.36 GiB RSS; the 5,000-file folder and 10,000 Analytics entries were visible, while the remaining 995,000 files and Notes still reported zero. This is a point-in-time observation, not a peak or a completed before/after result. The local host load average was 33.72, 31.87, 29.68. The one-hour run remains in progress; I will report its final result and limitation separately.
Author
Owner

Follow-up to the provisional local profile note: /proc/229648/status reports VmHWM: 1510240 kB (about 1.44 GiB peak RSS so far) at 39 minutes. VmRSS was 1429760 kB at the same sample. Counts remain unchanged: 5,000 folder rows and 10,000 Analytics entries; zero for remaining Files and Notes. This is a peak-so-far snapshot under heavy local host load, not the completed profile peak.

Follow-up to the provisional local profile note: `/proc/229648/status` reports `VmHWM: 1510240 kB` (about 1.44 GiB peak RSS so far) at 39 minutes. `VmRSS` was 1429760 kB at the same sample. Counts remain unchanged: 5,000 folder rows and 10,000 Analytics entries; zero for remaining Files and Notes. This is a peak-so-far snapshot under heavy local host load, not the completed profile peak.
Author
Owner

Local profile high-water update: at 51 minutes, /proc/229648/status reports VmHWM: 1658620 kB (about 1.58 GiB peak RSS so far), up from the 39-minute sample. VmRSS was 1440816 kB. Index counters remain at 5,000 folder rows, 10,000 Analytics entries, and zero for the remaining Files and Notes. The profile is still incomplete.

Local profile high-water update: at 51 minutes, `/proc/229648/status` reports `VmHWM: 1658620 kB` (about 1.58 GiB peak RSS so far), up from the 39-minute sample. `VmRSS` was 1440816 kB. Index counters remain at 5,000 folder rows, 10,000 Analytics entries, and zero for the remaining Files and Notes. The profile is still incomplete.
Author
Owner

#503 implementation is committed. Head SHA: 94d5d08658175bc7b8d25500fdcf7750d63523de.

Built

  • Bounded SQLite pools to 8 reader connections, 1 MiB page cache per connection, and 32 cached statements. Added regression coverage for reader clamping and cache pragmas.
  • Reworked semantic Home reconciliation to page directory entries into a bounded queue and fixed batches. It records paths in SQLite and removes stale rows with SQL anti-joins instead of Home-sized Rust collections.
  • Made semantic and CLIP ONNX sessions lazy and shared. Deleting vectors does not initialize a model.
  • Added an opt-in PERF_HOME_SKIP_PHOTO_WAIT=1 large-Home profile for one concurrent Photos rebuild while measuring Files and Notes.

Gates

  • cargo fmt --all --check: exit 0, no output.
  • cargo clippy -p calternal-db --all-targets -- -D warnings: passed, exit 0.
  • cargo test -p calternal-db: passed; 10 unit tests, queue integration 16 passed / 1 ignored, sqlite_limits 1 passed, doc tests 0.
  • cargo clippy -p calternal-embed --all-targets -- -D warnings: passed. Output:
        Checking calternal-embed v0.0.1 (/home/kayg/Developer/calternal-wt/fix-503/crates/calternal-embed)
        Finished `dev` profile [unoptimized + debuginfo] target(s) in 6.10s
    
  • cargo test -p calternal-embed:
    test result: ok. 36 passed; 0 failed; 4 ignored; 0 measured; 0 filtered out; finished in 35.52s
    
       Doc-tests calternal_embed
    
    running 0 tests
    
    test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s
    
  • bun run check: svelte-check found 0 errors and 0 warnings.
  • bun run test: Test Files 140 passed (140) and Tests 930 passed (930).
  • node --check apps/web/e2e/photos-perf.mjs, bash -n bench/run.sh, and git diff --check: exit 0, no output.

Performance evidence and gap

The baseline in docs/perf/2026-09-30.md reports 4.20 GB and 6.65 GB peak RSS for the 1M-file Home; Files reached 995k at about 1,330 seconds, Notes at about 1,520 seconds, and Photos did not finish.

The perf VM was unavailable: flock -n /root/perf.lock exited 1 because the lock was busy. I ran one local fallback with the same fixture. At server runtime 60:18 it still reported 5,000 folder rows and 10,000 Analytics entries, with 0 of the remaining 995,000 Files and 0 Notes indexed. The server's kernel high-water RSS was 1,658,620 kB (about 1.58 GiB) at the last sample. Local host load was about 30–34. This is an incomplete, loaded-host observation, not a comparable after peak. The browser remained stuck on an in-flight fetch past the indexing deadline; I stopped the benchmark server, after which its sampler logged /proc ENOENT errors. No after-run JSON was produced, and I did not update the performance baseline with this partial result.

Search rebuild map streaming remains with #496. Its current patch conflicts with the per-User isolation and Sidecar behavior in this branch, so I did not duplicate that change.

Decisions outside DESIGN

  • Set the SQLite reader, page-cache, and statement-cache limits above.
  • Store the per-pass seen paths in each User's SQLite file and prune stale semantic rows with anti-joins.
  • Initialize one shared inference session on first use; inference stays behind the existing model mutex.
  • Make the Photos-wait bypass an opt-in benchmark flag because the baseline Photos queue did not finish.

Files: crates/calternal-db/src/db.rs, crates/calternal-db/src/lib.rs, crates/calternal-db/tests/sqlite_limits.rs, crates/calternal-embed/Cargo.toml, crates/calternal-embed/src/store.rs, crates/calternal-embed/src/model.rs, crates/calternal-embed/src/clip_store.rs, Cargo.lock, apps/web/e2e/photos-perf.mjs, and bench/run.sh.

#503 implementation is committed. Head SHA: `94d5d08658175bc7b8d25500fdcf7750d63523de`. ## Built - Bounded SQLite pools to 8 reader connections, 1 MiB page cache per connection, and 32 cached statements. Added regression coverage for reader clamping and cache pragmas. - Reworked semantic Home reconciliation to page directory entries into a bounded queue and fixed batches. It records paths in SQLite and removes stale rows with SQL anti-joins instead of Home-sized Rust collections. - Made semantic and CLIP ONNX sessions lazy and shared. Deleting vectors does not initialize a model. - Added an opt-in `PERF_HOME_SKIP_PHOTO_WAIT=1` large-Home profile for one concurrent Photos rebuild while measuring Files and Notes. ## Gates - `cargo fmt --all --check`: exit 0, no output. - `cargo clippy -p calternal-db --all-targets -- -D warnings`: passed, exit 0. - `cargo test -p calternal-db`: passed; 10 unit tests, queue integration 16 passed / 1 ignored, `sqlite_limits` 1 passed, doc tests 0. - `cargo clippy -p calternal-embed --all-targets -- -D warnings`: passed. Output: ``` Checking calternal-embed v0.0.1 (/home/kayg/Developer/calternal-wt/fix-503/crates/calternal-embed) Finished `dev` profile [unoptimized + debuginfo] target(s) in 6.10s ``` - `cargo test -p calternal-embed`: ``` test result: ok. 36 passed; 0 failed; 4 ignored; 0 measured; 0 filtered out; finished in 35.52s Doc-tests calternal_embed running 0 tests test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s ``` - `bun run check`: `svelte-check found 0 errors and 0 warnings`. - `bun run test`: `Test Files 140 passed (140)` and `Tests 930 passed (930)`. - `node --check apps/web/e2e/photos-perf.mjs`, `bash -n bench/run.sh`, and `git diff --check`: exit 0, no output. ## Performance evidence and gap The baseline in `docs/perf/2026-09-30.md` reports 4.20 GB and 6.65 GB peak RSS for the 1M-file Home; Files reached 995k at about 1,330 seconds, Notes at about 1,520 seconds, and Photos did not finish. The perf VM was unavailable: `flock -n /root/perf.lock` exited 1 because the lock was busy. I ran one local fallback with the same fixture. At server runtime 60:18 it still reported 5,000 folder rows and 10,000 Analytics entries, with 0 of the remaining 995,000 Files and 0 Notes indexed. The server's kernel high-water RSS was 1,658,620 kB (about 1.58 GiB) at the last sample. Local host load was about 30–34. This is an incomplete, loaded-host observation, not a comparable after peak. The browser remained stuck on an in-flight fetch past the indexing deadline; I stopped the benchmark server, after which its sampler logged `/proc` ENOENT errors. No after-run JSON was produced, and I did not update the performance baseline with this partial result. Search rebuild map streaming remains with #496. Its current patch conflicts with the per-User isolation and Sidecar behavior in this branch, so I did not duplicate that change. ## Decisions outside DESIGN - Set the SQLite reader, page-cache, and statement-cache limits above. - Store the per-pass seen paths in each User's SQLite file and prune stale semantic rows with anti-joins. - Initialize one shared inference session on first use; inference stays behind the existing model mutex. - Make the Photos-wait bypass an opt-in benchmark flag because the baseline Photos queue did not finish. Files: `crates/calternal-db/src/db.rs`, `crates/calternal-db/src/lib.rs`, `crates/calternal-db/tests/sqlite_limits.rs`, `crates/calternal-embed/Cargo.toml`, `crates/calternal-embed/src/store.rs`, `crates/calternal-embed/src/model.rs`, `crates/calternal-embed/src/clip_store.rs`, `Cargo.lock`, `apps/web/e2e/photos-perf.mjs`, and `bench/run.sh`.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#503
No description provided.