Benchmark: ort (ONNX Runtime, C++) vs candle (pure Rust) for semantic search embeddings #74

Closed
opened 2026-09-24 19:40:18 +00:00 by kayg · 56 comments
Owner

Owner decision needed on data, not opinion: should semantic search (#59, crates/calternal-embed) keep ort (ONNX Runtime, a C++ library linked into the server, which needs the C++ cross toolchain and libstdc++) or switch to candle (Hugging Face, pure Rust)? The owner said the benchmark results matter. Measure carefully and report raw numbers; do not argue from blog posts.

What to build

  • A standalone bench crate outside the workspace at bench/embed-backends/, with its own Cargo.toml, so it doesn't touch the server. It has two binaries sharing one harness: bench-ort (the exact ort version, features and quantized model assets calternal-embed uses today; read crates/calternal-embed/src/model.rs and the model manifest) and bench-candle (candle-core + candle-nn + candle-transformers BERT, same model all-MiniLM-L6-v2, same tokenizer via the tokenizers crate, mean pooling + L2 normalize exactly like calternal-embed).
  • Run both backends in f32 (the apples-to-apples row) and in each backend's best CPU option (ort: the quantized model calternal ships; candle: its best available, e.g. f16/bf16 or quantized if a BERT path supports it; say what you used). Record model files and SHA-256.

Corpus (fixed, committed, deterministic)

  • bench/embed-backends/corpus/: 2,000 texts generated deterministically (fixed seed) to look like calternal data. 40% short log entries (5–15 words), 40% note paragraphs (60–200 words), 20% long chunks at the chunker's max token length. Use the chunker from calternal-embed to decide lengths. Plus 200 queries (2–8 words) and a 50-pair relevance set (query → the one text it should find, cross-vocabulary like "flat search in berlin" → "apartment hunting").

Measurements (each backend, each precision)

  1. Query latency: single query, batch 1, after 20 warm-up runs. 1,000 runs. Report p50/p95/p99/max in microseconds. Separately report the cold first query after process start.
  2. Indexing throughput: embed the 2,000 texts in batches of 1, 8 and 32; report texts/second and total wall time.
  3. Model load time: process start to ready, median of 10 fresh processes.
  4. Memory: peak RSS (/usr/bin/time -v) for the query run and for the batch-32 run.
  5. Threads: run with 1 thread and with 4 threads (set each backend's intra-op threads explicitly; state how).
  6. Output parity: cosine similarity between ort-f32 and candle-f32 embeddings on all 2,000 texts: min / mean. Must be ≥ 0.999 min, or explain the difference.
  7. Quality: recall@1 and recall@10 on the 50-pair set, per backend and precision.
  8. Build and deploy cost: clean release build time for x86_64 and a clean cross build for aarch64-unknown-linux-gnu (linker aarch64-linux-gnu-gcc; ort needs CXX_aarch64_unknown_linux_gnu=aarch64-linux-gnu-g++); stripped binary size of each bench binary; extra system libraries needed at runtime (ldd/readelf -d).

Noise control (this machine is shared and loaded, load average ~30 on 8 cores)

  • Pin every run with taskset -c 6,7 (4-thread runs: taskset -c 4-7) and nice -n -5 if permitted, otherwise plain.
  • Interleave: run A, B, A, B … for 5 rounds per measurement, not all-A then all-B, and report the median of the 5 round medians plus the spread (min–max of round medians). Record uptime load averages before and after each round.
  • If the spread of round medians exceeds 20%, rerun that measurement later in the night when load is lower; say which numbers are noisy.

Also on the real target (arm64)

o2 is the production host (arm64, the deploy target). Cross-build both bench binaries, copy them with the model assets to o2 via netbird ssh --no-browser --user kayg 10.69.69.52 (use cat > file over ssh; put everything under ~/bench-embed/, delete it when done) and run measurements 1, 3, 4 and 6 there with 1 and 4 threads, nice -n 10, the same interleaving. Do not touch any running service or container on o2. Keep it under 15 minutes of CPU there.

Report

Commit bench/embed-backends/RESULTS.md with one table per host (x86 VM, o2 arm64). Columns: backend, precision, threads, query p50/p95/p99, cold query, load time, texts/s at batch 1/8/32, peak RSS, recall@1/@10, binary size, cross-build notes. Include the exact commands, CPU model (lscpu), load averages, crate versions and git SHA. End with a recommendation using this rule:

  • Prefer candle if its best-option query p95 on o2 is within 1.5× of ort's, recall@10 is within 2 points, and indexing throughput is within 2×. Otherwise prefer ort.
  • State the rule's verdict plainly, even if you'd argue otherwise, and then add at most three sentences of judgment.
    Do not change crates/calternal-embed in this job; the switch, if any, is a follow-up.
Owner decision needed on data, not opinion: should semantic search (#59, `crates/calternal-embed`) keep **`ort`** (ONNX Runtime, a C++ library linked into the server, which needs the C++ cross toolchain and libstdc++) or switch to **`candle`** (Hugging Face, pure Rust)? The owner said the benchmark results matter. Measure carefully and report raw numbers; do not argue from blog posts. ## What to build - A standalone bench crate **outside the workspace** at `bench/embed-backends/`, with its own `Cargo.toml`, so it doesn't touch the server. It has two binaries sharing one harness: `bench-ort` (the exact `ort` version, features and quantized model assets `calternal-embed` uses today; read `crates/calternal-embed/src/model.rs` and the model manifest) and `bench-candle` (`candle-core` + `candle-nn` + `candle-transformers` BERT, same model **all-MiniLM-L6-v2**, same tokenizer via the `tokenizers` crate, mean pooling + L2 normalize exactly like calternal-embed). - Run both backends in **f32** (the apples-to-apples row) and in **each backend's best CPU option** (ort: the quantized model calternal ships; candle: its best available, e.g. f16/bf16 or quantized if a BERT path supports it; say what you used). Record model files and SHA-256. ## Corpus (fixed, committed, deterministic) - `bench/embed-backends/corpus/`: 2,000 texts generated deterministically (fixed seed) to look like calternal data. 40% short log entries (5–15 words), 40% note paragraphs (60–200 words), 20% long chunks at the chunker's max token length. Use the chunker from `calternal-embed` to decide lengths. Plus 200 queries (2–8 words) and a 50-pair relevance set (query → the one text it should find, cross-vocabulary like "flat search in berlin" → "apartment hunting"). ## Measurements (each backend, each precision) 1. **Query latency**: single query, batch 1, after 20 warm-up runs. 1,000 runs. Report p50/p95/p99/max in microseconds. Separately report the **cold first query** after process start. 2. **Indexing throughput**: embed the 2,000 texts in batches of 1, 8 and 32; report texts/second and total wall time. 3. **Model load time**: process start to ready, median of 10 fresh processes. 4. **Memory**: peak RSS (`/usr/bin/time -v`) for the query run and for the batch-32 run. 5. **Threads**: run with 1 thread and with 4 threads (set each backend's intra-op threads explicitly; state how). 6. **Output parity**: cosine similarity between ort-f32 and candle-f32 embeddings on all 2,000 texts: min / mean. Must be ≥ 0.999 min, or explain the difference. 7. **Quality**: recall@1 and recall@10 on the 50-pair set, per backend and precision. 8. **Build and deploy cost**: clean release build time for x86_64 and a clean cross build for `aarch64-unknown-linux-gnu` (linker `aarch64-linux-gnu-gcc`; ort needs `CXX_aarch64_unknown_linux_gnu=aarch64-linux-gnu-g++`); stripped binary size of each bench binary; extra system libraries needed at runtime (`ldd`/`readelf -d`). ## Noise control (this machine is shared and loaded, load average ~30 on 8 cores) - Pin every run with `taskset -c 6,7` (4-thread runs: `taskset -c 4-7`) and `nice -n -5` if permitted, otherwise plain. - **Interleave**: run A, B, A, B … for 5 rounds per measurement, not all-A then all-B, and report the median of the 5 round medians plus the spread (min–max of round medians). Record `uptime` load averages before and after each round. - If the spread of round medians exceeds 20%, rerun that measurement later in the night when load is lower; say which numbers are noisy. ## Also on the real target (arm64) o2 is the production host (arm64, the deploy target). Cross-build both bench binaries, copy them with the model assets to o2 via `netbird ssh --no-browser --user kayg 10.69.69.52` (use `cat > file` over ssh; put everything under `~/bench-embed/`, delete it when done) and run measurements 1, 3, 4 and 6 there with 1 and 4 threads, `nice -n 10`, the same interleaving. Do not touch any running service or container on o2. Keep it under 15 minutes of CPU there. ## Report Commit `bench/embed-backends/RESULTS.md` with one table per host (x86 VM, o2 arm64). Columns: backend, precision, threads, query p50/p95/p99, cold query, load time, texts/s at batch 1/8/32, peak RSS, recall@1/@10, binary size, cross-build notes. Include the exact commands, CPU model (`lscpu`), load averages, crate versions and git SHA. End with a **recommendation using this rule**: - Prefer candle if its best-option query p95 on o2 is within 1.5× of ort's, recall@10 is within 2 points, and indexing throughput is within 2×. Otherwise prefer ort. - State the rule's verdict plainly, even if you'd argue otherwise, and then add at most three sentences of judgment. Do not change `crates/calternal-embed` in this job; the switch, if any, is a follow-up.
Author
Owner

Starting issue #74 in worktree branch job/embed-bench at base SHA 8c216759f01d704e69ecf37ef8293f422abf5552. I have read the issue; it currently has no comments. I am reading the binding repository docs and will follow the measurement and commit order in the issue.

Starting issue #74 in worktree branch `job/embed-bench` at base SHA `8c216759f01d704e69ecf37ef8293f422abf5552`. I have read the issue; it currently has no comments. I am reading the binding repository docs and will follow the measurement and commit order in the issue.
Author
Owner

Finding: the production manifest pins all-MiniLM-L6-v2 revision 1110a243fdf4706b3f48f1d95db1a4f5529b4d41, max length 256, and tokenizer SHA-256 be50c3628f2bf5bb5e3a7f17b1f74611b2561a3a27eeab05e5aa30f411572037. It selects ONNX SHA-256 b941bf19f1f1283680f449fa6a7336bb5600bdcd5f84d10ddc5cd72218a0fd21 on x86 AVX2 and 4278337fd0ff3c68bfb6291042cad8ab363e1d9fbc43dcb499fe91c871902474 on arm64; the unquantized fallback is 6fd5d72fe4589f189f8ebc006442dbb529bb7ce38f8082112682524616046452. ort resolves to 2.0.0-rc.13 with its default features, matching the production workspace dependency. Cargo reports Candle core, nn, and transformers 0.11.0; the Transformers crate includes a BERT model but no quantized BERT implementation, so I am evaluating CPU f16 as Candle's best-option candidate alongside f32.

Finding: the production manifest pins all-MiniLM-L6-v2 revision `1110a243fdf4706b3f48f1d95db1a4f5529b4d41`, max length 256, and tokenizer SHA-256 `be50c3628f2bf5bb5e3a7f17b1f74611b2561a3a27eeab05e5aa30f411572037`. It selects ONNX SHA-256 `b941bf19f1f1283680f449fa6a7336bb5600bdcd5f84d10ddc5cd72218a0fd21` on x86 AVX2 and `4278337fd0ff3c68bfb6291042cad8ab363e1d9fbc43dcb499fe91c871902474` on arm64; the unquantized fallback is `6fd5d72fe4589f189f8ebc006442dbb529bb7ce38f8082112682524616046452`. `ort` resolves to 2.0.0-rc.13 with its default features, matching the production workspace dependency. Cargo reports Candle core, nn, and transformers 0.11.0; the Transformers crate includes a BERT model but no quantized BERT implementation, so I am evaluating CPU f16 as Candle's best-option candidate alongside f32.
Author
Owner

Finding: I downloaded each file from the immutable production revision and verified it with sha256sum. The tokenizer and all three ONNX hashes match the manifest exactly. The same-revision Candle safetensors file hashes to 53aa51172d142c89d9012cce15ae4d6cc0ca6895895114379cacb4fab128d9db; its config.json hashes to 953f9c0d463486b10a6871cc2fd59f223b2c70184f49815e7efbcab5d8908b41. /usr/bin/time is absent on the x86 VM. GNU time 1.10 is available from Nix at /nix/store/23s9fdv7c4fai0bygx433a87arrqkykf-time-1.10/bin/time, and time -v true returned its verbose RSS and CPU fields. I will use that GNU time binary for equivalent verbose accounting and record the command in raw results.

Finding: I downloaded each file from the immutable production revision and verified it with `sha256sum`. The tokenizer and all three ONNX hashes match the manifest exactly. The same-revision Candle safetensors file hashes to `53aa51172d142c89d9012cce15ae4d6cc0ca6895895114379cacb4fab128d9db`; its `config.json` hashes to `953f9c0d463486b10a6871cc2fd59f223b2c70184f49815e7efbcab5d8908b41`. `/usr/bin/time` is absent on the x86 VM. GNU time 1.10 is available from Nix at `/nix/store/23s9fdv7c4fai0bygx433a87arrqkykf-time-1.10/bin/time`, and `time -v true` returned its verbose RSS and CPU fields. I will use that GNU time binary for equivalent verbose accounting and record the command in raw results.
Author
Owner

Finding: o2 is a 4-core Neoverse-N1 system (CPUs 0–3), load average 0.39/0.52/0.66 when inspected. The x86 pin masks 6,7 and 4–7 do not exist there. I will pin one-thread o2 runs to CPUs 0–1 and four-thread runs to CPUs 0–3, while retaining nice -n 10 and the required interleaving.

Finding: o2 is a 4-core Neoverse-N1 system (CPUs 0–3), load average 0.39/0.52/0.66 when inspected. The x86 pin masks 6,7 and 4–7 do not exist there. I will pin one-thread o2 runs to CPUs 0–1 and four-thread runs to CPUs 0–3, while retaining `nice -n 10` and the required interleaving.
Author
Owner

Finding: the pinned tokenizer JSON has fixed padding length 128 and an embedded truncation length 128. Production model.rs overrides truncation to 256 but leaves fixed padding in place. A direct encode("a", false) therefore returns 128 IDs with only one active attention-mask token. The corpus generator must count active attention-mask entries rather than ID-array length, and the benchmark must preserve the tokenizer's fixed padding to match production. I am correcting this before corpus generation.

Finding: the pinned tokenizer JSON has fixed padding length 128 and an embedded truncation length 128. Production `model.rs` overrides truncation to 256 but leaves fixed padding in place. A direct `encode("a", false)` therefore returns 128 IDs with only one active attention-mask token. The corpus generator must count active attention-mask entries rather than ID-array length, and the benchmark must preserve the tokenizer's fixed padding to match production. I am correcting this before corpus generation.
Author
Owner

Harness and measurement runner are committed on job/embed-bench.

  • Harness: f66610d
  • Runner/build support: 00cf55f25e35d53f1fc705ab63b648da4beb95b1
  • Standalone crate gates pass: cargo fmt --check; all-feature tests (7 passed); all-target/all-feature Clippy with -D warnings.

The O2 runner now checkpoints each backend run and applies a 120 CPU-second process limit. It stops launching at 720 accounted seconds, with a 900-second aggregate ceiling. The ARM matrix follows requested measurements 1, 3, 4, and 6; indexing is not part of the O2 run.

Harness and measurement runner are committed on `job/embed-bench`. - Harness: `f66610d` - Runner/build support: `00cf55f25e35d53f1fc705ab63b648da4beb95b1` - Standalone crate gates pass: `cargo fmt --check`; all-feature tests (7 passed); all-target/all-feature Clippy with `-D warnings`. The O2 runner now checkpoints each backend run and applies a 120 CPU-second process limit. It stops launching at 720 accounted seconds, with a 900-second aggregate ceiling. The ARM matrix follows requested measurements 1, 3, 4, and 6; indexing is not part of the O2 run.
Author
Owner

Finding from the first clean-build pass: this host sets RUSTC_WRAPPER=/home/kayg/.nix-profile/bin/sccache. A cargo clean therefore did not ensure an uncached release build. The first x86 ORT run completed in 120.325896790 s, but I discarded that pass and stopped the Candle run before treating either timing as final. The build helper now sets SCCACHE_DISABLE=1 and records that setting; I am rerunning both native builds from clean targets.

Finding from the first clean-build pass: this host sets `RUSTC_WRAPPER=/home/kayg/.nix-profile/bin/sccache`. A `cargo clean` therefore did not ensure an uncached release build. The first x86 ORT run completed in 120.325896790 s, but I discarded that pass and stopped the Candle run before treating either timing as final. The build helper now sets `SCCACHE_DISABLE=1` and records that setting; I am rerunning both native builds from clean targets.
Author
Owner

Correction to the build-cache finding: setting SCCACHE_DISABLE=1 alone still spawned sccache client processes (observed multiple sccache rustc children with 0% CPU while the Cargo build waited). The updated helper removes RUSTC_WRAPPER and CARGO_BUILD_RUSTC_WRAPPER from the Cargo environment so it runs rustc directly. I stopped the two contaminated attempts and am discarding their timings.

Correction to the build-cache finding: setting `SCCACHE_DISABLE=1` alone still spawned sccache client processes (observed multiple `sccache rustc` children with 0% CPU while the Cargo build waited). The updated helper removes `RUSTC_WRAPPER` and `CARGO_BUILD_RUSTC_WRAPPER` from the Cargo environment so it runs rustc directly. I stopped the two contaminated attempts and am discarding their timings.
Author
Owner

Native x86 clean release builds are recorded in commit 88a83d66fcd0899f893bcc1cfeab5be14072c375 (raw/x86-build.json). With RUSTC_WRAPPER removed, ORT built in 98.999808024 s and Candle in 262.134217217 s. Stripped sizes are 27,059,424 bytes and 6,809,488 bytes. ldd shows ORT has no separate libonnxruntime.so dependency; it does need the host C++ runtime (libstdc++.so.6).

Native x86 clean release builds are recorded in commit `88a83d66fcd0899f893bcc1cfeab5be14072c375` (`raw/x86-build.json`). With `RUSTC_WRAPPER` removed, ORT built in 98.999808024 s and Candle in 262.134217217 s. Stripped sizes are 27,059,424 bytes and 6,809,488 bytes. `ldd` shows ORT has no separate `libonnxruntime.so` dependency; it does need the host C++ runtime (`libstdc++.so.6`).
Author
Owner

Candle precision finding, recorded in raw/x86-smoke.json (commit a0fb11a7b7bcb8c4312f96c28c056afa2bc8b35a): the release f16 candidate fails its first query with norm=NaN, finite_dimensions=0/384, and token_count=6. A BF16 CPU query fails with unsupported dtype BF16 for op matmul. Candle 0.11 has no quantized BERT path, so the only usable Candle CPU precision found is f32. The runner now compares that f32 path in both the baseline f32 pair and the interleaved best-CPU pair with ORT's host-specific quantized model.

Candle precision finding, recorded in `raw/x86-smoke.json` (commit `a0fb11a7b7bcb8c4312f96c28c056afa2bc8b35a`): the release f16 candidate fails its first query with `norm=NaN`, `finite_dimensions=0/384`, and `token_count=6`. A BF16 CPU query fails with `unsupported dtype BF16 for op matmul`. Candle 0.11 has no quantized BERT path, so the only usable Candle CPU precision found is f32. The runner now compares that f32 path in both the baseline f32 pair and the interleaved best-CPU pair with ORT's host-specific quantized model.
Author
Owner

Final clean build pass is recorded in commit 01ebd6c. Source revision: 2b0d0aefecc297a9a2f02e740b7869e573c57b1f. Results from raw/x86-build.json: x86_64 ORT 89.965430274 s / 27,063,072 bytes; x86_64 Candle 186.709400095 s / 6,808,920 bytes; aarch64 cross ORT 81.733230656 s / 27,082,760 bytes; aarch64 cross Candle 199.321161940 s / 5,446,024 bytes. All four builds exited 0. I am starting the x86 query/index/runtime measurement matrix now.

Final clean build pass is recorded in commit `01ebd6c`. Source revision: `2b0d0aefecc297a9a2f02e740b7869e573c57b1f`. Results from `raw/x86-build.json`: x86_64 ORT 89.965430274 s / 27,063,072 bytes; x86_64 Candle 186.709400095 s / 6,808,920 bytes; aarch64 cross ORT 81.733230656 s / 27,082,760 bytes; aarch64 cross Candle 199.321161940 s / 5,446,024 bytes. All four builds exited 0. I am starting the x86 query/index/runtime measurement matrix now.
Author
Owner

The first x86 query group (F32, 1 thread, 5 paired rounds) is complete. raw/x86-vm.json records p50 medians of 37.916634 ms for ORT F32 and 81.355545 ms for Candle F32. The per-round spread exceeded 20% for ORT p99 (20.16%) and max (42.88%), and Candle p95 (23.73%), p99 (31.31%), and max (124.47%). The runner will consider this group for the protocol’s lower-load nighttime recheck after the primary matrix finishes.

The first x86 query group (F32, 1 thread, 5 paired rounds) is complete. `raw/x86-vm.json` records p50 medians of 37.916634 ms for ORT F32 and 81.355545 ms for Candle F32. The per-round spread exceeded 20% for ORT p99 (20.16%) and max (42.88%), and Candle p95 (23.73%), p99 (31.31%), and max (124.47%). The runner will consider this group for the protocol’s lower-load nighttime recheck after the primary matrix finishes.
Author
Owner

Finding from the live O2 ARM64 raw file: in the first F32 query round at one thread, Candle reached the per-process 120 CPU second cap and exited 137 (cpu_seconds=120.44); ORT F32 completed in 91.10 CPU seconds. The O2 runner is charging actual child CPU and remains bounded by the 900 second host limit. I am letting it finish and will record any protocol steps that the cap prevents.

Finding from the live O2 ARM64 raw file: in the first F32 query round at one thread, Candle reached the per-process 120 CPU second cap and exited 137 (`cpu_seconds=120.44`); ORT F32 completed in 91.10 CPU seconds. The O2 runner is charging actual child CPU and remains bounded by the 900 second host limit. I am letting it finish and will record any protocol steps that the cap prevents.
Author
Owner

O2 ARM64 primary-run raw data is committed in 73a36e5. It used 726.46 child CPU seconds (726.48 including the runner) and stopped at the configured launch threshold with status partial_cpu_budget. It records three complete one-thread F32 query rounds plus the next ORT side: ORT F32 was 91.10, 91.03, and 91.31 CPU seconds; Candle F32 hit the 120 CPU second cap in all three rounds (exit 137). The fourth ORT F32 process was saved at 91.13 seconds. About 174 CPU seconds remain before the 900-second host ceiling; the uncompleted ARM measurements are not inferred.

O2 ARM64 primary-run raw data is committed in `73a36e5`. It used 726.46 child CPU seconds (726.48 including the runner) and stopped at the configured launch threshold with status `partial_cpu_budget`. It records three complete one-thread F32 query rounds plus the next ORT side: ORT F32 was 91.10, 91.03, and 91.31 CPU seconds; Candle F32 hit the 120 CPU second cap in all three rounds (exit 137). The fourth ORT F32 process was saved at 91.13 seconds. About 174 CPU seconds remain before the 900-second host ceiling; the uncompleted ARM measurements are not inferred.
Author
Owner

The O2 continuation is in ee7f616. It completed all model-load samples at 1 and 4 threads. Medians (process start to ready): ORT F32 218.988/218.198 ms; ORT qint8 164.452/170.706 ms; Candle F32 66.243/66.971 ms (1/4 threads). It attempted the batch-32 RSS workload: ORT F32 hit the 101.95 CPU-second process limit and exited 137 at 506,496 KiB RSS. The cumulative O2 total is 839.48 CPU seconds including the runner, under the 900-second limit; the reserve stopped further workloads before F32 output parity. No ARM indexing throughput, Candle batch-32 RSS, or parity number is available.

The O2 continuation is in `ee7f616`. It completed all model-load samples at 1 and 4 threads. Medians (process start to ready): ORT F32 218.988/218.198 ms; ORT qint8 164.452/170.706 ms; Candle F32 66.243/66.971 ms (1/4 threads). It attempted the batch-32 RSS workload: ORT F32 hit the 101.95 CPU-second process limit and exited 137 at 506,496 KiB RSS. The cumulative O2 total is 839.48 CPU seconds including the runner, under the 900-second limit; the reserve stopped further workloads before F32 output parity. No ARM indexing throughput, Candle batch-32 RSS, or parity number is available.
Author
Owner

The x86 F32, one-thread, batch-1 index group has completed and is very noisy. From its five measured round medians, ORT elapsed time median is 126777182331 ns (min 105686618304, max 353372403328; spread 195.37094962192973%). Candle median is 294466264647 ns (min 230178838407, max 718921832170; spread 165.9758866941498%). The runner will consider both for a later lower-load nighttime recheck; the group and its per-round load averages are in raw/x86-vm.json.

The x86 F32, one-thread, batch-1 index group has completed and is very noisy. From its five measured round medians, ORT elapsed time median is 126777182331 ns (min 105686618304, max 353372403328; spread 195.37094962192973%). Candle median is 294466264647 ns (min 230178838407, max 718921832170; spread 165.9758866941498%). The runner will consider both for a later lower-load nighttime recheck; the group and its per-round load averages are in `raw/x86-vm.json`.
Author
Owner

The x86 F32, one-thread, batch-8 index group has completed. Both backends exceed the 20% spread threshold across five rounds: ORT median elapsed time 455194756325 ns (spread 92.95905427497567%); Candle median 707694999796 ns (spread 55.758075977468614%). Load averages vary in the attached raw rounds. The runner will consider these measurements for a later lower-load night recheck.

The x86 F32, one-thread, batch-8 index group has completed. Both backends exceed the 20% spread threshold across five rounds: ORT median elapsed time 455194756325 ns (spread 92.95905427497567%); Candle median 707694999796 ns (spread 55.758075977468614%). Load averages vary in the attached raw rounds. The runner will consider these measurements for a later lower-load night recheck.
Author
Owner

The x86 F32, one-thread, batch-32 index group is complete. Both five-round elapsed-time spreads exceed 20%: ORT median 192474471121 ns (spread 123.55747674635533%); Candle median 474083626079 ns (spread 133.70964617229143%). The paired rounds record the varying host load; these groups are flagged for a lower-load nighttime recheck.

The x86 F32, one-thread, batch-32 index group is complete. Both five-round elapsed-time spreads exceed 20%: ORT median 192474471121 ns (spread 123.55747674635533%); Candle median 474083626079 ns (spread 133.70964617229143%). The paired rounds record the varying host load; these groups are flagged for a lower-load nighttime recheck.
Author
Owner

The x86 best-CPU query group at one thread is complete and exceeds the 20% spread threshold for both p95 series. ORT AVX2 quantized p95 median is 220245336 ns (spread 88.92835896420527%); Candle F32 p95 median is 309531774 ns (spread 126.04942069695242%). These results are flagged for lower-load nighttime rechecks. The raw rows also retain p50/p99 and cold first-query values.

The x86 best-CPU query group at one thread is complete and exceeds the 20% spread threshold for both p95 series. ORT AVX2 quantized p95 median is 220245336 ns (spread 88.92835896420527%); Candle F32 p95 median is 309531774 ns (spread 126.04942069695242%). These results are flagged for lower-load nighttime rechecks. The raw rows also retain p50/p99 and cold first-query values.
Author
Owner

The x86 best-CPU, one-thread, batch-1 index group has completed. ORT AVX2 quantized median elapsed time is 120332266408 ns with a 29.557883806828535% spread, exceeding the 20% threshold. Candle F32 median is 306799430903 ns with a 13.080809241686106% spread, below threshold. The ORT group is flagged for a lower-load recheck.

The x86 best-CPU, one-thread, batch-1 index group has completed. ORT AVX2 quantized median elapsed time is 120332266408 ns with a 29.557883806828535% spread, exceeding the 20% threshold. Candle F32 median is 306799430903 ns with a 13.080809241686106% spread, below threshold. The ORT group is flagged for a lower-load recheck.
Author
Owner

The x86 best-CPU, one-thread, batch-8 index group is complete. Both backends exceed the 20% spread threshold across five rounds: ORT AVX2 quantized median elapsed time is 137061616859 ns (spread 152.26552535980542%); Candle F32 median is 531153462959 ns (spread 55.28988852241915%). Both are flagged for the lower-load night recheck.

The x86 best-CPU, one-thread, batch-8 index group is complete. Both backends exceed the 20% spread threshold across five rounds: ORT AVX2 quantized median elapsed time is 137061616859 ns (spread 152.26552535980542%); Candle F32 median is 531153462959 ns (spread 55.28988852241915%). Both are flagged for the lower-load night recheck.
Author
Owner

Completed the x86-vm one-thread best-CPU batch-32 five-round interleaved set. Both backends exceed the 20% spread threshold: ORT quantized elapsed medians were [249851714577, 205684068186, 140736146459, 192928617955, 168050812587] ns, with 56.557481867957435% max/min spread; Candle f32 medians were [518037404332, 413834384949, 319570430720, 476553047356, 773599365451] ns, with 95.27353507653183% spread. The raw rounds are committed in bench/embed-backends/raw/x86-vm.json at 0d69d24. I will evaluate this set for a lower-load nighttime recheck after the primary matrix completes.

Completed the x86-vm one-thread best-CPU batch-32 five-round interleaved set. Both backends exceed the 20% spread threshold: ORT quantized elapsed medians were [249851714577, 205684068186, 140736146459, 192928617955, 168050812587] ns, with 56.557481867957435% max/min spread; Candle f32 medians were [518037404332, 413834384949, 319570430720, 476553047356, 773599365451] ns, with 95.27353507653183% spread. The raw rounds are committed in `bench/embed-backends/raw/x86-vm.json` at 0d69d24. I will evaluate this set for a lower-load nighttime recheck after the primary matrix completes.
Author
Owner

Completed the five-round x86-vm four-thread f32 query set. Query p95 round medians were ORT [231084569, 296962537, 180969080, 205790865, 161047634] ns (median 205790865 ns; max/min spread 66.04515851566103%) and Candle [388648761, 520355316, 401643453, 352389150, 360770138] ns (median 388648761 ns; spread 43.21798571229717%). p50, p99 and max also exceeded the 20% threshold for both backends. The raw rounds are committed in bench/embed-backends/raw/x86-vm.json at 5a2d1a2; they need a lower-load nighttime recheck after the primary matrix completes.

Completed the five-round x86-vm four-thread f32 query set. Query p95 round medians were ORT [231084569, 296962537, 180969080, 205790865, 161047634] ns (median 205790865 ns; max/min spread 66.04515851566103%) and Candle [388648761, 520355316, 401643453, 352389150, 360770138] ns (median 388648761 ns; spread 43.21798571229717%). p50, p99 and max also exceeded the 20% threshold for both backends. The raw rounds are committed in `bench/embed-backends/raw/x86-vm.json` at 5a2d1a2; they need a lower-load nighttime recheck after the primary matrix completes.
Author
Owner

Completed the five-round x86-vm four-thread f32 batch-1 indexing set. ORT elapsed round medians were [282080455925, 91975246764, 54244857034, 125821567721, 43209664882] ns (median 91975246764 ns; max/min spread 259.71204149733944%). Candle elapsed round medians were [316165915933, 272554731655, 295184966881, 261436780200, 159858849254] ns (median 272554731655 ns; spread 57.34887291439637%). Both exceed the 20% threshold. Raw data is committed in bench/embed-backends/raw/x86-vm.json at f27355a.

Completed the five-round x86-vm four-thread f32 batch-1 indexing set. ORT elapsed round medians were [282080455925, 91975246764, 54244857034, 125821567721, 43209664882] ns (median 91975246764 ns; max/min spread 259.71204149733944%). Candle elapsed round medians were [316165915933, 272554731655, 295184966881, 261436780200, 159858849254] ns (median 272554731655 ns; spread 57.34887291439637%). Both exceed the 20% threshold. Raw data is committed in `bench/embed-backends/raw/x86-vm.json` at f27355a.
Author
Owner

Completed the five-round x86-vm four-thread f32 batch-8 indexing set. ORT elapsed round medians were [48493144160, 41940047777, 49093175650, 41987715893, 149692655803] ns (median 48493144160 ns; max/min spread 222.20173571438724%). Candle elapsed round medians were [144029888431, 161117145231, 219695268419, 246162379968, 206348273024] ns (median 206348273024 ns; spread 49.49520053658077%). Both exceed the 20% threshold. Raw data is committed in bench/embed-backends/raw/x86-vm.json at 6e7be87.

Completed the five-round x86-vm four-thread f32 batch-8 indexing set. ORT elapsed round medians were [48493144160, 41940047777, 49093175650, 41987715893, 149692655803] ns (median 48493144160 ns; max/min spread 222.20173571438724%). Candle elapsed round medians were [144029888431, 161117145231, 219695268419, 246162379968, 206348273024] ns (median 206348273024 ns; spread 49.49520053658077%). Both exceed the 20% threshold. Raw data is committed in `bench/embed-backends/raw/x86-vm.json` at 6e7be87.
Author
Owner

Completed the five-round x86-vm four-thread f32 batch-32 indexing set. ORT elapsed round medians were [117033462354, 127242850736, 90314049899, 60519191918, 167522053247] ns (median 117033462354 ns; max/min spread 91.42928798033874%). Candle elapsed round medians were [311792876023, 324664635249, 269974095773, 281736763671, 337434078764] ns (median 311792876023 ns; spread 21.63615277278615%). Both exceed the 20% threshold. Raw data is committed in bench/embed-backends/raw/x86-vm.json at cb9c946.

Completed the five-round x86-vm four-thread f32 batch-32 indexing set. ORT elapsed round medians were [117033462354, 127242850736, 90314049899, 60519191918, 167522053247] ns (median 117033462354 ns; max/min spread 91.42928798033874%). Candle elapsed round medians were [311792876023, 324664635249, 269974095773, 281736763671, 337434078764] ns (median 311792876023 ns; spread 21.63615277278615%). Both exceed the 20% threshold. Raw data is committed in `bench/embed-backends/raw/x86-vm.json` at cb9c946.
Author
Owner

Completed the five-round x86-vm four-thread best-CPU query comparison. Query p95 round medians were ORT quantized [36761019, 22327773, 22574058, 22212109, 22745553] ns (median 22574058 ns; max/min spread 64.44968822176322%) and Candle f32 [119591055, 72558816, 70410813, 70551339, 76456941] ns (median 72558816 ns; spread 67.77982981420205%). p50, p99 and max also exceeded the 20% threshold for both backends. Raw data is committed in bench/embed-backends/raw/x86-vm.json at cb9c946.

Completed the five-round x86-vm four-thread best-CPU query comparison. Query p95 round medians were ORT quantized [36761019, 22327773, 22574058, 22212109, 22745553] ns (median 22574058 ns; max/min spread 64.44968822176322%) and Candle f32 [119591055, 72558816, 70410813, 70551339, 76456941] ns (median 72558816 ns; spread 67.77982981420205%). p50, p99 and max also exceeded the 20% threshold for both backends. Raw data is committed in `bench/embed-backends/raw/x86-vm.json` at cb9c946.
Author
Owner

X86 four-thread best-CPU index, batch 1, completed five paired rounds. Round medians in ns were ORT [47271389889, 60329752698, 126474479552, 84052686836, 48230092484] (median 60329752698, spread 131.2836305818732%) and Candle [188725179668, 300129654085, 314732679606, 234266030507, 198479995994] (median 234266030507, spread 53.78820807493676%). Both exceed the protocol's 20% noise threshold. Raw checkpoint: commit 702ba15. The primary run continues with batch sizes 8 and 32, then parity, quality, and load measurements.

X86 four-thread best-CPU index, batch 1, completed five paired rounds. Round medians in ns were ORT `[47271389889, 60329752698, 126474479552, 84052686836, 48230092484]` (median `60329752698`, spread `131.2836305818732%`) and Candle `[188725179668, 300129654085, 314732679606, 234266030507, 198479995994]` (median `234266030507`, spread `53.78820807493676%`). Both exceed the protocol's 20% noise threshold. Raw checkpoint: commit `702ba15`. The primary run continues with batch sizes 8 and 32, then parity, quality, and load measurements.
Author
Owner

Progress on job/embed-bench: x86 four-thread best-CPU batch 1 finished all five pairs and was checkpointed in 702ba15; both measured round-median spreads exceed 20%. Four-thread best-CPU batch 8 is on round 4 of 5. The corpus runs are serialized and each process is taking several minutes; the runner saves raw data after each process. The host matrix, parity/quality, load measurements, report, and lower-load nighttime rechecks remain outstanding.

Progress on `job/embed-bench`: x86 four-thread best-CPU batch 1 finished all five pairs and was checkpointed in `702ba15`; both measured round-median spreads exceed 20%. Four-thread best-CPU batch 8 is on round 4 of 5. The corpus runs are serialized and each process is taking several minutes; the runner saves raw data after each process. The host matrix, parity/quality, load measurements, report, and lower-load nighttime rechecks remain outstanding.
Author
Owner

X86 four-thread best-CPU index, batch 8, completed five paired rounds. Round medians in ns were ORT [52776695318, 68946822887, 67923669247, 64403899965, 71994672636] (median 67923669247, spread 28.293491107076502%) and Candle [222098279804, 231003376297, 219969484513, 273943377838, 185059768855] (median 222098279804, spread 40.019944801661275%). Both exceed the protocol's 20% noise threshold. Raw checkpoint: commit 6ac8f93. Four-thread batch 32 and parity/quality/load stages remain.

X86 four-thread best-CPU index, batch 8, completed five paired rounds. Round medians in ns were ORT `[52776695318, 68946822887, 67923669247, 64403899965, 71994672636]` (median `67923669247`, spread `28.293491107076502%`) and Candle `[222098279804, 231003376297, 219969484513, 273943377838, 185059768855]` (median `222098279804`, spread `40.019944801661275%`). Both exceed the protocol's 20% noise threshold. Raw checkpoint: commit `6ac8f93`. Four-thread batch 32 and parity/quality/load stages remain.
Author
Owner

Progress: x86 four-thread best-CPU batch 32 is on paired round 4 of 5. The first three rounds are complete; the current Candle process is still active with CPU time advancing. The raw JSON checkpoints each completed backend result. A five-round spread summary, parity/quality/load stages, and the required nighttime rechecks remain.

Progress: x86 four-thread best-CPU batch 32 is on paired round 4 of 5. The first three rounds are complete; the current Candle process is still active with CPU time advancing. The raw JSON checkpoints each completed backend result. A five-round spread summary, parity/quality/load stages, and the required nighttime rechecks remain.
Author
Owner

X86 four-thread best-CPU index, batch 32, completed five paired rounds. Round medians in ns were ORT [89769484354, 66913119869, 85054002039, 78218461558, 81855455973] (median 81855455973, spread 27.922835702655135%) and Candle [264272034283, 194181720956, 410839482344, 308739002247, 225998139326] (median 264272034283, spread 81.9828560278113%). Both exceed the protocol's 20% noise threshold. Raw checkpoint: commit 9924425. The runner has started x86 parity vector generation at one thread.

X86 four-thread best-CPU index, batch 32, completed five paired rounds. Round medians in ns were ORT `[89769484354, 66913119869, 85054002039, 78218461558, 81855455973]` (median `81855455973`, spread `27.922835702655135%`) and Candle `[264272034283, 194181720956, 410839482344, 308739002247, 225998139326]` (median `264272034283`, spread `81.9828560278113%`). Both exceed the protocol's 20% noise threshold. Raw checkpoint: commit `9924425`. The runner has started x86 parity vector generation at one thread.
Author
Owner

X86 one-thread F32 embedding comparison completed across all 2,000 texts. Minimum cosine similarity was 0.9999999999977327; mean was 0.9999999999994899. On the 50 relevance pairs, ORT F32 and Candle F32 each measured Recall@1 43/50 and Recall@10 47/50. Raw checkpoint: commit 6076dfe. ORT best-CPU vectors are being generated for its quality score; four-thread parity remains.

X86 one-thread F32 embedding comparison completed across all 2,000 texts. Minimum cosine similarity was `0.9999999999977327`; mean was `0.9999999999994899`. On the 50 relevance pairs, ORT F32 and Candle F32 each measured Recall@1 `43/50` and Recall@10 `47/50`. Raw checkpoint: commit `6076dfe`. ORT best-CPU vectors are being generated for its quality score; four-thread parity remains.
Author
Owner

X86 ORT AVX2 quint8 best-CPU quality completed on the same 50 relevance pairs: Recall@1 45/50; Recall@10 48/50. The F32 ORT and Candle scores remain 43/50 and 47/50. Raw checkpoint: commit 01df3d9. Four-thread Candle F32 vectors are currently running for the remaining parity check.

X86 ORT AVX2 quint8 best-CPU quality completed on the same 50 relevance pairs: Recall@1 `45/50`; Recall@10 `48/50`. The F32 ORT and Candle scores remain `43/50` and `47/50`. Raw checkpoint: commit `01df3d9`. Four-thread Candle F32 vectors are currently running for the remaining parity check.
Author
Owner

Reporting finding: the O2 raw file contains valid continuation load data, but summarize.py selects only night_recheck or primary. A reproduction returned selected set primary and zero load samples for ort-f32 at one thread, while the raw continuation summary has 10 samples and median 218988168.5 ns. The O2 memory_rss_only batch-32 ORT F32 sample records 506496 KB before exit 137; the helper filters it out because it has no completed embedded_texts payload. I’m correcting the summary extraction so RESULTS.md includes those measured values and labels the incomplete runs accurately.

Reporting finding: the O2 raw file contains valid continuation load data, but `summarize.py` selects only `night_recheck` or `primary`. A reproduction returned selected set `primary` and zero load samples for `ort-f32` at one thread, while the raw `continuation` summary has 10 samples and median `218988168.5 ns`. The O2 `memory_rss_only` batch-32 ORT F32 sample records `506496 KB` before exit 137; the helper filters it out because it has no completed `embedded_texts` payload. I’m correcting the summary extraction so RESULTS.md includes those measured values and labels the incomplete runs accurately.
Author
Owner

The O2 summary extraction correction is committed as 0b6858d. python3 scripts/test_summarize.py reproduced 2 failures and 1 error before the fix; afterward it reports Ran 3 tests ... OK. O2 model-load rows now show continuation medians (218.988†, 66.243, 164.452, 218.198, 66.971, 170.706 ms), and the batch-32 RSS-only sample appears as 506,496 KB despite process exit 137. summarize.py also includes peak RSS from query processes that hit their CPU cap.

The O2 summary extraction correction is committed as `0b6858d`. `python3 scripts/test_summarize.py` reproduced 2 failures and 1 error before the fix; afterward it reports `Ran 3 tests ... OK`. O2 model-load rows now show continuation medians (`218.988†`, `66.243`, `164.452`, `218.198`, `66.971`, `170.706` ms), and the batch-32 RSS-only sample appears as `506,496 KB` despite process exit 137. `summarize.py` also includes peak RSS from query processes that hit their CPU cap.
Author
Owner

The completed x86 primary data has 20 groups with a five-round spread above 20%: 4 query groups, 12 indexing groups, and 4 model-load groups. The runner's --rerun-noisy path will check each after 22:00 local time and only rerun it when the start load average is below that group's recorded baseline. The 20-group count is computed from the committed raw/x86-vm.json.

The completed x86 primary data has 20 groups with a five-round spread above 20%: 4 query groups, 12 indexing groups, and 4 model-load groups. The runner's `--rerun-noisy` path will check each after 22:00 local time and only rerun it when the start load average is below that group's recorded baseline. The 20-group count is computed from the committed `raw/x86-vm.json`.
Author
Owner

O2 recheck audit: noisy_primary_groups() returns an empty list for the raw O2 host, although the continuation F32 one-thread ORT load median is 218988168.5 ns with 26.03212328340926% spread. The recheck path also looks up its load baseline under primary, while those load records use continuation. The host has 839.4807249999998 CPU seconds including the runner against the 900-second limit. I’m adding a bounded recheck path that includes continuation data and keeps a CPU reserve.

O2 recheck audit: `noisy_primary_groups()` returns an empty list for the raw O2 host, although the continuation F32 one-thread ORT load median is `218988168.5 ns` with `26.03212328340926%` spread. The recheck path also looks up its load baseline under `primary`, while those load records use `continuation`. The host has `839.4807249999998` CPU seconds including the runner against the 900-second limit. I’m adding a bounded recheck path that includes continuation data and keeps a CPU reserve.
Author
Owner

O2 nighttime recheck fix committed as f749104 (Recheck noisy O2 loads within CPU budget). The regression test covers the continuation load baseline and the bounded O2 budget. The existing noisy F32 one-thread ORT load group used 2.9 child CPU seconds for ten fresh processes; current accounted use is 839.4807249999998 seconds. The retry runner stops launching at 870 seconds, leaving 30 seconds below the issue cap. Python checks pass: python3 scripts/test_run.py (2 passed) and python3 scripts/test_summarize.py (4 passed).

O2 nighttime recheck fix committed as `f749104` (`Recheck noisy O2 loads within CPU budget`). The regression test covers the continuation load baseline and the bounded O2 budget. The existing noisy F32 one-thread ORT load group used 2.9 child CPU seconds for ten fresh processes; current accounted use is `839.4807249999998` seconds. The retry runner stops launching at 870 seconds, leaving 30 seconds below the issue cap. Python checks pass: `python3 scripts/test_run.py` (2 passed) and `python3 scripts/test_summarize.py` (4 passed).
Author
Owner

Recheck gate audit: the raw thresholds for all currently noisy groups are x86 load below 3.19 and O2 load below 0.94. At 2026-09-25 11:04 CEST, x86 was 16.55 and O2 (UTC host) was 1.00, so no daytime retry qualifies. The runner records the recheck source revision and exact argv in the host JSON; source metadata fix is committed as 341fefd. The overlapping host-local night window is 22:00–05:59 UTC on O2 and 22:00–05:59 CEST on x86.

Recheck gate audit: the raw thresholds for all currently noisy groups are x86 load below `3.19` and O2 load below `0.94`. At 2026-09-25 11:04 CEST, x86 was `16.55` and O2 (UTC host) was `1.00`, so no daytime retry qualifies. The runner records the recheck source revision and exact argv in the host JSON; source metadata fix is committed as `341fefd`. The overlapping host-local night window is 22:00–05:59 UTC on O2 and 22:00–05:59 CEST on x86.
Author
Owner

Standalone bench crate gates pass: cargo fmt --manifest-path bench/embed-backends/Cargo.toml --check (exit 0, no output); cargo clippy --manifest-path bench/embed-backends/Cargo.toml --all-targets --all-features -- -D warnings (Finished dev profile [unoptimized + debuginfo] target(s) in 0.71s); and cargo test --manifest-path bench/embed-backends/Cargo.toml --all-features (7 passed). The root workspace Clippy gate could not compile calternal-server because ignored build output apps/web/build/ is absent; I left paths outside this job's ownership unchanged.

Standalone bench crate gates pass: `cargo fmt --manifest-path bench/embed-backends/Cargo.toml --check` (exit 0, no output); `cargo clippy --manifest-path bench/embed-backends/Cargo.toml --all-targets --all-features -- -D warnings` (`Finished `dev` profile [unoptimized + debuginfo] target(s) in 0.71s`); and `cargo test --manifest-path bench/embed-backends/Cargo.toml --all-features` (7 passed). The root workspace Clippy gate could not compile `calternal-server` because ignored build output `apps/web/build/` is absent; I left paths outside this job's ownership unchanged.
Author
Owner

Preflight fix committed as 3246b23 (Import Python runtime for recheck metadata). A focused regression test caught the missing sys import in recheck argv capture before any nighttime measurement started. python3 scripts/test_run.py now passes 3 tests and python3 scripts/test_summarize.py passes 4 tests.

Preflight fix committed as `3246b23` (`Import Python runtime for recheck metadata`). A focused regression test caught the missing `sys` import in recheck argv capture before any nighttime measurement started. `python3 scripts/test_run.py` now passes 3 tests and `python3 scripts/test_summarize.py` passes 4 tests.
Author
Owner

Recheck load sampling fix committed as b71686e (Check load before each noisy rerun). The old code reused the initial 0.50 load for a group even when a fresh immediately-before-group sample was 0.40; the regression test failed on that evidence, then passed after the fix. Each group now records and compares its own current load against its source-set median baseline.

Recheck load sampling fix committed as `b71686e` (`Check load before each noisy rerun`). The old code reused the initial `0.50` load for a group even when a fresh immediately-before-group sample was `0.40`; the regression test failed on that evidence, then passed after the fix. Each group now records and compares its own current load against its source-set median baseline.
Author
Owner

Primary measurements are now reported in commit 5025561 (Report primary embedding backend results). It follows the already committed x86 and O2 raw JSON and marks every noisy summary with †; its Night rechecks section states that only the primary set is selected so far. The runner session is waiting for each host's local night window and a lower per-group load sample. After eligible rechecks finish, I will commit the updated raw JSON per host, regenerate both summary tables from raw, update RESULTS.md, and post those tables here.

Primary measurements are now reported in commit `5025561` (`Report primary embedding backend results`). It follows the already committed x86 and O2 raw JSON and marks every noisy summary with `†`; its Night rechecks section states that only the primary set is selected so far. The runner session is waiting for each host's local night window and a lower per-group load sample. After eligible rechecks finish, I will commit the updated raw JSON per host, regenerate both summary tables from raw, update RESULTS.md, and post those tables here.
Author
Owner

O2 raw query audit: the F32/1-thread ORT row contains four successful 1,000-run samples, not five. The first three Candle runs used 120.44, 120.83, and 120.62 CPU seconds and exited 137 with zero latency samples; round four's Candle side was not launched, and no fifth or four-thread query ran. RESULTS.md now records this exact partial set in commit ab99360 (Clarify O2 partial query rounds).

O2 raw query audit: the F32/1-thread ORT row contains four successful 1,000-run samples, not five. The first three Candle runs used `120.44`, `120.83`, and `120.62` CPU seconds and exited `137` with zero latency samples; round four's Candle side was not launched, and no fifth or four-thread query ran. RESULTS.md now records this exact partial set in commit `ab99360` (`Clarify O2 partial query rounds`).
Author
Owner

Report format audit found that query p95/p99 cells showed only each metric's median, while p50 showed the five-round range. The raw summaries already retain per-round p95/p99 values. Commit 83a48b2 updates the renderer and RESULTS.md so every query percentile, max, and cold query shows its median and min–max across successful rounds; regression tests pass (3 run tests, 5 summary tests).

Report format audit found that query p95/p99 cells showed only each metric's median, while p50 showed the five-round range. The raw summaries already retain per-round p95/p99 values. Commit `83a48b2` updates the renderer and RESULTS.md so every query percentile, max, and cold query shows its median and min–max across successful rounds; regression tests pass (3 run tests, 5 summary tests).
Author
Owner

The model-load cells now show the median and measured min–max of all ten fresh-process samples, alongside the existing >20% noise marker. This came from the raw samples already recorded; python3 scripts/test_summarize.py passes 6 tests. Commit: a065082 (Show fresh-process load ranges).

The model-load cells now show the median and measured min–max of all ten fresh-process samples, alongside the existing >20% noise marker. This came from the raw samples already recorded; `python3 scripts/test_summarize.py` passes 6 tests. Commit: `a065082` (`Show fresh-process load ranges`).
Author
Owner

Decision not specified by the issue: O2 runs use taskset -c 0-1 for the one-thread configuration and taskset -c 0-3 for the four-thread configuration. The report records the rationale: use the first two logical CPUs for the one-thread configuration and all four logical CPUs for the four-thread configuration. This appears in commit 80990f2 (Explain O2 thread affinity choice).

Decision not specified by the issue: O2 runs use `taskset -c 0-1` for the one-thread configuration and `taskset -c 0-3` for the four-thread configuration. The report records the rationale: use the first two logical CPUs for the one-thread configuration and all four logical CPUs for the four-thread configuration. This appears in commit `80990f2` (`Explain O2 thread affinity choice`).
Author
Owner

License audit: cargo metadata resolved 96 packages; no GPL-only, proprietary, or non-commercial dependency license appeared. The pinned all-MiniLM-L6-v2 revision 1110a243fdf4706b3f48d6b90e8cd65606993a20 declares Apache-2.0 on its Hugging Face revision page; RESULTS.md records that source in commit 45674d2 (Record pinned model license).

License audit: `cargo metadata` resolved 96 packages; no GPL-only, proprietary, or non-commercial dependency license appeared. The pinned `all-MiniLM-L6-v2` revision `1110a243fdf4706b3f48d6b90e8cd65606993a20` declares Apache-2.0 on its Hugging Face revision page; RESULTS.md records that source in commit `45674d2` (`Record pinned model license`).
Author
Owner

Report audit found the measured parity values in raw were not rendered. RESULTS.md now reports each x86 thread setting against all 2,000 texts: minimum cosine 0.9999999999977327, mean 0.9999999999994899 at both 1 and 4 threads. O2 parity remains unmeasured because its CPU budget ended before vector inference. Commit f129946 (Report embedding parity measurements).

Report audit found the measured parity values in raw were not rendered. RESULTS.md now reports each x86 thread setting against all 2,000 texts: minimum cosine `0.9999999999977327`, mean `0.9999999999994899` at both 1 and 4 threads. O2 parity remains unmeasured because its CPU budget ended before vector inference. Commit `f129946` (`Report embedding parity measurements`).
Author
Owner

Night-boundary fix committed as 33a485d (Enforce recheck window for every group). The runner now checks the local-hour window before each group, after its load probe, and immediately before launch; a group reaching 06:00 is recorded as skipped for a later night. Raw host metadata keeps every recheck invocation's argv and Git SHA, and group decisions link to the invocation. Regression evidence: the test supplies hour 22 at startup then 6 before the group and confirms that the backend is not launched; run tests pass (4).

Night-boundary fix committed as `33a485d` (`Enforce recheck window for every group`). The runner now checks the local-hour window before each group, after its load probe, and immediately before launch; a group reaching 06:00 is recorded as skipped for a later night. Raw host metadata keeps every recheck invocation's argv and Git SHA, and group decisions link to the invocation. Regression evidence: the test supplies hour 22 at startup then 6 before the group and confirms that the backend is not launched; run tests pass (4).
Author
Owner

Each group decision now stores its load-check local hour and launch local hour, along with baseline, observed load, and invocation ID. This makes partial runs that cross 06:00 auditable per group. Commit 5398f4d (Record per-group recheck times); Python regression suites pass (4 run tests, 6 summary tests).

Each group decision now stores its load-check local hour and launch local hour, along with baseline, observed load, and invocation ID. This makes partial runs that cross 06:00 auditable per group. Commit `5398f4d` (`Record per-group recheck times`); Python regression suites pass (4 run tests, 6 summary tests).
Author
Owner

Found a recheck recovery gap while auditing the overnight runner: one_group saves each round before launching its child processes, but a later run treated any night_recheck measurement entry as a completed group. A process stop could therefore leave a partial raw group that would never be retried. Added a completion-ledger check in rerun_noisy_groups and a regression test. The test failed before the change (calls == []), then passed after it. Commit: 6d93a22f4be8057217a590be5575cfc0272435b7.

Found a recheck recovery gap while auditing the overnight runner: `one_group` saves each round before launching its child processes, but a later run treated any `night_recheck` measurement entry as a completed group. A process stop could therefore leave a partial raw group that would never be retried. Added a completion-ledger check in `rerun_noisy_groups` and a regression test. The test failed before the change (`calls == []`), then passed after it. Commit: `6d93a22f4be8057217a590be5575cfc0272435b7`.
Author
Owner

Added per-sample recheck_invocation_id metadata to query/index rounds, model-load samples, and their summaries. This links every nighttime raw sample to the exact invocation record containing its source SHA and argv, including if a partial run is retried later. The regression test exercises the query and load recording paths and checks all five query rounds and ten load samples. Test suites pass. Commit: 9f03b43f2cee09b55c4973ef4ecf1cc4b61cf70e.

Added per-sample `recheck_invocation_id` metadata to query/index rounds, model-load samples, and their summaries. This links every nighttime raw sample to the exact invocation record containing its source SHA and argv, including if a partial run is retried later. The regression test exercises the query and load recording paths and checks all five query rounds and ten load samples. Test suites pass. Commit: `9f03b43f2cee09b55c4973ef4ecf1cc4b61cf70e`.
Author
Owner

Found a retry aggregation defect while hardening raw recovery: the summarizer included successful rounds from an older night_recheck invocation alongside a later retry, and selected the first noise summary. A two-invocation fixture reproduced it (query_rounds returned p50 values [1000, 2000] instead of only [2000]). The summarizer now selects the latest invocation ID from the matching noise summary and filters query, indexing, RSS, load, and noise samples to that invocation. Both Python suites pass. Commit: 22fe86516d98e94ad22011f8de2a741287284e2d.

Found a retry aggregation defect while hardening raw recovery: the summarizer included successful rounds from an older `night_recheck` invocation alongside a later retry, and selected the first noise summary. A two-invocation fixture reproduced it (`query_rounds` returned p50 values `[1000, 2000]` instead of only `[2000]`). The summarizer now selects the latest invocation ID from the matching noise summary and filters query, indexing, RSS, load, and noise samples to that invocation. Both Python suites pass. Commit: `22fe86516d98e94ad22011f8de2a741287284e2d`.
Author
Owner

Before the night rechecks, I traced a scheduler ordering problem: O2 uses UTC local time, while x86 uses Europe/Berlin. A synchronous scheduler could start x86 at 22:00 Berlin and remain blocked in a multi-hour group, missing O2's 22:00–05:59 UTC window. The job scheduler now monitors x86 and O2 independently, so each host can start only when its own night and per-group load gates pass. This is operational orchestration only; no recheck has started yet. I will document the scheduling choice with the final results.

Before the night rechecks, I traced a scheduler ordering problem: O2 uses UTC local time, while x86 uses Europe/Berlin. A synchronous scheduler could start x86 at 22:00 Berlin and remain blocked in a multi-hour group, missing O2's 22:00–05:59 UTC window. The job scheduler now monitors x86 and O2 independently, so each host can start only when its own night and per-group load gates pass. This is operational orchestration only; no recheck has started yet. I will document the scheduling choice with the final results.
kayg closed this issue 2026-09-25 12:11:00 +00:00
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#74
No description provided.