EVAL: self-hosted decision model (pplx-decider-v1-27b) for tag suggestions, mail categories and document types #656

Open
opened 2026-10-02 04:20:31 +00:00 by kayg · 11 comments
Owner

Owner question (2026-10-02)

Can the open-weight decision model perplexity-ai/pplx-decider-v1-27b (Apache-2.0, fine-tuned from Qwen 3.8 27B, multimodal, outputs a probability distribution over a fixed answer set) serve as a self-hosted replacement for a hosted decision API like the one calternal.js used for tag suggestions?

Background

  • calternal.js had Jev tag suggestions in the composer (DESIGN 2026-09-18). The composer sent one event title plus the vault's area/* tag names, and got back a score per tag. A score ≥ 0.90 auto-applied the tag, ≥ 0.50 suggested it, and vague titles got nothing. An 80-title golden set measured 97.8% precision at 0.90. Code: calternal.js/packages/sync-client/src/jev.ts, .claude/skills/jev-eval/ (golden set references/golden.json).
  • calternal did not port it on purpose: "Jev tag suggestions fold into the AI plugin's intent layer later (one AI path)" (DESIGN, Composer bullet). Composer.svelte notes it as not ported.
  • The hosted service answers in 70–500 ms. A self-hosted 27B model on our CPU-only hosts will be much slower, unless the constant part of the prompt (instructions + the User's tag list) is prefix-cached, so each request only processes the short title.

Evaluation (measurement only, no product code)

On the perf VM (CPU only; flock /root/perf.lock), with a 4-bit build of pplx-decider-v1-27b in a llama.cpp-class runtime with prompt/prefix caching:

  1. Tag suggestions: the calternal.js 80-title golden set with the same question shape and the 0.50/0.90 decision rule. Report precision of auto-applied tags, silence on vague titles, and recall, compared with the documented hosted results (97.8% at 0.90).
  2. Latency: cold, warm and prefix-cached p50/p95 per decision, peak RSS, and CPU seconds. Target for interactive use ≤ 500 ms p95 warm; anything slower is background-only.
  3. Other decisions on the private corpora under ~/calternal-private/, reporting aggregates only and never content: mail category (People/Newsletters/Notifications/Receipts) and statement detection vs the current rules; document type after OCR (#584).
  4. Licences: confirm Apache-2.0 for the weights and the base model, and that the runtime is AGPL-compatible.

Decision rule (fixed now): interactive features ship only if p95 warm ≤ 500 ms on the perf VM AND tag precision ≥ 97% at 0.90; otherwise background-only (opt-in, bounded worker like OCR), or wait for a smaller model.

## Owner question (2026-10-02) Can the open-weight decision model `perplexity-ai/pplx-decider-v1-27b` (Apache-2.0, fine-tuned from Qwen 3.8 27B, multimodal, outputs a probability distribution over a fixed answer set) serve as a **self-hosted replacement for a hosted decision API** like the one calternal.js used for tag suggestions? ## Background - calternal.js had **Jev tag suggestions** in the composer (DESIGN 2026-09-18). The composer sent one event title plus the vault's `area/*` tag names, and got back a score per tag. A score ≥ 0.90 auto-applied the tag, ≥ 0.50 suggested it, and vague titles got nothing. An 80-title golden set measured 97.8% precision at 0.90. Code: `calternal.js/packages/sync-client/src/jev.ts`, `.claude/skills/jev-eval/` (golden set `references/golden.json`). - calternal did not port it on purpose: "Jev tag suggestions fold into the AI plugin's intent layer later (one AI path)" (DESIGN, Composer bullet). `Composer.svelte` notes it as not ported. - The hosted service answers in 70–500 ms. A self-hosted 27B model on our CPU-only hosts will be much slower, unless the constant part of the prompt (instructions + the User's tag list) is prefix-cached, so each request only processes the short title. ## Evaluation (measurement only, no product code) On the perf VM (CPU only; `flock /root/perf.lock`), with a 4-bit build of pplx-decider-v1-27b in a llama.cpp-class runtime with prompt/prefix caching: 1. **Tag suggestions:** the calternal.js 80-title golden set with the same question shape and the 0.50/0.90 decision rule. Report precision of auto-applied tags, silence on vague titles, and recall, compared with the documented hosted results (97.8% at 0.90). 2. **Latency:** cold, warm and prefix-cached p50/p95 per decision, peak RSS, and CPU seconds. Target for interactive use ≤ 500 ms p95 warm; anything slower is background-only. 3. **Other decisions** on the private corpora under `~/calternal-private/`, reporting aggregates only and never content: mail category (People/Newsletters/Notifications/Receipts) and statement detection vs the current rules; document type after OCR (#584). 4. **Licences:** confirm Apache-2.0 for the weights and the base model, and that the runtime is AGPL-compatible. Decision rule (fixed now): interactive features ship only if p95 warm ≤ 500 ms on the perf VM AND tag precision ≥ 97% at 0.90; otherwise background-only (opt-in, bounded worker like OCR), or wait for a smaller model.
Author
Owner

Scope update (owner, 2026-10-02): evaluate three candidates that may fit our CPU-only hardware, smallest first, all Apache-2.0:

  1. Laya (421M, ModernBERT-large encoder + typed decision head; open weights on Hugging Face, reference code on GitHub). Candidate for the interactive tier (composer tag suggestions, mail category on arrival): expected tens of ms on CPU.
  2. Clef-flash (Qwen 3.5 9B backbone + rank-256 LoRA, prefill-only, 64k context; Hugging Face Cloudflare/clef-flash). Candidate for the background tier, with a 4-bit build (~5–6 GB).
  3. pplx-decider-v1-27b and Clef (27B): reference only, background at best on CPU.
    Same golden set, decision rule and private corpora as above (aggregates only). Add a two-tier option to the verdict: the small model decides when confident, and only uncertain items go to the larger model in the background. Report RSS and per-decision CPU ms for each on the perf VM.
Scope update (owner, 2026-10-02): evaluate three candidates that may fit our CPU-only hardware, smallest first, all Apache-2.0: 1. **Laya** (421M, ModernBERT-large encoder + typed decision head; open weights on Hugging Face, reference code on GitHub). Candidate for the interactive tier (composer tag suggestions, mail category on arrival): expected tens of ms on CPU. 2. **Clef-flash** (Qwen 3.5 9B backbone + rank-256 LoRA, prefill-only, 64k context; Hugging Face `Cloudflare/clef-flash`). Candidate for the background tier, with a 4-bit build (~5–6 GB). 3. **pplx-decider-v1-27b** and **Clef** (27B): reference only, background at best on CPU. Same golden set, decision rule and private corpora as above (aggregates only). Add a two-tier option to the verdict: the small model decides when confident, and only uncertain items go to the larger model in the background. Report RSS and per-decision CPU ms for each on the perf VM.
Author
Owner

Starting the measurement-only evaluation on branch job/decider-656, based on dev at c4a61e8cf0. I read the issue body and 2026-10-02 scope update. I will evaluate Laya first, then Clef-flash, and attempt 27B references if time permits; all model weights stay on the perf VM.

Starting the measurement-only evaluation on branch job/decider-656, based on dev at c4a61e8cf090170f35b1bed3350d9de20c83ecd5. I read the issue body and 2026-10-02 scope update. I will evaluate Laya first, then Clef-flash, and attempt 27B references if time permits; all model weights stay on the perf VM.
Author
Owner

Finding from the permitted corpus inventory: the perf VM staging directory /root/private currently contains 0 files. The local document metadata has 614 records, with document_type unset on all 614; its tags include 12 statement-tagged records, and the private card-statements set has 14 PDFs. I found no mail-named corpus folder or EML/MBOX/MSG file. I did not print or transfer any corpus content. Please point me to a separate mail corpus if one exists; I am proceeding with the golden set and available statement data.

Finding from the permitted corpus inventory: the perf VM staging directory /root/private currently contains 0 files. The local document metadata has 614 records, with document_type unset on all 614; its tags include 12 statement-tagged records, and the private card-statements set has 14 PDFs. I found no mail-named corpus folder or EML/MBOX/MSG file. I did not print or transfer any corpus content. Please point me to a separate mail corpus if one exists; I am proceeding with the golden set and available statement data.
Author
Owner

Finding: model and runtime license checks passed for the intended candidates. Hugging Face lists Laya and its ModernBERT-large base as Apache-2.0; PyPI's latest Laya package is 0.3.23 (released 2026-10-01, Apache-2.0). Cloudflare lists Clef-flash and its Qwen3.5-9B base as Apache-2.0. The CPU OpenVINO conversion retains the joint decision head and is Apache-2.0; it is an unofficial conversion. OpenVINO itself is Apache-2.0; PyPI's current release is 2026.4.1, above the conversion's documented 2026.2.0 minimum. The CPU PyTorch 2.11.0 wheel exists for the VM's CPython 3.13, matching the Clef card's tested PyTorch 2.11 / Transformers 5.10.2 versions. The 27B Perplexity weights and Qwen3.8-27B base are Apache-2.0, but the VM has 7.7 GiB total RAM; the 27B 4-bit weights alone have a 13.5 GB lower-bound footprint, so I will not attempt that reference run.

Sources: https://huggingface.co/convaiinnovations/laya ; https://huggingface.co/answerdotai/ModernBERT-large ; https://pypi.org/project/laya/0.3.23/ ; https://huggingface.co/Cloudflare/clef-flash ; https://huggingface.co/Qwen/Qwen3.5-9B ; https://huggingface.co/meossistant/clef-flash-openvino ; https://pypi.org/project/openvino/ ; https://download.pytorch.org/whl/cpu/torch/ ; https://huggingface.co/perplexity-ai/pplx-decider-v1-27b ; https://huggingface.co/Qwen/Qwen3.8-27B .

Finding: model and runtime license checks passed for the intended candidates. Hugging Face lists Laya and its ModernBERT-large base as Apache-2.0; PyPI's latest Laya package is 0.3.23 (released 2026-10-01, Apache-2.0). Cloudflare lists Clef-flash and its Qwen3.5-9B base as Apache-2.0. The CPU OpenVINO conversion retains the joint decision head and is Apache-2.0; it is an unofficial conversion. OpenVINO itself is Apache-2.0; PyPI's current release is 2026.4.1, above the conversion's documented 2026.2.0 minimum. The CPU PyTorch 2.11.0 wheel exists for the VM's CPython 3.13, matching the Clef card's tested PyTorch 2.11 / Transformers 5.10.2 versions. The 27B Perplexity weights and Qwen3.8-27B base are Apache-2.0, but the VM has 7.7 GiB total RAM; the 27B 4-bit weights alone have a 13.5 GB lower-bound footprint, so I will not attempt that reference run. Sources: https://huggingface.co/convaiinnovations/laya ; https://huggingface.co/answerdotai/ModernBERT-large ; https://pypi.org/project/laya/0.3.23/ ; https://huggingface.co/Cloudflare/clef-flash ; https://huggingface.co/Qwen/Qwen3.5-9B ; https://huggingface.co/meossistant/clef-flash-openvino ; https://pypi.org/project/openvino/ ; https://download.pytorch.org/whl/cpu/torch/ ; https://huggingface.co/perplexity-ai/pplx-decider-v1-27b ; https://huggingface.co/Qwen/Qwen3.8-27B .
Author
Owner

Finding from Laya 0.3.23 on the perf VM: its checkpoint reports an invalid calibration temperature (0.10058, below the supported 0.5–5 range). The library clamps it to 0.5 and warns that these confidence values are uncalibrated. This 31-option Jev question uses that bucket. I am keeping the fixed 0.50/0.90 rule for the measurement and will report this calibration limitation with the results.

Finding from Laya 0.3.23 on the perf VM: its checkpoint reports an invalid calibration temperature (0.10058, below the supported 0.5–5 range). The library clamps it to 0.5 and warns that these confidence values are uncalibrated. This 31-option Jev question uses that bucket. I am keeping the fixed 0.50/0.90 rule for the measurement and will report this calibration limitation with the results.
Author
Owner

Correction to the preceding Laya finding: the affected calibration bucket is choice:11+ (the page's generic formatting dropped the bucket name). This is the bucket used for 30 tags plus none.

Correction to the preceding Laya finding: the affected calibration bucket is `choice:11+` (the page's generic formatting dropped the bucket name). This is the bucket used for 30 tags plus none.
Author
Owner

Laya CPU golden-set finding (aggregate only): on the 58 eligible decisions, it returned 9 apply decisions (9/9 correct, 100.0% precision) and 6 suggestions (2/6 correct, 33.3% precision). It returned a decision on 11/52 labeled asked cases (21.15% recall); 6/6 vague cases were silent. Warm p50/p95 were 531.4/716.4 ms, above the 500 ms interactive p95 target. Peak RSS was 2,826.1 MiB; CPU time averaged 2.184 s/decision. Prefix caching is unsupported by this bidirectional encoder path. The 31-choice confidence bucket raised a calibration warning and was clamped, so confidence values are uncalibrated. This is not an interactive-tier pass; proceed to the background tier. Historical hosted baseline cited in the issue is 97.8% precision at the apply threshold and about 90.4% coverage, so local Laya coverage is materially lower.

Laya CPU golden-set finding (aggregate only): on the 58 eligible decisions, it returned 9 apply decisions (9/9 correct, 100.0% precision) and 6 suggestions (2/6 correct, 33.3% precision). It returned a decision on 11/52 labeled asked cases (21.15% recall); 6/6 vague cases were silent. Warm p50/p95 were 531.4/716.4 ms, above the 500 ms interactive p95 target. Peak RSS was 2,826.1 MiB; CPU time averaged 2.184 s/decision. Prefix caching is unsupported by this bidirectional encoder path. The 31-choice confidence bucket raised a calibration warning and was clamped, so confidence values are uncalibrated. This is not an interactive-tier pass; proceed to the background tier. Historical hosted baseline cited in the issue is 97.8% precision at the apply threshold and about 90.4% coverage, so local Laya coverage is materially lower.
Author
Owner

Laya CPU evaluation finding (synthetic 80-title Jev golden set; 58 eligible, 52 labeled): the Torch path made 9 apply decisions, all correct (100% on n=9), 6 suggestions of which 2 were correct (33.3%), returned a correct tag for 11/52 labeled cases (21.15% recall), and stayed silent on all 6 vague titles. Hosted baseline in the issue is 97.8% precision at the apply threshold and about 90.4% coverage. All three measured warm p95 values missed the 500 ms target: Torch 716.4 ms (p50 531.4), FP32 ONNX 740.9 ms (p50 464.4), and Torch repeat for per-case cascade scoring 902.2 ms (p50 570.4). The Torch repeat had higher VM load (0.86/0.58/0.58 at start); ONNX used less RSS (1,939 MiB vs Torch 2,780–2,826 MiB) but did not improve p95. ONNX and Torch made the same tag/action decision on all 58 cases. Prefix caching is unavailable in both bidirectional encoder paths.

The checkpoint's choice:11+ temperature is 0.10058, outside Laya's accepted [0.5, 5] range; both runtimes clamp it to 0.5 and warn that these confidence values are uncalibrated. The 31-option tag request uses this bucket. Therefore the 9/9 apply precision is an observed small sample, not evidence that the fixed confidence threshold is calibrated. The CPU interactive tier does not pass on either latency or calibrated-confidence evidence; continue with Clef-flash as the bounded background tier.

Each measurement held flock /root/perf.lock; uptime/load at starts were: Torch first run 15:19:23 uptime 10:16 load 0.00/0.14/0.84; ONNX 13:32:43 UTC uptime 10:30 load 0.01/0.09/0.45; Torch repeat 13:35:54 UTC uptime 10:33 load 0.86/0.58/0.58.

Laya CPU evaluation finding (synthetic 80-title Jev golden set; 58 eligible, 52 labeled): the Torch path made 9 apply decisions, all correct (100% on n=9), 6 suggestions of which 2 were correct (33.3%), returned a correct tag for 11/52 labeled cases (21.15% recall), and stayed silent on all 6 vague titles. Hosted baseline in the issue is 97.8% precision at the apply threshold and about 90.4% coverage. All three measured warm p95 values missed the 500 ms target: Torch 716.4 ms (p50 531.4), FP32 ONNX 740.9 ms (p50 464.4), and Torch repeat for per-case cascade scoring 902.2 ms (p50 570.4). The Torch repeat had higher VM load (0.86/0.58/0.58 at start); ONNX used less RSS (1,939 MiB vs Torch 2,780–2,826 MiB) but did not improve p95. ONNX and Torch made the same tag/action decision on all 58 cases. Prefix caching is unavailable in both bidirectional encoder paths. The checkpoint's `choice:11+` temperature is 0.10058, outside Laya's accepted [0.5, 5] range; both runtimes clamp it to 0.5 and warn that these confidence values are uncalibrated. The 31-option tag request uses this bucket. Therefore the 9/9 apply precision is an observed small sample, not evidence that the fixed confidence threshold is calibrated. The CPU interactive tier does not pass on either latency or calibrated-confidence evidence; continue with Clef-flash as the bounded background tier. Each measurement held `flock /root/perf.lock`; uptime/load at starts were: Torch first run 15:19:23 uptime 10:16 load 0.00/0.14/0.84; ONNX 13:32:43 UTC uptime 10:30 load 0.01/0.09/0.45; Torch repeat 13:35:54 UTC uptime 10:33 load 0.86/0.58/0.58.
Author
Owner

Clef-flash setup finding: the pinned OpenVINO conversion (meossistant/clef-flash-openvino, revision f244f287938ecb92cd72efbe9109db7a1648bf87) compiled its OpenVINO backbone and joint head inside the 7 GiB systemd memory scope, but its bundled loader then failed at AutoProcessor.from_pretrained(path). Evidence: Transformers 5.10.2 raised ValueError: Unrecognized image processor; this snapshot has tokenizer and processor metadata but no config.json/preprocessor config for AutoProcessor resolution. The evaluation only sends text titles, and the converter's encode_record returns before touching the image processor when no images/videos are present. I will load the bundled tokenizer directly for this text-only run. The failed load produced no inference results and will not be counted as a measurement.

Clef-flash setup finding: the pinned OpenVINO conversion (`meossistant/clef-flash-openvino`, revision `f244f287938ecb92cd72efbe9109db7a1648bf87`) compiled its OpenVINO backbone and joint head inside the 7 GiB systemd memory scope, but its bundled loader then failed at `AutoProcessor.from_pretrained(path)`. Evidence: Transformers 5.10.2 raised `ValueError: Unrecognized image processor`; this snapshot has tokenizer and processor metadata but no `config.json`/preprocessor config for AutoProcessor resolution. The evaluation only sends text titles, and the converter's `encode_record` returns before touching the image processor when no images/videos are present. I will load the bundled tokenizer directly for this text-only run. The failed load produced no inference results and will not be counted as a measurement.
Author
Owner

Clef-flash background-tier finding: the 6.32 GB mixed INT4/INT8 OpenVINO backbone plus 243.5 MB joint head staged successfully on the VM. On the 4-core i9-9900K perf VM (7.7 GiB RAM, no swap), the run used a systemd scope capped at 7 GiB. It printed MODEL_LOAD_DONE, then the scoped process was killed (exit 137) before returning the first decision; no latency or quality result was produced. The conversion maintainer's card estimates ~8.5 GB peak RAM. I will not remove the cap on this host. The 9B tier remains unscored; rerun on a host with at least 16 GiB RAM to allow headroom. The attempted phase held /root/perf.lock; start uptime 15:59:10 CEST / 13:59:10 UTC, 10:56 uptime, load 0.14/0.20/0.32.

Clef-flash background-tier finding: the 6.32 GB mixed INT4/INT8 OpenVINO backbone plus 243.5 MB joint head staged successfully on the VM. On the 4-core i9-9900K perf VM (7.7 GiB RAM, no swap), the run used a systemd scope capped at 7 GiB. It printed `MODEL_LOAD_DONE`, then the scoped process was killed (exit 137) before returning the first decision; no latency or quality result was produced. The conversion maintainer's card estimates ~8.5 GB peak RAM. I will not remove the cap on this host. The 9B tier remains unscored; rerun on a host with at least 16 GiB RAM to allow headroom. The attempted phase held `/root/perf.lock`; start uptime 15:59:10 CEST / 13:59:10 UTC, 10:56 uptime, load 0.14/0.20/0.32.
Author
Owner

Final report — measurement only

No product code changed. The worktree is clean on job/decider-656; head SHA: c4a61e8cf090170f35b1bed3350d9de20c83ecd5. No commit was needed. Model files and staging data were removed from the perf VM after measurement. Local helpers and aggregate-only outputs remain under ~/calternal-private/decider-656/.

Jev golden set

The original parser produced 80 titles: 58 eligible, 52 labeled, 6 vague, and 22 skipped by the existing eligibility rule. Each eligible title was scored as one request with 30 tags plus none. Cold is the first decision after model load; warm percentiles cover the other 57 calls. CPU figures are process CPU time during inference, divided by 58.

Model/runtime Load ms / first decision ms Warm p50 / p95 ms Peak RSS MiB CPU total / per decision s Apply precision Suggest precision Correct recall / labeled Vague silent
Hosted Jev documented baseline — 70–500 service range — — 97.8% at ≥0.90 — about 90.4% coverage vague titles got no tag
Laya Torch, run 1 2381.8 / 521.6 531.4 / 716.4 2826.1 126.648 / 2.184 9/9 (100%) 2/6 (33.3%) 11/52 (21.15%) 6/6
Laya FP32 ONNX 3794.4 / 643.4 464.4 / 740.9 1939.0 114.846 / 1.980 9/9 (100%) 2/6 (33.3%) 11/52 (21.15%) 6/6
Laya Torch, repeat for cascade rows 2793.4 / 657.2 570.4 / 902.2 2779.9 141.235 / 2.435 9/9 (100%) 2/6 (33.3%) 11/52 (21.15%) 6/6
Clef-flash mixed INT4/INT8 OpenVINO Load completed; first call killed — — — — — — —

Every Laya run missed the fixed 500 ms warm-p95 limit. The ONNX path reduced median latency and RSS, but not p95. Torch and ONNX returned the same tag and action on all 58 cases. Prefix caching is not supported by Laya's bidirectional encoder; the Clef conversion loader also encodes each title per call and provides no prefix-cache path.

The Laya checkpoint warned that its choice:11+ temperature (0.10058) is outside the supported [0.5, 5] range and was clamped to 0.5. The 31-option request uses this bucket, so its confidence is uncalibrated. The observed 9/9 apply precision is a small sample and is not sufficient calibration evidence. Laya is background-only under the fixed latency rule; its observed recall/coverage is far below the hosted baseline.

Clef-flash and two-tier verdict

The tested OpenVINO conversion was pinned to f244f287938ecb92cd72efbe9109db7a1648bf87; it contains a 6.32 GB mixed INT4/INT8 backbone and a 243.5 MB joint decision head. Its maintainer estimates about 8.5 GB peak RAM. On the 4-core i9-9900K VM (7.7 GiB RAM, no swap), the model loaded inside a 7 GiB systemd memory scope. The process was killed with exit 137 before the first decision, so there is no Clef accuracy, latency, RSS, or cascade score. I kept the cap to protect the shared VM. A repeat needs a host with at least 16 GiB RAM.

For a later cascade evaluation, the simplest interpretation is to keep Laya apply decisions and route its suggestions and silent decisions to Clef (49/58 cases, 84.5%). The current Laya confidence warning means those apply decisions should also be treated as unvalidated until calibrated on separate data. Since Clef returned no decisions, no end-to-end two-tier result is claimed. Current verdict: do not ship an interactive self-hosted replacement or claim a validated cascade from this VM. Re-run Clef-flash on a larger host; then score the cascade and calibration on held-out data. The optional 27B references were not downloaded or tested because this VM has 7.7 GiB RAM; the published pplx-decider-v1-27b and Qwen3.8-27B source repos are 52.2 GB and 55.6 GB respectively.

Other decisions and licenses

No mail corpus was present in the VM's private staging area. The available local document-label aggregate has 614 records and no document_type labels. The owned code has mail-category rules but no statement-document detector. Mail categories, statement detection, and post-OCR document-type accuracy could not be scored. No private document or mail content was printed or copied to the VM.

The checked model weights and runtimes are permissive and AGPL-compatible: Laya model and Laya 0.3.23 package are Apache-2.0; ModernBERT-large is Apache-2.0. The Laya paths used PyTorch 2.11.0+cpu (BSD-3-Clause) and ONNX Runtime 1.30.0 (MIT). Clef-flash and Qwen3.5-9B, plus the OpenVINO conversion, are Apache-2.0; OpenVINO 2026.4.1 is Apache-2.0, and PyTorch is BSD-3-Clause. The untested 27B weights are Apache-2.0 and their Qwen base is Apache-2.0.

Runtime versions: Laya 0.3.23, PyTorch 2.11.0+cpu, ONNX Runtime 1.30.0, OpenVINO 2026.4.1, Transformers 5.10.2.

Files, gates, and decisions

No repository files changed. Measurement helpers/results: ~/calternal-private/decider-656/{prepare_golden.ts,eval_laya.py,export_onnx.py,eval_clef.py,fetch_clef.py,golden.parsed.json,laya.*.results.json}. The downloaded models, virtualenv, and run artifacts were deleted from the perf VM.

Rust and web gates were not run because this job changed no product files. cargo clean output:

     Removed 1 file, 356B total

Decisions not specified by the design: use Laya's official FP32 ONNX export as the CPU comparison path, not INT8 because the exporter documents decision drift; count warm percentiles after the first inference and show model-load time separately; for a future cascade, keep only Laya apply and route suggestions/silence to Clef; cap the Clef process at 7 GiB on this host. These choices do not validate product behavior.

## Final report — measurement only No product code changed. The worktree is clean on `job/decider-656`; head SHA: `c4a61e8cf090170f35b1bed3350d9de20c83ecd5`. No commit was needed. Model files and staging data were removed from the perf VM after measurement. Local helpers and aggregate-only outputs remain under `~/calternal-private/decider-656/`. ### Jev golden set The original parser produced 80 titles: 58 eligible, 52 labeled, 6 vague, and 22 skipped by the existing eligibility rule. Each eligible title was scored as one request with 30 tags plus `none`. Cold is the first decision after model load; warm percentiles cover the other 57 calls. CPU figures are process CPU time during inference, divided by 58. | Model/runtime | Load ms / first decision ms | Warm p50 / p95 ms | Peak RSS MiB | CPU total / per decision s | Apply precision | Suggest precision | Correct recall / labeled | Vague silent | |---|---:|---:|---:|---:|---:|---:|---:|---:| | Hosted Jev documented baseline | — | 70–500 service range | — | — | 97.8% at ≥0.90 | — | about 90.4% coverage | vague titles got no tag | | Laya Torch, run 1 | 2381.8 / 521.6 | 531.4 / 716.4 | 2826.1 | 126.648 / 2.184 | 9/9 (100%) | 2/6 (33.3%) | 11/52 (21.15%) | 6/6 | | Laya FP32 ONNX | 3794.4 / 643.4 | 464.4 / 740.9 | 1939.0 | 114.846 / 1.980 | 9/9 (100%) | 2/6 (33.3%) | 11/52 (21.15%) | 6/6 | | Laya Torch, repeat for cascade rows | 2793.4 / 657.2 | 570.4 / 902.2 | 2779.9 | 141.235 / 2.435 | 9/9 (100%) | 2/6 (33.3%) | 11/52 (21.15%) | 6/6 | | Clef-flash mixed INT4/INT8 OpenVINO | Load completed; first call killed | — | — | — | — | — | — | — | Every Laya run missed the fixed 500 ms warm-p95 limit. The ONNX path reduced median latency and RSS, but not p95. Torch and ONNX returned the same tag and action on all 58 cases. Prefix caching is not supported by Laya's bidirectional encoder; the Clef conversion loader also encodes each title per call and provides no prefix-cache path. The Laya checkpoint warned that its `choice:11+` temperature (0.10058) is outside the supported `[0.5, 5]` range and was clamped to 0.5. The 31-option request uses this bucket, so its confidence is uncalibrated. The observed 9/9 apply precision is a small sample and is not sufficient calibration evidence. Laya is background-only under the fixed latency rule; its observed recall/coverage is far below the hosted baseline. ### Clef-flash and two-tier verdict The tested OpenVINO conversion was pinned to `f244f287938ecb92cd72efbe9109db7a1648bf87`; it contains a 6.32 GB mixed INT4/INT8 backbone and a 243.5 MB joint decision head. Its maintainer estimates about 8.5 GB peak RAM. On the 4-core i9-9900K VM (7.7 GiB RAM, no swap), the model loaded inside a 7 GiB systemd memory scope. The process was killed with exit 137 before the first decision, so there is no Clef accuracy, latency, RSS, or cascade score. I kept the cap to protect the shared VM. A repeat needs a host with at least 16 GiB RAM. For a later cascade evaluation, the simplest interpretation is to keep Laya `apply` decisions and route its suggestions and silent decisions to Clef (49/58 cases, 84.5%). The current Laya confidence warning means those apply decisions should also be treated as unvalidated until calibrated on separate data. Since Clef returned no decisions, no end-to-end two-tier result is claimed. **Current verdict: do not ship an interactive self-hosted replacement or claim a validated cascade from this VM.** Re-run Clef-flash on a larger host; then score the cascade and calibration on held-out data. The optional 27B references were not downloaded or tested because this VM has 7.7 GiB RAM; the published `pplx-decider-v1-27b` and Qwen3.8-27B source repos are 52.2 GB and 55.6 GB respectively. ### Other decisions and licenses No mail corpus was present in the VM's private staging area. The available local document-label aggregate has 614 records and no `document_type` labels. The owned code has mail-category rules but no statement-document detector. Mail categories, statement detection, and post-OCR document-type accuracy could not be scored. No private document or mail content was printed or copied to the VM. The checked model weights and runtimes are permissive and AGPL-compatible: [Laya model](https://huggingface.co/convaiinnovations/laya) and [Laya 0.3.23 package](https://pypi.org/project/laya/0.3.23/) are Apache-2.0; ModernBERT-large is [Apache-2.0](https://huggingface.co/answerdotai/ModernBERT-large). The Laya paths used PyTorch 2.11.0+cpu (BSD-3-Clause) and ONNX Runtime 1.30.0 (MIT). Clef-flash and [Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B), plus the [OpenVINO conversion](https://huggingface.co/meossistant/clef-flash-openvino), are Apache-2.0; OpenVINO 2026.4.1 is [Apache-2.0](https://pypi.org/project/openvino/2026.4.1/), and PyTorch is BSD-3-Clause. The untested 27B weights are [Apache-2.0](https://huggingface.co/perplexity-ai/pplx-decider-v1-27b) and their Qwen base is [Apache-2.0](https://huggingface.co/Qwen/Qwen3.8-27B). Runtime versions: Laya 0.3.23, PyTorch 2.11.0+cpu, ONNX Runtime 1.30.0, OpenVINO 2026.4.1, Transformers 5.10.2. ### Files, gates, and decisions No repository files changed. Measurement helpers/results: `~/calternal-private/decider-656/{prepare_golden.ts,eval_laya.py,export_onnx.py,eval_clef.py,fetch_clef.py,golden.parsed.json,laya.*.results.json}`. The downloaded models, virtualenv, and run artifacts were deleted from the perf VM. Rust and web gates were not run because this job changed no product files. `cargo clean` output: ```text Removed 1 file, 356B total ``` Decisions not specified by the design: use Laya's official FP32 ONNX export as the CPU comparison path, not INT8 because the exporter documents decision drift; count warm percentiles after the first inference and show model-load time separately; for a future cascade, keep only Laya `apply` and route suggestions/silence to Clef; cap the Clef process at 7 GiB on this host. These choices do not validate product behavior.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#656
No description provided.