Documents: local recognition research on the owner's corpus (OCR, calibrated classifier, Laya, doc-vs-photo detector) #417

Open
opened 2026-09-29 09:09:23 +00:00 by kayg · 42 comments
Owner

Part of #341 (Documents). Owner decisions, 2026-09-29:

  • Recognition runs locally only. "Find a solid local way to do that, even write our own if it's not too much work, and test it on my docs."
  • Q6: auto-apply a value only above a per-field confidence bar, calibrated on the User's own corrections. The target is about 95% precision per User. Each rejection raises that tag's bar. On a new account, only high-certainty signals apply.
  • Q7: below the bar, the value is a quiet suggestion in the inspector (Cmd+I). Accepting it trains the model.
  • Q8: a photo of paper is detected from OCR text coverage, a paper rectangle filling the frame, a CLIP zero-shot label and source hints. When in doubt, the image stays in Photos.
  • Q9: screenshots go to Photos → Screenshots, except receipts, bookings and payments, which go to Documents.
  • The owner asked whether a decision model such as Jev (hosted by TypeSafe, so it does not qualify) or Laya (Apache-2.0 open weights, about 421M parameters, 1–2 GB RAM, runs on CPU; its calibration reportedly trails Jev's, ECE 0.213 vs 0.144) would help with confident decisions.

Research and prototype (evidence, not opinions)

Corpus: the owner's real documents at ~/calternal-private/docs-corpus/ (614 files, 192 MB, read-only). Privacy is a hard rule. This repo and its issue tracker are public. Never commit, post or log any document content, file name, name, amount or number from the corpus. Report aggregate metrics only. Keep all intermediate files (OCR text, features, models) under ~/calternal-private/recog-work/ (mode 700), never in the worktree. Labels: the folder path of each file. The paperless import turned tags into folders; derive the type and sender labels from it and describe the mapping in aggregate.

Measure on the corpus with k-fold cross-validation, per field (type, sender, tags):

  1. OCR: Tesseract 5 (subprocess), ocrs (pure Rust, MIT) and PaddleOCR via ONNX (Apache-2.0). For each: text quality on a sample you check by eye (report counts only), CPU seconds per page, and peak RSS on a 4-core machine. Also measure how many PDFs already have a text layer.
  2. Classifiers:
    • (a) paperless-ngx's approach (bag of words plus MLP) as the baseline;
    • (b) TF-IDF plus a linear model (logistic regression or linear SVM) with isotonic or Platt calibration, written in Rust with no heavy framework;
    • (c) Laya zero-shot, and Laya fine-tuned if cheap, on the same fields;
    • (d) CLIP embeddings, which Photos already uses, where they help the type field.
      Report: accuracy, coverage at 95% precision (the share of values the app could auto-apply), expected calibration error, training time, per-document inference time, and memory. Also simulate the cold start: the first 10, 25, 50 and 100 corrections.
  3. Extraction: document date, amounts and due dates. Compare rule-based extraction (with locale-aware parsing) and Laya. Report precision and recall on a sample you label by hand, counts only.
  4. Doc-vs-photo detector (Q8 and Q9): build the four-signal detector and measure it on the corpus images plus a set of ordinary photos. Use public test images or the repo's photo fixtures, not the owner's Photos library.
  5. Licences: every model and its weights must be compatible with AGPL-3.0 (no non-commercial weights).
  6. Runtime budget: the Instance must stay light, so test on the perf-test VM (netbird ssh --no-browser --user root 10.69.69.63: 4 vCPU, 7 GB). You may copy the corpus there under /root/private/ (mode 700) and delete it when done.

Deliverable

  • A report posted on this issue: a table per option, a recommendation for each field (OCR, classifier, extraction, detector), and an estimate of how much work "write our own" is compared with using the options above.
  • On the job branch: the prototype as a benchmark tool under bench/recog/ (code only, no corpus data), so it can be rerun. Do not build the product feature yet.
  • Also record the XMP decisions below in DESIGN.md §"Sidecars" if the orchestrator's comment on #341 confirms them.
Part of #341 (Documents). Owner decisions, 2026-09-29: - Recognition runs **locally only**. "Find a solid local way to do that, even write our own if it's not too much work, and test it on my docs." - Q6: auto-apply a value only above a per-field confidence bar, calibrated on the User's own corrections. The target is about 95% precision per User. Each rejection raises that tag's bar. On a new account, only high-certainty signals apply. - Q7: below the bar, the value is a quiet suggestion in the inspector (Cmd+I). Accepting it trains the model. - Q8: a photo of paper is detected from OCR text coverage, a paper rectangle filling the frame, a CLIP zero-shot label and source hints. When in doubt, the image stays in Photos. - Q9: screenshots go to Photos → Screenshots, except receipts, bookings and payments, which go to Documents. - The owner asked whether a decision model such as Jev (hosted by TypeSafe, so it does not qualify) or **Laya** (Apache-2.0 open weights, about 421M parameters, 1–2 GB RAM, runs on CPU; its calibration reportedly trails Jev's, ECE 0.213 vs 0.144) would help with confident decisions. ## Research and prototype (evidence, not opinions) Corpus: the owner's real documents at `~/calternal-private/docs-corpus/` (614 files, 192 MB, read-only). **Privacy is a hard rule.** This repo and its issue tracker are public. Never commit, post or log any document content, file name, name, amount or number from the corpus. Report aggregate metrics only. Keep all intermediate files (OCR text, features, models) under `~/calternal-private/recog-work/` (mode 700), never in the worktree. Labels: the folder path of each file. The paperless import turned tags into folders; derive the type and sender labels from it and describe the mapping in aggregate. Measure on the corpus with k-fold cross-validation, per field (type, sender, tags): 1. **OCR:** Tesseract 5 (subprocess), `ocrs` (pure Rust, MIT) and PaddleOCR via ONNX (Apache-2.0). For each: text quality on a sample you check by eye (report counts only), CPU seconds per page, and peak RSS on a 4-core machine. Also measure how many PDFs already have a text layer. 2. **Classifiers:** - (a) paperless-ngx's approach (bag of words plus MLP) as the baseline; - (b) TF-IDF plus a linear model (logistic regression or linear SVM) with isotonic or Platt calibration, written in Rust with no heavy framework; - (c) Laya zero-shot, and Laya fine-tuned if cheap, on the same fields; - (d) CLIP embeddings, which Photos already uses, where they help the type field. Report: accuracy, **coverage at 95% precision** (the share of values the app could auto-apply), expected calibration error, training time, per-document inference time, and memory. Also simulate the cold start: the first 10, 25, 50 and 100 corrections. 3. **Extraction:** document date, amounts and due dates. Compare rule-based extraction (with locale-aware parsing) and Laya. Report precision and recall on a sample you label by hand, counts only. 4. **Doc-vs-photo detector** (Q8 and Q9): build the four-signal detector and measure it on the corpus images plus a set of ordinary photos. Use public test images or the repo's photo fixtures, not the owner's Photos library. 5. **Licences:** every model and its weights must be compatible with AGPL-3.0 (no non-commercial weights). 6. **Runtime budget:** the Instance must stay light, so test on the perf-test VM (`netbird ssh --no-browser --user root 10.69.69.63`: 4 vCPU, 7 GB). You may copy the corpus there under `/root/private/` (mode 700) and delete it when done. ## Deliverable - A report posted on this issue: a table per option, a recommendation for each field (OCR, classifier, extraction, detector), and an estimate of how much work "write our own" is compared with using the options above. - On the job branch: the prototype as a benchmark tool under `bench/recog/` (code only, no corpus data), so it can be rerun. Do not build the product feature yet. - Also record the XMP decisions below in DESIGN.md §"Sidecars" if the orchestrator's comment on #341 confirms them.
Author
Owner

Starting on job/recog-417, based on 99ee1fc9bbb0b3dd1f2d5b35e15529d1d0b3633c. I will keep all corpus reads and intermediate results private, commit only benchmark code, and report aggregate measurements. The #341 comments currently do not confirm an XMP Sidecar decision, so I will leave DESIGN §Sidecars unchanged unless that changes.

Starting on `job/recog-417`, based on `99ee1fc9bbb0b3dd1f2d5b35e15529d1d0b3633c`. I will keep all corpus reads and intermediate results private, commit only benchmark code, and report aggregate measurements. The #341 comments currently do not confirm an XMP Sidecar decision, so I will leave DESIGN §Sidecars unchanged unless that changes.
Author
Owner

Finding: the provided private corpus has 614 files (455 PDF, 140 DOCX, 15 XLSX, 3 XLS, 1 PNG), all directly in its root. The stated folder-path label source therefore contains zero type, sender, or tag classes. I am checking for a separate private label export while building the rerunnable benchmark. I will not fabricate supervised accuracy or 95% precision coverage from unlabeled data.

Finding: the provided private corpus has 614 files (455 PDF, 140 DOCX, 15 XLSX, 3 XLS, 1 PNG), all directly in its root. The stated folder-path label source therefore contains zero type, sender, or tag classes. I am checking for a separate private label export while building the rerunnable benchmark. I will not fabricate supervised accuracy or 95% precision coverage from unlabeled data.
Author
Owner

Finding: the ocrs code is MIT/Apache-2.0, but its published .rten weight files have no explicit license in the model repository. The upstream licensing question is still open (https://github.com/robertknight/ocrs-models/issues/34). Therefore I can benchmark the CLI, but I cannot recommend bundling those weights in an AGPL-3.0-only release until the owner grants compatible terms. Paperless-ngx's current classifier source uses CountVectorizer word 1–2 grams and separate scikit-learn MLPClassifier heads for type, correspondent and tags (https://github.com/paperless-ngx/paperless-ngx/blob/dev/src/documents/classifier.py).

Finding: the `ocrs` code is MIT/Apache-2.0, but its published `.rten` weight files have no explicit license in the model repository. The upstream licensing question is still open (https://github.com/robertknight/ocrs-models/issues/34). Therefore I can benchmark the CLI, but I cannot recommend bundling those weights in an AGPL-3.0-only release until the owner grants compatible terms. Paperless-ngx's current classifier source uses CountVectorizer word 1–2 grams and separate scikit-learn MLPClassifier heads for type, correspondent and tags (https://github.com/paperless-ngx/paperless-ngx/blob/dev/src/documents/classifier.py).
Author
Owner

Finding: I ran the same deterministic 12-page PDF sample with Tesseract 5 on the 4-vCPU, 7-GB perf-test VM. All 12 pages produced nonempty text. OCR used 1.806 CPU seconds/page and 0.834 wall seconds/page; peak child-process RSS upper bound was 64,560 KiB. On the loaded build host the same sample measured 24.534 CPU seconds/page and 20.708 wall seconds/page, so VM figures are the useful runtime estimate. The PDFs were copied privately under /root/private/recog-sample and will be removed after measurements.

Finding: I ran the same deterministic 12-page PDF sample with Tesseract 5 on the 4-vCPU, 7-GB perf-test VM. All 12 pages produced nonempty text. OCR used 1.806 CPU seconds/page and 0.834 wall seconds/page; peak child-process RSS upper bound was 64,560 KiB. On the loaded build host the same sample measured 24.534 CPU seconds/page and 20.708 wall seconds/page, so VM figures are the useful runtime estimate. The PDFs were copied privately under `/root/private/recog-sample` and will be removed after measurements.
Author
Owner

Measurement correction: the initial 12-page OCR sample all had a first-page text layer. I have changed the benchmark to select equal numbers of text-layer and textless first pages. I will use the mixed sample for OCR recommendations; the earlier Tesseract numbers remain a text-layer-only reference.

Measurement correction: the initial 12-page OCR sample all had a first-page text layer. I have changed the benchmark to select equal numbers of text-layer and textless first pages. I will use the mixed sample for OCR recommendations; the earlier Tesseract numbers remain a text-layer-only reference.
Author
Owner

Finding on the initial text-layer-only sample: PaddleOCR 3.7 with PP-OCRv5 mobile ONNX models on the 4-vCPU VM found text on all 12 pages, but used 20.361 CPU seconds/page, 7.324 wall seconds/page, and 1.802 GB observed RSS. Its model load took 0.586 seconds after weights were cached. Tesseract 5 used 1.806 CPU seconds/page, 0.834 wall seconds/page, and at most 64,560 KiB child-process RSS on those same pages. I am repeating both on a 6 text-layer / 6 textless-page sample before choosing an OCR backend.

Finding on the initial text-layer-only sample: PaddleOCR 3.7 with PP-OCRv5 mobile ONNX models on the 4-vCPU VM found text on all 12 pages, but used 20.361 CPU seconds/page, 7.324 wall seconds/page, and 1.802 GB observed RSS. Its model load took 0.586 seconds after weights were cached. Tesseract 5 used 1.806 CPU seconds/page, 0.834 wall seconds/page, and at most 64,560 KiB child-process RSS on those same pages. I am repeating both on a 6 text-layer / 6 textless-page sample before choosing an OCR backend.
Author
Owner

Finding: the four-signal prototype combines OCR word-box coverage, a large quadrilateral, the existing MIT-licensed CLIP model's zero-shot image/text embeddings, and actual source provenance. On the available labeled images (1 corpus document image plus 4 CC0 Photos fixtures), with the corpus item's real document-import source hint, it classified 1/1 document and 4/4 negatives, with no errors. The rectangle signal fired on 0/5 images; the single document required OCR + CLIP + source. This sample is far too small to estimate field precision or screenshot routing, and it contains no screenshot. Without the source hint, the document stayed in Photos, as the conservative rule requires.

Finding: the four-signal prototype combines OCR word-box coverage, a large quadrilateral, the existing MIT-licensed CLIP model's zero-shot image/text embeddings, and actual source provenance. On the available labeled images (1 corpus document image plus 4 CC0 Photos fixtures), with the corpus item's real document-import source hint, it classified 1/1 document and 4/4 negatives, with no errors. The rectangle signal fired on 0/5 images; the single document required OCR + CLIP + source. This sample is far too small to estimate field precision or screenshot routing, and it contains no screenshot. Without the source hint, the document stayed in Photos, as the conservative rule requires.
Author
Owner

Finding on the mixed 6 text-layer / 6 textless-page sample, on the 4-vCPU VM: Tesseract 5 processed 12/12 pages with nonempty text. It used 25.512 CPU seconds/page and 11.790 wall seconds/page; peak child-process RSS upper bound was 64,392 KiB. Mean token recall against the six extractable PDF text layers was 0.887. Token recall is a proxy, not hand-checked OCR accuracy; the six textless pages have no ground-truth transcript. I am running the same sample through ocrs and PaddleOCR ONNX.

Finding on the mixed 6 text-layer / 6 textless-page sample, on the 4-vCPU VM: Tesseract 5 processed 12/12 pages with nonempty text. It used 25.512 CPU seconds/page and 11.790 wall seconds/page; peak child-process RSS upper bound was 64,392 KiB. Mean token recall against the six extractable PDF text layers was 0.887. Token recall is a proxy, not hand-checked OCR accuracy; the six textless pages have no ground-truth transcript. I am running the same sample through `ocrs` and PaddleOCR ONNX.
Author
Owner

Finding on the same mixed 12-page VM sample: ocrs 0.13.1 found text on 12/12, used 4.166 CPU seconds/page and 1.852 wall seconds/page, with a 528,744 KiB peak child-process RSS upper bound. Mean token recall against the six text layers was 0.881, close to Tesseract's 0.887 on this proxy. The CLI loads its models per page, so a persistent library worker may improve throughput. The published ocrs weights still lack explicit redistribution terms, which blocks bundling them in the AGPL release even if their benchmark is favorable.

Finding on the same mixed 12-page VM sample: `ocrs` 0.13.1 found text on 12/12, used 4.166 CPU seconds/page and 1.852 wall seconds/page, with a 528,744 KiB peak child-process RSS upper bound. Mean token recall against the six text layers was 0.881, close to Tesseract's 0.887 on this proxy. The CLI loads its models per page, so a persistent library worker may improve throughput. The published `ocrs` weights still lack explicit redistribution terms, which blocks bundling them in the AGPL release even if their benchmark is favorable.
Author
Owner

I found the new owner Q11 comment on #341. I recorded its confirmed Sidecar rules in DESIGN §4: standard names, hidden by default, the Hidden files setting, and carrying Sidecars with a moved or renamed parent. Q10 (which XMP fields Documents should use) remains OPEN, so I did not specify a Document XMP schema or build that behavior. Commit: 0d64b2ba.

I found the new owner Q11 comment on #341. I recorded its confirmed Sidecar rules in DESIGN §4: standard names, hidden by default, the Hidden files setting, and carrying Sidecars with a moved or renamed parent. Q10 (which XMP fields Documents should use) remains OPEN, so I did not specify a Document XMP schema or build that behavior. Commit: `0d64b2ba`.
Author
Owner

Finding on the mixed 12-page 4-vCPU VM sample: PaddleOCR 3.7 / PP-OCRv5 mobile ONNX processed 12/12 pages with nonempty text. It used 13.934 CPU seconds/page and 4.022 wall seconds/page, with 1.524 GB observed RSS and 0.753 seconds model load from cached weights. Mean token recall against the six existing PDF text layers was 0.968. The matched values are Tesseract: 25.512 CPU s/page, 11.790 wall s/page, ≤64,392 KiB RSS, recall 0.887; ocrs: 4.166 CPU s/page, 1.852 wall s/page, ≤528,744 KiB RSS, recall 0.881. This proxy says nothing about transcription correctness on the six scans without text layers. Paddle offers the strongest measured text recovery with an Apache-2.0 model, but needs one bounded background worker because its memory is much larger. ocrs is fastest but its weight license remains unresolved.

Finding on the mixed 12-page 4-vCPU VM sample: PaddleOCR 3.7 / PP-OCRv5 mobile ONNX processed 12/12 pages with nonempty text. It used 13.934 CPU seconds/page and 4.022 wall seconds/page, with 1.524 GB observed RSS and 0.753 seconds model load from cached weights. Mean token recall against the six existing PDF text layers was 0.968. The matched values are Tesseract: 25.512 CPU s/page, 11.790 wall s/page, ≤64,392 KiB RSS, recall 0.887; `ocrs`: 4.166 CPU s/page, 1.852 wall s/page, ≤528,744 KiB RSS, recall 0.881. This proxy says nothing about transcription correctness on the six scans without text layers. Paddle offers the strongest measured text recovery with an Apache-2.0 model, but needs one bounded background worker because its memory is much larger. `ocrs` is fastest but its weight license remains unresolved.
Author
Owner

Finding from the private corpus using the rule extractor on first-page PDF text and DOCX text (DMY dates, comma decimals): 579/614 files were readable by this path; 35 were unsupported or unreadable. It emitted document-date candidates for 205, due-date candidates for 4, and amount candidates for 20. These are candidate counts, not precision or recall. A hand-labeled sample is still needed before any field can auto-apply. The conservative amount rule requires a currency marker; it misses unmarked amounts by design.

Finding from the private corpus using the rule extractor on first-page PDF text and DOCX text (DMY dates, comma decimals): 579/614 files were readable by this path; 35 were unsupported or unreadable. It emitted document-date candidates for 205, due-date candidates for 4, and amount candidates for 20. These are candidate counts, not precision or recall. A hand-labeled sample is still needed before any field can auto-apply. The conservative amount rule requires a currency marker; it misses unmarked amounts by design.
Author
Owner

Perf-test VM detector round: the same 1 document image and 4 CC0 Photos fixtures gave 1 true positive, 4 true negatives, 0 errors, at 7.369 wall seconds/image after CLIP model initialization. The rectangle signal fired on none of these five images. This is a tiny functional check, not a precision estimate. The detector must run as bounded background work and should short-circuit cheap negative signals before CLIP in a product build. The benchmark keeps all four signals visible for comparison.

Perf-test VM detector round: the same 1 document image and 4 CC0 Photos fixtures gave 1 true positive, 4 true negatives, 0 errors, at 7.369 wall seconds/image after CLIP model initialization. The rectangle signal fired on none of these five images. This is a tiny functional check, not a precision estimate. The detector must run as bounded background work and should short-circuit cheap negative signals before CLIP in a product build. The benchmark keeps all four signals visible for comparison.
Author
Owner

Merged dev once at 853fcc51 before final gates. The one conflict was an unrelated Svelte selector in AccountRow.svelte. I retained dev's :global(:nth-child(2)) fix for the nested action count and made no other web edits. The recognition benchmark and confirmed Sidecar text remain intact.

Merged `dev` once at `853fcc51` before final gates. The one conflict was an unrelated Svelte selector in `AccountRow.svelte`. I retained `dev`'s `:global(:nth-child(2))` fix for the nested action count and made no other web edits. The recognition benchmark and confirmed Sidecar text remain intact.
Author
Owner

#417 report — local document recognition prototype

Branch job/recog-417; final gate output will follow in a separate comment. The code is under bench/recog/. It reads the corpus in place, puts model and OCR work under a mode-0700 private directory, and prints aggregate results only. The copies on the perf-test VM have been removed.

Corpus and labels

The private corpus has 614 files: 455 PDF, 140 DOCX, 15 XLSX, 3 XLS, and 1 PNG. All 614 files are at the corpus root. Thus the specified folder-path mapping has one constant class and yields no type, sender, or tag labels. There is no usable k-fold ground truth or correction order. I did not infer labels from names or invent accuracy, calibration, or 95%-precision coverage. A private Paperless export or hand-labeled manifest is required to complete those comparisons.

Of 455 PDFs, 396 have at least one extractable-text page. The PDF inventory reports 1,557 pages, of which 1,450 meet the text threshold. Sixteen PDFs failed inspection; the remaining 107 pages are a mix of textless and uninspected pages. The product should use existing text before OCR and queue OCR only for pages that need it.

OCR: 4-vCPU, 7-GB VM

The matched sample has 12 first pages: six with text layers and six without. Every engine returned nonempty text on all 12. Token recall compares unique OCR tokens with the six existing text layers. It is a proxy, not hand-checked transcription quality on scans.

Engine CPU s/page Wall s/page Peak RSS Token recall proxy Weight terms
Tesseract 5 25.512 11.790 ≤64,392 KiB child peak 0.887 Apache-2.0
ocrs 0.13.1 4.166 1.852 ≤528,744 KiB child peak 0.881 Unclear for published weights
PaddleOCR 3.7, PP-OCRv5 mobile ONNX 13.934 4.022 1.524 GB observed process RSS 0.968 Apache-2.0

The ocrs CLI loads models for each page, while the Paddle process stays loaded. These RSS measures use different methods. Tesseract and PaddleOCR use Apache-2.0 and Apache-2.0 Paddle detection weights and recognition weights. The ocrs code is MIT/Apache-2.0, but the published weight license remains unanswered. Do not bundle those weights yet.

OCR recommendation: use the PDF text layer first. For missing text, provisionally use PaddleOCR mobile ONNX in one bounded background worker and unload it when idle; keep Tesseract as the low-memory path. The 1.5-GB active Paddle footprint is much larger than the 148-MB server idle RSS baseline, so it requires an Instance limit and a separate worker process. Before a product choice, hand-check scan transcriptions on a private labeled sample. ocrs is the fastest measured option but cannot ship with the present weight-license evidence.

Classification, confidence, and cold start

The benchmark includes a Paperless-style CountVectorizer 1–2 gram + MLP baseline and a small Rust TF-IDF softmax regression model with Platt scaling. The Rust tool fits vocabulary, model, calibrator, and confidence bar inside training folds. It reports type and sender separately, each tag as a private binary task, and first-10/25/50/100-correction simulations. Neither model has been measured on the owner's fields because the labels are absent.

Option Type / sender / tags accuracy Coverage at 95% precision ECE Training / inference / RAM on corpus Assessment
Paperless-style bag of words + MLP Not measurable Not measurable Not measurable Not measurable Rerunnable baseline; Python/scikit-learn runtime
Rust TF-IDF + linear model + Platt Not measurable Not measurable Not measurable Not measurable Smallest local starting point once corrections exist
Laya zero-shot Not measurable Not measurable Not measurable Not measured on VM Open Apache-2.0 weights, but upstream base result is near chance on its own typed-decision benchmark
Laya fine-tuned Not measurable Not measurable Not measurable No corpus training run Needs labels and a separate held-out calibration set; CPU-only fine-tuning is not a cheap first step
Existing CLIP embeddings for image type Not measurable Not measurable Not measurable Detector runtime below Useful visual signal; insufficient labeled document images to test type
Field Usable folder labels K-fold accuracy / ECE / 95% coverage Cold start (10 / 25 / 50 / 100)
Type 0 Not measurable Not measurable
Sender 0 Not measurable Not measurable
Tags 0 Not measurable Not measurable

The Laya authors report 0.362 zero-shot accuracy for the base English checkpoint and 0.766 for a tuned checkpoint on their typed-decision benchmark; these are not corpus results. On that same benchmark, they report ECE 0.175 for the base checkpoint, 0.213 for the tuned checkpoint, and 0.144 for hosted Jev. These are not corpus calibration values; Jev does not meet the local-only rule. They report roughly 4–5 hours on two T4 GPUs for 30,000 training questions. I did not run Laya on the corpus because the folder labels contain no target values. Cold-start accuracy and coverage at 10, 25, 50, and 100 corrections are also unmeasurable on this corpus.

Classifier recommendation: start with the local Rust linear model once a private label export exists. Calibrate each field for each User from their corrections. Use a separate held-out set to measure achieved precision and coverage; raise a tag's bar after a rejection. On a new User, apply only deterministic high-certainty values. Put other values in the inspector as quiet suggestions. Compare the Paperless baseline and Laya on the same folds before selecting a product model.

Extraction and image routing

The locale-aware rule prototype reads first-page PDF text and DOCX text. It read 579/614 files; 35 were unsupported or unreadable. There is no private hand-labeled date/amount sample. Laya is a typed decision model and would need rule-generated candidates before it could choose a date or amount.

Extraction option Corpus candidates Precision / recall Runtime / RAM Assessment
Locale-aware rules 205 document dates; 4 due dates; 20 amounts Not measurable Not separately measured Conservative candidate generator; misses amounts without currency markers
Laya candidate selection Not run Not measurable Not measured on VM No evidence it improves rules without labeled candidate choices

Candidate counts are not precision or recall.

The detector uses OCR word-box coverage, a large quadrilateral, CLIP zero-shot labels, and source hints. With the real document-import hint, it routed the one corpus document image to Documents and four CC0 Photos fixtures to Photos: 1 true positive, 4 true negatives, 0 errors. It took 7.369 wall seconds/image on the VM after CLIP initialization. The rectangle signal fired on 0/5. With no source hint, the one document remained in Photos. There is no screenshot in this set, so the Screenshots exception is tested only with synthetic decision tests. Five images cannot establish precision.

Detector option Labeled set Result VM runtime Assessment
Four signals with actual import hint 1 document + 4 CC0 fixtures 1 TP, 4 TN, 0 errors 7.369 wall s/image after model load Functional check only
Four signals with source unknown Same five images 0 TP, 4 TN, 1 FN Not measured on VM Conservative Photos fallback
CLIP alone Same five images Not scored as a route Not isolated Insufficient for a routing precision claim

Extraction recommendation: show rule candidates as suggestions until dates, due dates, and amounts have hand-labeled precision/recall. Detector recommendation: keep the conservative four-signal rule and its Photos fallback; run it as bounded background work. Keep ordinary screenshots in Photos → Screenshots, and show only high-certainty receipt, booking, or payment screenshots in Documents. Classification does not move files. Build a larger public negative set and private positive set before auto-routing.

License review

Asset Evidence AGPL-3.0-only decision
Tesseract 5 code and official language data Apache-2.0 project documentation Compatible
PP-OCRv5 mobile ONNX detection and recognition weights Detection card, recognition card: Apache-2.0 Compatible
Laya code and checkpoint Apache-2.0 upstream Compatible, but not selected on accuracy evidence
Pinned Photos CLIP Existing model manifest: MIT Already approved in the repo; reuse it
ocrs code and published models Code is MIT/Apache-2.0; weight terms unanswered Do not bundle weights

The Paperless-style baseline uses scikit-learn but copies no Paperless source code. No hosted decision service or non-commercial weights were used.

Work estimate and decisions

The existing models avoid writing an OCR engine. Turning the prototype into a production Rust linear classifier with private correction storage, calibration, retraining, and per-User thresholds is roughly 2–3 engineer-weeks after labels exist. Integrating Paddle ONNX preprocessing and one bounded worker is roughly 2–4 weeks; a Tesseract subprocess path is roughly 3–5 days. A reliable multi-language OCR model from scratch would take months plus licensed training data. Hardening rule extraction and the image detector is roughly 1–3 weeks each once labeled validation sets exist. These are engineering estimates, not measured project schedules.

The design did not specify the benchmark sample. I chose six first pages with text and six without, selected by a stable path hash. The rule extractor uses the first PDF page and DOCX text. The classifier treats private JSONL order as correction order and requires at least 20 calibration predictions before a 95% bar can be set. The detector requires three signals for an unknown-source image and keeps uncertain images in Photos. None of these prototype thresholds is a product decision.

The new owner comment on #341 confirms Sidecars keep standard names, stay hidden by default, and move with their parent file. I recorded those rules in DESIGN §4. Document XMP fields remain OPEN.

Known gaps

  • A private label export is needed for valid k-fold type/sender/tag results and cold-start curves.
  • Hand-labeled scan transcripts and date/amount values are needed for OCR quality and extraction precision/recall. I did not show corpus content in logs or issue comments to do a visual review.
  • The detector sample has only one positive and four public negatives; it has no real screenshot or broad photo set.
  • Laya, Laya fine-tuning, and CLIP image-type classification were not measured on the owner's fields.
  • The ocrs model owner must state redistribution terms before those weights can ship.
## #417 report — local document recognition prototype Branch `job/recog-417`; final gate output will follow in a separate comment. The code is under `bench/recog/`. It reads the corpus in place, puts model and OCR work under a mode-0700 private directory, and prints aggregate results only. The copies on the perf-test VM have been removed. ### Corpus and labels The private corpus has 614 files: 455 PDF, 140 DOCX, 15 XLSX, 3 XLS, and 1 PNG. All 614 files are at the corpus root. Thus the specified folder-path mapping has **one constant class** and yields **no type, sender, or tag labels**. There is no usable k-fold ground truth or correction order. I did not infer labels from names or invent accuracy, calibration, or 95%-precision coverage. A private Paperless export or hand-labeled manifest is required to complete those comparisons. Of 455 PDFs, 396 have at least one extractable-text page. The PDF inventory reports 1,557 pages, of which 1,450 meet the text threshold. Sixteen PDFs failed inspection; the remaining 107 pages are a mix of textless and uninspected pages. The product should use existing text before OCR and queue OCR only for pages that need it. ### OCR: 4-vCPU, 7-GB VM The matched sample has 12 first pages: six with text layers and six without. Every engine returned nonempty text on all 12. Token recall compares unique OCR tokens with the six existing text layers. It is a proxy, **not** hand-checked transcription quality on scans. | Engine | CPU s/page | Wall s/page | Peak RSS | Token recall proxy | Weight terms | |---|---:|---:|---:|---:|---| | Tesseract 5 | 25.512 | 11.790 | ≤64,392 KiB child peak | 0.887 | Apache-2.0 | | `ocrs` 0.13.1 | 4.166 | 1.852 | ≤528,744 KiB child peak | 0.881 | **Unclear** for published weights | | PaddleOCR 3.7, PP-OCRv5 mobile ONNX | 13.934 | 4.022 | 1.524 GB observed process RSS | 0.968 | Apache-2.0 | The `ocrs` CLI loads models for each page, while the Paddle process stays loaded. These RSS measures use different methods. Tesseract and PaddleOCR use [Apache-2.0](https://tesseract-ocr.github.io/tessdoc/) and [Apache-2.0 Paddle detection weights](https://huggingface.co/PaddlePaddle/PP-OCRv5_mobile_det_onnx) and [recognition weights](https://huggingface.co/PaddlePaddle/PP-OCRv5_mobile_rec_onnx). The [`ocrs` code is MIT/Apache-2.0](https://github.com/robertknight/ocrs), but the [published weight license remains unanswered](https://github.com/robertknight/ocrs-models/issues/34). Do not bundle those weights yet. **OCR recommendation:** use the PDF text layer first. For missing text, provisionally use PaddleOCR mobile ONNX in one bounded background worker and unload it when idle; keep Tesseract as the low-memory path. The 1.5-GB active Paddle footprint is much larger than the [148-MB server idle RSS baseline](https://git.kayg.org/kayg/calternal/src/branch/dev/docs/perf/baseline.json), so it requires an Instance limit and a separate worker process. Before a product choice, hand-check scan transcriptions on a private labeled sample. `ocrs` is the fastest measured option but cannot ship with the present weight-license evidence. ### Classification, confidence, and cold start The benchmark includes a [Paperless-style CountVectorizer 1–2 gram + MLP baseline](https://github.com/paperless-ngx/paperless-ngx/blob/dev/src/documents/classifier.py) and a small Rust TF-IDF softmax regression model with Platt scaling. The Rust tool fits vocabulary, model, calibrator, and confidence bar inside training folds. It reports type and sender separately, each tag as a private binary task, and first-10/25/50/100-correction simulations. Neither model has been measured on the owner's fields because the labels are absent. | Option | Type / sender / tags accuracy | Coverage at 95% precision | ECE | Training / inference / RAM on corpus | Assessment | |---|---|---|---|---|---| | Paperless-style bag of words + MLP | Not measurable | Not measurable | Not measurable | Not measurable | Rerunnable baseline; Python/scikit-learn runtime | | Rust TF-IDF + linear model + Platt | Not measurable | Not measurable | Not measurable | Not measurable | Smallest local starting point once corrections exist | | Laya zero-shot | Not measurable | Not measurable | Not measurable | Not measured on VM | Open Apache-2.0 weights, but upstream base result is near chance on its own typed-decision benchmark | | Laya fine-tuned | Not measurable | Not measurable | Not measurable | No corpus training run | Needs labels and a separate held-out calibration set; CPU-only fine-tuning is not a cheap first step | | Existing CLIP embeddings for image type | Not measurable | Not measurable | Not measurable | Detector runtime below | Useful visual signal; insufficient labeled document images to test type | | Field | Usable folder labels | K-fold accuracy / ECE / 95% coverage | Cold start (10 / 25 / 50 / 100) | |---|---:|---|---| | Type | 0 | Not measurable | Not measurable | | Sender | 0 | Not measurable | Not measurable | | Tags | 0 | Not measurable | Not measurable | The [Laya authors](https://github.com/NandhaKishorM/laya) report 0.362 zero-shot accuracy for the base English checkpoint and 0.766 for a tuned checkpoint on **their** typed-decision benchmark; these are not corpus results. On that same benchmark, they report ECE 0.175 for the base checkpoint, 0.213 for the tuned checkpoint, and 0.144 for hosted Jev. These are not corpus calibration values; Jev does not meet the local-only rule. They report roughly 4–5 hours on two T4 GPUs for 30,000 training questions. I did not run Laya on the corpus because the folder labels contain no target values. Cold-start accuracy and coverage at 10, 25, 50, and 100 corrections are also unmeasurable on this corpus. **Classifier recommendation:** start with the local Rust linear model once a private label export exists. Calibrate each field for each User from their corrections. Use a separate held-out set to measure achieved precision and coverage; raise a tag's bar after a rejection. On a new User, apply only deterministic high-certainty values. Put other values in the inspector as quiet suggestions. Compare the Paperless baseline and Laya on the same folds before selecting a product model. ### Extraction and image routing The locale-aware rule prototype reads first-page PDF text and DOCX text. It read 579/614 files; 35 were unsupported or unreadable. There is no private hand-labeled date/amount sample. Laya is a typed decision model and would need rule-generated candidates before it could choose a date or amount. | Extraction option | Corpus candidates | Precision / recall | Runtime / RAM | Assessment | |---|---|---|---|---| | Locale-aware rules | 205 document dates; 4 due dates; 20 amounts | Not measurable | Not separately measured | Conservative candidate generator; misses amounts without currency markers | | Laya candidate selection | Not run | Not measurable | Not measured on VM | No evidence it improves rules without labeled candidate choices | Candidate counts are **not precision or recall**. The detector uses OCR word-box coverage, a large quadrilateral, CLIP zero-shot labels, and source hints. With the real document-import hint, it routed the one corpus document image to Documents and four CC0 Photos fixtures to Photos: 1 true positive, 4 true negatives, 0 errors. It took 7.369 wall seconds/image on the VM after CLIP initialization. The rectangle signal fired on 0/5. With no source hint, the one document remained in Photos. There is no screenshot in this set, so the Screenshots exception is tested only with synthetic decision tests. Five images cannot establish precision. | Detector option | Labeled set | Result | VM runtime | Assessment | |---|---|---|---|---| | Four signals with actual import hint | 1 document + 4 CC0 fixtures | 1 TP, 4 TN, 0 errors | 7.369 wall s/image after model load | Functional check only | | Four signals with source unknown | Same five images | 0 TP, 4 TN, 1 FN | Not measured on VM | Conservative Photos fallback | | CLIP alone | Same five images | Not scored as a route | Not isolated | Insufficient for a routing precision claim | **Extraction recommendation:** show rule candidates as suggestions until dates, due dates, and amounts have hand-labeled precision/recall. **Detector recommendation:** keep the conservative four-signal rule and its Photos fallback; run it as bounded background work. Keep ordinary screenshots in Photos → Screenshots, and show only high-certainty receipt, booking, or payment screenshots in Documents. Classification does not move files. Build a larger public negative set and private positive set before auto-routing. ### License review | Asset | Evidence | AGPL-3.0-only decision | |---|---|---| | Tesseract 5 code and official language data | [Apache-2.0 project documentation](https://tesseract-ocr.github.io/tessdoc/) | Compatible | | PP-OCRv5 mobile ONNX detection and recognition weights | [Detection card](https://huggingface.co/PaddlePaddle/PP-OCRv5_mobile_det_onnx), [recognition card](https://huggingface.co/PaddlePaddle/PP-OCRv5_mobile_rec_onnx): Apache-2.0 | Compatible | | Laya code and checkpoint | [Apache-2.0 upstream](https://huggingface.co/convaiinnovations/laya) | Compatible, but not selected on accuracy evidence | | Pinned Photos CLIP | [Existing model manifest](https://git.kayg.org/kayg/calternal/src/branch/dev/crates/calternal-embed/models/manifest.json): MIT | Already approved in the repo; reuse it | | `ocrs` code and published models | [Code is MIT/Apache-2.0](https://github.com/robertknight/ocrs); [weight terms unanswered](https://github.com/robertknight/ocrs-models/issues/34) | Do not bundle weights | The Paperless-style baseline uses scikit-learn but copies no Paperless source code. No hosted decision service or non-commercial weights were used. ### Work estimate and decisions The existing models avoid writing an OCR engine. Turning the prototype into a production Rust linear classifier with private correction storage, calibration, retraining, and per-User thresholds is roughly 2–3 engineer-weeks after labels exist. Integrating Paddle ONNX preprocessing and one bounded worker is roughly 2–4 weeks; a Tesseract subprocess path is roughly 3–5 days. A reliable multi-language OCR model from scratch would take months plus licensed training data. Hardening rule extraction and the image detector is roughly 1–3 weeks each once labeled validation sets exist. These are engineering estimates, not measured project schedules. The design did not specify the benchmark sample. I chose six first pages with text and six without, selected by a stable path hash. The rule extractor uses the first PDF page and DOCX text. The classifier treats private JSONL order as correction order and requires at least 20 calibration predictions before a 95% bar can be set. The detector requires three signals for an unknown-source image and keeps uncertain images in Photos. None of these prototype thresholds is a product decision. The new owner comment on #341 confirms Sidecars keep standard names, stay hidden by default, and move with their parent file. I recorded those rules in DESIGN §4. Document XMP fields remain **OPEN**. ### Known gaps - A private label export is needed for valid k-fold type/sender/tag results and cold-start curves. - Hand-labeled scan transcripts and date/amount values are needed for OCR quality and extraction precision/recall. I did not show corpus content in logs or issue comments to do a visual review. - The detector sample has only one positive and four public negatives; it has no real screenshot or broad photo set. - Laya, Laya fine-tuning, and CLIP image-type classification were not measured on the owner's fields. - The `ocrs` model owner must state redistribution terms before those weights can ship.
Author
Owner

#417 final branch and gates

Head: b1199ff5f0f11a323d045c29e349596bac785b93 on job/recog-417. Working tree clean. The branch contains the rerunnable bench/recog/ prototype and the confirmed Sidecars decision in DESIGN §4. No product recognition route or UI was built, so no API adversarial round or screenshots apply.

Full workspace Rust gates were launched once after the dev merge. On the shared host, Clippy spent over two hours compiling dependencies; cargo test --workspace stayed on the Cargo build lock. I stopped both at the job timebox, exit 130. There were no diagnostics before interruption. These two workspace gates are incomplete, not passed. Their exact final log excerpts:

    Checking zstd v0.13.3
    Checking tantivy-fst v0.5.0
    Checking dashmap v6.2.1
    Checking rand v0.10.3
    Checking smallstr v0.3.1
    Checking murmurhash32 v0.3.1
   Compiling typetag v0.2.23
    Checking unicode-bidi v0.3.18
    Blocking waiting for file lock on build directory

Changed benchmark crate gates: Clippy exit 0; Rust test exit 0. Exact output:

cargo fmt --manifest-path bench/recog/Cargo.toml --check: exit 0, no output.

    Finished `dev` profile [unoptimized + debuginfo] target(s) in 0.59s
    Finished `test` profile [unoptimized + debuginfo] target(s) in 0.25s
     Running unittests src/main.rs (/mnt/hdd/targets/jobs/recog-417/debug/deps/recog_bench-c785a733d7ec9970)

running 2 tests
test tests::softmax_probabilities_sum_to_one ... ok
test tests::vectorizer_uses_training_terms_only ... ok

test result: ok. 2 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s

Root cargo fmt --check: exit 0, no output. bun run check and bun run test: exit 0. Exact output:

$ node scripts/check-type-tokens.mjs && svelte-kit sync && svelte-check --tsconfig ./tsconfig.json
Text sizes use shared role tokens.
Loading svelte-check in workspace: /home/kayg/Developer/calternal-wt/recog-417/apps/web
Getting Svelte diagnostics...

svelte-check found 0 errors and 0 warnings
 Test Files  125 passed (125)
      Tests  800 passed (800)
   Start at  12:01:30
   Duration  107.37s (transform 54%, environment 17%, import 15%, tests 10%, setup 4%)

Python tests: exit 0. Exact output:

.......
----------------------------------------------------------------------
Ran 7 tests in 0.001s

OK

Cleanup, exact output:

     Removed 6910 files, 2.2GiB total

apps/web/.svelte-kit was deleted. The perf-VM corpus and model copies were deleted after measurement. The remaining gate work for the orchestrator is the full workspace Clippy and Cargo test on a quieter or larger build host; the branch is ready for that review.

## #417 final branch and gates Head: `b1199ff5f0f11a323d045c29e349596bac785b93` on `job/recog-417`. Working tree clean. The branch contains the rerunnable `bench/recog/` prototype and the confirmed Sidecars decision in DESIGN §4. No product recognition route or UI was built, so no API adversarial round or screenshots apply. Full workspace Rust gates were launched once after the `dev` merge. On the shared host, Clippy spent over two hours compiling dependencies; `cargo test --workspace` stayed on the Cargo build lock. I stopped both at the job timebox, exit 130. There were no diagnostics before interruption. **These two workspace gates are incomplete**, not passed. Their exact final log excerpts: ```text Checking zstd v0.13.3 Checking tantivy-fst v0.5.0 Checking dashmap v6.2.1 Checking rand v0.10.3 Checking smallstr v0.3.1 Checking murmurhash32 v0.3.1 Compiling typetag v0.2.23 Checking unicode-bidi v0.3.18 ``` ```text Blocking waiting for file lock on build directory ``` Changed benchmark crate gates: Clippy exit 0; Rust test exit 0. Exact output: `cargo fmt --manifest-path bench/recog/Cargo.toml --check`: exit 0, no output. ```text Finished `dev` profile [unoptimized + debuginfo] target(s) in 0.59s ``` ```text Finished `test` profile [unoptimized + debuginfo] target(s) in 0.25s Running unittests src/main.rs (/mnt/hdd/targets/jobs/recog-417/debug/deps/recog_bench-c785a733d7ec9970) running 2 tests test tests::softmax_probabilities_sum_to_one ... ok test tests::vectorizer_uses_training_terms_only ... ok test result: ok. 2 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s ``` Root `cargo fmt --check`: exit 0, no output. `bun run check` and `bun run test`: exit 0. Exact output: ```text $ node scripts/check-type-tokens.mjs && svelte-kit sync && svelte-check --tsconfig ./tsconfig.json Text sizes use shared role tokens. Loading svelte-check in workspace: /home/kayg/Developer/calternal-wt/recog-417/apps/web Getting Svelte diagnostics... svelte-check found 0 errors and 0 warnings ``` ```text Test Files 125 passed (125) Tests 800 passed (800) Start at 12:01:30 Duration 107.37s (transform 54%, environment 17%, import 15%, tests 10%, setup 4%) ``` Python tests: exit 0. Exact output: ```text ....... ---------------------------------------------------------------------- Ran 7 tests in 0.001s OK ``` Cleanup, exact output: ```text Removed 6910 files, 2.2GiB total ``` `apps/web/.svelte-kit` was deleted. The perf-VM corpus and model copies were deleted after measurement. The remaining gate work for the orchestrator is the full workspace Clippy and Cargo test on a quieter or larger build host; the branch is ready for that review.
Author
Owner

Resuming round 2 on job/recog-417 from b1199ff5f0f11a323d045c29e349596bac785b93 (current dev: 191b179baac3ef4f5bebfe07ce91c4b7a887ace2). I will evaluate pdf-inspector and current local layout-aware OCR on the 4-vCPU perf-test VM, keep corpus artifacts under the private directory, and report aggregate results only. The classifier remains unchanged while R3 labels are pending.

Resuming round 2 on `job/recog-417` from `b1199ff5f0f11a323d045c29e349596bac785b93` (current `dev`: `191b179baac3ef4f5bebfe07ce91c4b7a887ace2`). I will evaluate pdf-inspector and current local layout-aware OCR on the 4-vCPU perf-test VM, keep corpus artifacts under the private directory, and report aggregate results only. The classifier remains unchanged while R3 labels are pending.
Author
Owner

Round-2 finding on the 4-vCPU, 7-GB perf-test VM: pdf-inspector 1.25.2 processed 439/455 PDFs; the same 16 failed the prior Poppler inspection. Against the prior 20-byte-per-page text-layer threshold, 1,557 pages yielded 1,426 correct text routes, 102 correct OCR routes, 24 text pages sent to OCR, and 5 textless pages missed. Full Markdown extraction across these PDFs took 14.443 CPU s and 14.528 wall s (0.0093 s/page), with a 105,880 KiB peak child-process RSS upper bound. Across 383 PDFs with an extractable text layer and nonempty Markdown, mean unique-token recall against Poppler text was 0.9909. This is a text-layer proxy, not a scan transcription score; the five missed pages require a fallback check before product routing. The rerunnable aggregate-only benchmark is commit 662de426 under bench/recog/.

Round-2 finding on the 4-vCPU, 7-GB perf-test VM: pdf-inspector 1.25.2 processed 439/455 PDFs; the same 16 failed the prior Poppler inspection. Against the prior 20-byte-per-page text-layer threshold, 1,557 pages yielded 1,426 correct text routes, 102 correct OCR routes, 24 text pages sent to OCR, and 5 textless pages missed. Full Markdown extraction across these PDFs took 14.443 CPU s and 14.528 wall s (0.0093 s/page), with a 105,880 KiB peak child-process RSS upper bound. Across 383 PDFs with an extractable text layer and nonempty Markdown, mean unique-token recall against Poppler text was 0.9909. This is a text-layer proxy, not a scan transcription score; the five missed pages require a fallback check before product routing. The rerunnable aggregate-only benchmark is commit `662de426` under `bench/recog/`.
Author
Owner

Round-2 finding on the 4-vCPU VM: Docling.rs 1.74.1 (MIT code, permissively licensed ONNX assets) processed the matched 12 one-page PDFs with one warm worker in 83.08 CPU s and 38.28 wall s (6.92 CPU and 3.19 wall s/page), peak RSS 3,355,964 KiB. It returned nonempty text on 12/12. Against the six text layers, mean unique-token recall was 0.9232 and adjacent-token recall was 0.8410; it emitted Markdown table syntax on 5/12 pages and no HTML tables. These are proxies, not hand-labeled scan or table accuracy. The ongoing 107-textless-page run has already reached about 6 GB RSS, so the final recommendation must use its worst-case memory, not the mixed sample alone. Granite-Docling produced only an embedded image and zero text on one textless page (42.67 wall s, 2.81 GB peak RSS). SmolDocling did not finish a warm one-page conversion within 75 wall seconds, so it fails this VM's 60 s/page screen.

License correction: the MinerU 2.5 model card marks the weights AGPL-3.0, but the current upstream MinerU code LICENSE.md adds commercial-use conditions to Apache-2.0. That current code is not an acceptable dependency for this AGPL-3.0-only repository. I am recording this distinction rather than relying on the older code-license description.

Round-2 finding on the 4-vCPU VM: Docling.rs 1.74.1 (MIT code, permissively licensed ONNX assets) processed the matched 12 one-page PDFs with one warm worker in 83.08 CPU s and 38.28 wall s (6.92 CPU and 3.19 wall s/page), peak RSS 3,355,964 KiB. It returned nonempty text on 12/12. Against the six text layers, mean unique-token recall was 0.9232 and adjacent-token recall was 0.8410; it emitted Markdown table syntax on 5/12 pages and no HTML tables. These are proxies, not hand-labeled scan or table accuracy. The ongoing 107-textless-page run has already reached about 6 GB RSS, so the final recommendation must use its worst-case memory, not the mixed sample alone. Granite-Docling produced only an embedded image and zero text on one textless page (42.67 wall s, 2.81 GB peak RSS). SmolDocling did not finish a warm one-page conversion within 75 wall seconds, so it fails this VM's 60 s/page screen. License correction: the MinerU 2.5 model card marks the weights AGPL-3.0, but the current upstream MinerU code LICENSE.md adds commercial-use conditions to Apache-2.0. That current code is not an acceptable dependency for this AGPL-3.0-only repository. I am recording this distinction rather than relying on the older code-license description.
Author
Owner

Docling.rs corpus stress finding: a continuous one-worker run retained about 6.7 GB RSS by 70/107 textless pages, so I stopped it before memory exhaustion on the 7.7-GB VM. Restarting the process in 12-page batches reduced the first batch's peak RSS to 1.99 GB. The second batch peaked at 5.08 GB and converted 11/12 pages; it rejected one PDF page whose declared render size exceeded its built-in per-side safety cap. I am keeping that cap and counting the page as a fallback case. This shows why a bounded, recyclable worker and a lower-resolution OCR fallback are needed for extreme page geometry. No corpus path or content is in this comment.

Docling.rs corpus stress finding: a continuous one-worker run retained about 6.7 GB RSS by 70/107 textless pages, so I stopped it before memory exhaustion on the 7.7-GB VM. Restarting the process in 12-page batches reduced the first batch's peak RSS to 1.99 GB. The second batch peaked at 5.08 GB and converted 11/12 pages; it rejected one PDF page whose declared render size exceeded its built-in per-side safety cap. I am keeping that cap and counting the page as a fallback case. This shows why a bounded, recyclable worker and a lower-resolution OCR fallback are needed for extreme page geometry. No corpus path or content is in this comment.
Author
Owner

Completed the 107 textless-page Docling.rs CPU run on the 4-vCPU VM with a fresh process every 12 pages. Nine batches took 1,270.02 CPU s and 617.75 wall s in total (11.87 CPU and 5.77 wall s per input page). Peak batch RSS was 5,078,068 KiB. The highest successful per-page wall time was 25.1 s; none exceeded 60 s. It produced Markdown for 106/107 pages and nonempty text for 95/106 outputs. Twenty-four outputs had Markdown table syntax and none had HTML table syntax. One page hit the parser's render-size safety cap; it needs a lower-resolution OCR fallback. Because the scans lack hand transcripts, nonempty text and table syntax are coverage checks, not accuracy scores. The earlier continuous run retained about 6.7 GB RSS, so process recycling is required on this VM.

Completed the 107 textless-page Docling.rs CPU run on the 4-vCPU VM with a fresh process every 12 pages. Nine batches took 1,270.02 CPU s and 617.75 wall s in total (11.87 CPU and 5.77 wall s per input page). Peak batch RSS was 5,078,068 KiB. The highest successful per-page wall time was 25.1 s; none exceeded 60 s. It produced Markdown for 106/107 pages and nonempty text for 95/106 outputs. Twenty-four outputs had Markdown table syntax and none had HTML table syntax. One page hit the parser's render-size safety cap; it needs a lower-resolution OCR fallback. Because the scans lack hand transcripts, nonempty text and table syntax are coverage checks, not accuracy scores. The earlier continuous run retained about 6.7 GB RSS, so process recycling is required on this VM.
Author
Owner

Final pdf-inspector 1.25.2 routing measurement on the 4-vCPU VM: classification-only for 439 readable PDFs took 3.296 CPU s and 3.424 wall s. Full Markdown extraction took 14.535 CPU s and 14.143 wall s for 1,557 pages, peak child-process RSS ≤106,448 KiB. At PDF level it identified 383/396 text-bearing PDFs and called 0/43 textless PDFs text-bearing. At page level it kept 1,426/1,450 text-bearing pages on the text path, routed 102/107 textless pages to OCR, sent 24 text-bearing pages to OCR, and missed 5 textless pages. On 383 text-bearing PDFs with nonempty Markdown, mean unique-token recall against Poppler text was 0.9909 and adjacent-token recall was 0.9267. Markdown had table syntax in 285 PDFs, but table-cell correctness is unverified. The default Rust crate has MIT code and 86 normal dependencies; the checked dependency licenses contained no GPL-only, non-commercial, or unspecified license. The API accepts PDF bytes in memory, which can come from calternal-fs without constructing a user path.

Final pdf-inspector 1.25.2 routing measurement on the 4-vCPU VM: classification-only for 439 readable PDFs took 3.296 CPU s and 3.424 wall s. Full Markdown extraction took 14.535 CPU s and 14.143 wall s for 1,557 pages, peak child-process RSS ≤106,448 KiB. At PDF level it identified 383/396 text-bearing PDFs and called 0/43 textless PDFs text-bearing. At page level it kept 1,426/1,450 text-bearing pages on the text path, routed 102/107 textless pages to OCR, sent 24 text-bearing pages to OCR, and missed 5 textless pages. On 383 text-bearing PDFs with nonempty Markdown, mean unique-token recall against Poppler text was 0.9909 and adjacent-token recall was 0.9267. Markdown had table syntax in 285 PDFs, but table-cell correctness is unverified. The default Rust crate has MIT code and 86 normal dependencies; the checked dependency licenses contained no GPL-only, non-commercial, or unspecified license. The API accepts PDF bytes in memory, which can come from calternal-fs without constructing a user path.
Author
Owner

PP-StructureV3 3.7.0 / PaddlePaddle CPU 3.3.1 finding on the 4-vCPU VM: its full layout/table pipeline loaded about 1.7 GB of official model assets but failed inside the Paddle static CPU predictor before returning one page. With oneDNN enabled it raised NotImplementedError. With oneDNN disabled it raised MemoryError at 5,958,472 KiB peak RSS. I then kept text/layout/table recognition, disabled formula/chart/seal models, and rendered the same page to a 1600-px maximum side; it still raised MemoryError at 4,952,736 KiB peak RSS after 35.27 wall seconds. No text or table output was produced, so there is no valid corpus recall or seconds-per-successful-page figure. This configuration is not a usable worker on this VM. The sampled official model cards checked so far declare Apache-2.0; I will record the complete model/license list in the report.

PP-StructureV3 3.7.0 / PaddlePaddle CPU 3.3.1 finding on the 4-vCPU VM: its full layout/table pipeline loaded about 1.7 GB of official model assets but failed inside the Paddle static CPU predictor before returning one page. With oneDNN enabled it raised NotImplementedError. With oneDNN disabled it raised MemoryError at 5,958,472 KiB peak RSS. I then kept text/layout/table recognition, disabled formula/chart/seal models, and rendered the same page to a 1600-px maximum side; it still raised MemoryError at 4,952,736 KiB peak RSS after 35.27 wall seconds. No text or table output was produced, so there is no valid corpus recall or seconds-per-successful-page figure. This configuration is not a usable worker on this VM. The sampled official model cards checked so far declare Apache-2.0; I will record the complete model/license list in the report.
Author
Owner

Round 2 result

Branch job/recog-417 has been merged with local dev; head SHA and final gate output follow below. The aggregate-only report is docs/research/document-recognition-417.md. Benchmark code is in bench/recog/.

Route / extractor Corpus result Time and memory Quality proxy
pdf-inspector 1.25.2 (MIT) 439/455 PDFs inspected; 1,426/1,450 text pages kept; 102/107 textless pages sent to OCR; 5 textless pages missed Detect 3.296 CPU / 3.424 wall s; Markdown extraction 14.535 CPU / 14.143 wall s; ≤106,448 KiB RSS Mean unique-token recall 0.9909; adjacent-token recall 0.9267 vs Poppler on 383 text PDFs; Markdown table syntax in 285 PDFs
Docling.rs 1.74.1 (MIT; permissive weights) Mixed 12/12 nonempty; textless 106/107 outputs and 95 nonempty; 24 Markdown-table pages Mixed 83.08 CPU / 38.28 wall s, 3.36 GB peak RSS; textless 107: 1,270.02 CPU / 617.75 wall s, 5.08 GB peak RSS in 12-page fresh-process batches Mixed text-layer token recall 0.923; adjacent-token recall 0.841; table syntax is not a correctness score
Other CPU model Result on same VM
Granite-Docling-258M (Apache-2.0 weights) An older Docling run used 65.91 CPU / 42.67 wall s and 2.81 GB RSS on one scan but produced no text; current Docling timed out after 90 wall s without output.
SmolDocling-256M-preview (CDLA-Permissive-2.0 weights) No output by 180 s initial and 75 s warm limits; current warm run used 89.68 CPU s, 75.06 wall s, 1.75 GB RSS.
PP-StructureV3 (Apache-2.0 code/models checked) oneDNN failure; then MemoryError even after disabling formula, chart, and seal models and limiting rendered side to 1,600 px (4.95 GB peak RSS). No page completed.
PaddleOCR-VL 1.6, 0.9B (Apache-2.0 weights/code) 61.68 s model load, then OOM kill at 6.88 GB RSS before one page completed.
MinerU 2.5 (AGPL-3.0 weights) Current upstream code has extra commercial-use conditions; excluded under the repository dependency rule.

pdf-inspector is a local Rust crate with an in-memory API and 86 normal dependencies in the checked default closure. Its default build excludes OCR and PDFium. AnyDoc is MIT, runs office/PDF conversion locally, and depends on pdf-inspector. Its optional hosted OCR sends the PDF to Firecrawl Parse, so it is outside the local-only path.

Decision: use pdf-inspector for routing and the PDF text layer for text-bearing pages. Retry empty extracted text even when routing says text. Use one Docling.rs background worker for textless pages, unload it when idle, recycle after at most 12 pages, cap it at 6 GiB RSS and 60 s per page. Use the previously measured PP-OCRv5 mobile ONNX path at lower resolution for the one render-cap failure and 11 blank pages. The measured PDF work plus estimated fallback is about 12–15 minutes for the 439 readable PDFs on the VM. The 16 unreadable PDFs and office extraction are outside that estimate. No transcripts or table/reading-order gold labels exist, so quality remains a proxy. The type/sender/tag classifier is unchanged pending the private R3 Paperless-ngx label export.

Known gaps: no product recognition path was implemented; no valid CPU throughput/quality figures exist for models that failed or exceeded the one-page time limit. All staged corpus and model copies were removed from the VM. No per-document data is in this report.

Gates

Head SHA: 31a523176c2b91bbd1bac006732b36e1fb16a359.

cargo fmt --check: exit 0, no output.

cargo clippy --manifest-path bench/recog/Cargo.toml --all-targets -- -D warnings (exit 0):

    Checking serde_core v1.0.229
    Checking zmij v1.0.23
    Checking itoa v1.0.18
    Checking memchr v2.8.3
    Checking serde v1.0.229
    Checking serde_json v1.0.151
    Checking recog-bench v0.1.0 (/home/kayg/Developer/calternal-wt/recog-417/bench/recog)
    Finished `dev` profile [unoptimized + debuginfo] target(s) in 4m 00s

cargo test --manifest-path bench/recog/Cargo.toml (exit 0):

   Compiling recog-bench v0.1.0 (/home/kayg/Developer/calternal-wt/recog-417/bench/recog)
    Finished `test` profile [unoptimized + debuginfo] target(s) in 3m 15s
     Running unittests src/main.rs (/mnt/hdd/targets/jobs/recog-417/debug/deps/recog_bench-c785a733d7ec9970)

running 2 tests
test tests::vectorizer_uses_training_terms_only ... ok
test tests::softmax_probabilities_sum_to_one ... ok

test result: ok. 2 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.01s

bun run check in apps/web (exit 0):

$ node scripts/check-type-tokens.mjs && svelte-kit sync && svelte-check --tsconfig ./tsconfig.json
Text sizes use shared role tokens.
Loading svelte-check in workspace: /home/kayg/Developer/calternal-wt/recog-417/apps/web
Getting Svelte diagnostics...

svelte-check found 0 errors and 0 warnings

bun run test in apps/web (exit 0; final test summary):

 Test Files  126 passed (126)
      Tests  801 passed (801)
   Start at  15:25:00
   Duration  102.61s (transform 55%, environment 17%, import 15%, tests 9%, setup 3%)

python3 -m py_compile bench/recog/*.py and the local privacy/extraction tests passed. Full local Python discovery could not import numpy for the earlier detector test because model dependencies are installed in the private VM environment. This does not affect the new routing/layout scripts. The latest host rule requires Rust gates per touched crate; the orchestrator runs full-workspace gates at merge time. No API was changed, so the adversarial API round did not apply. cargo clean removed 214 files, 105.1 MiB; web build output was removed. No push, deployment, or merge into dev was made.

## Round 2 result Branch `job/recog-417` has been merged with local `dev`; head SHA and final gate output follow below. The aggregate-only report is `docs/research/document-recognition-417.md`. Benchmark code is in `bench/recog/`. | Route / extractor | Corpus result | Time and memory | Quality proxy | |---|---:|---:|---| | pdf-inspector 1.25.2 (MIT) | 439/455 PDFs inspected; 1,426/1,450 text pages kept; 102/107 textless pages sent to OCR; 5 textless pages missed | Detect 3.296 CPU / 3.424 wall s; Markdown extraction 14.535 CPU / 14.143 wall s; ≤106,448 KiB RSS | Mean unique-token recall 0.9909; adjacent-token recall 0.9267 vs Poppler on 383 text PDFs; Markdown table syntax in 285 PDFs | | Docling.rs 1.74.1 (MIT; permissive weights) | Mixed 12/12 nonempty; textless 106/107 outputs and 95 nonempty; 24 Markdown-table pages | Mixed 83.08 CPU / 38.28 wall s, 3.36 GB peak RSS; textless 107: 1,270.02 CPU / 617.75 wall s, 5.08 GB peak RSS in 12-page fresh-process batches | Mixed text-layer token recall 0.923; adjacent-token recall 0.841; table syntax is not a correctness score | | Other CPU model | Result on same VM | |---|---| | Granite-Docling-258M (Apache-2.0 weights) | An older Docling run used 65.91 CPU / 42.67 wall s and 2.81 GB RSS on one scan but produced no text; current Docling timed out after 90 wall s without output. | | SmolDocling-256M-preview (CDLA-Permissive-2.0 weights) | No output by 180 s initial and 75 s warm limits; current warm run used 89.68 CPU s, 75.06 wall s, 1.75 GB RSS. | | PP-StructureV3 (Apache-2.0 code/models checked) | oneDNN failure; then `MemoryError` even after disabling formula, chart, and seal models and limiting rendered side to 1,600 px (4.95 GB peak RSS). No page completed. | | PaddleOCR-VL 1.6, 0.9B (Apache-2.0 weights/code) | 61.68 s model load, then OOM kill at 6.88 GB RSS before one page completed. | | MinerU 2.5 (AGPL-3.0 weights) | Current upstream code has extra commercial-use conditions; excluded under the repository dependency rule. | pdf-inspector is a local Rust crate with an in-memory API and 86 normal dependencies in the checked default closure. Its default build excludes OCR and PDFium. AnyDoc is MIT, runs office/PDF conversion locally, and depends on pdf-inspector. Its optional hosted OCR sends the PDF to Firecrawl Parse, so it is outside the local-only path. Decision: use pdf-inspector for routing and the PDF text layer for text-bearing pages. Retry empty extracted text even when routing says text. Use one Docling.rs background worker for textless pages, unload it when idle, recycle after at most 12 pages, cap it at 6 GiB RSS and 60 s per page. Use the previously measured PP-OCRv5 mobile ONNX path at lower resolution for the one render-cap failure and 11 blank pages. The measured PDF work plus estimated fallback is about 12–15 minutes for the 439 readable PDFs on the VM. The 16 unreadable PDFs and office extraction are outside that estimate. No transcripts or table/reading-order gold labels exist, so quality remains a proxy. The type/sender/tag classifier is unchanged pending the private R3 Paperless-ngx label export. Known gaps: no product recognition path was implemented; no valid CPU throughput/quality figures exist for models that failed or exceeded the one-page time limit. All staged corpus and model copies were removed from the VM. No per-document data is in this report. ## Gates Head SHA: `31a523176c2b91bbd1bac006732b36e1fb16a359`. `cargo fmt --check`: exit 0, no output. `cargo clippy --manifest-path bench/recog/Cargo.toml --all-targets -- -D warnings` (exit 0): ```text Checking serde_core v1.0.229 Checking zmij v1.0.23 Checking itoa v1.0.18 Checking memchr v2.8.3 Checking serde v1.0.229 Checking serde_json v1.0.151 Checking recog-bench v0.1.0 (/home/kayg/Developer/calternal-wt/recog-417/bench/recog) Finished `dev` profile [unoptimized + debuginfo] target(s) in 4m 00s ``` `cargo test --manifest-path bench/recog/Cargo.toml` (exit 0): ```text Compiling recog-bench v0.1.0 (/home/kayg/Developer/calternal-wt/recog-417/bench/recog) Finished `test` profile [unoptimized + debuginfo] target(s) in 3m 15s Running unittests src/main.rs (/mnt/hdd/targets/jobs/recog-417/debug/deps/recog_bench-c785a733d7ec9970) running 2 tests test tests::vectorizer_uses_training_terms_only ... ok test tests::softmax_probabilities_sum_to_one ... ok test result: ok. 2 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.01s ``` `bun run check` in `apps/web` (exit 0): ```text $ node scripts/check-type-tokens.mjs && svelte-kit sync && svelte-check --tsconfig ./tsconfig.json Text sizes use shared role tokens. Loading svelte-check in workspace: /home/kayg/Developer/calternal-wt/recog-417/apps/web Getting Svelte diagnostics... svelte-check found 0 errors and 0 warnings ``` `bun run test` in `apps/web` (exit 0; final test summary): ```text Test Files 126 passed (126) Tests 801 passed (801) Start at 15:25:00 Duration 102.61s (transform 55%, environment 17%, import 15%, tests 9%, setup 3%) ``` `python3 -m py_compile bench/recog/*.py` and the local privacy/extraction tests passed. Full local Python discovery could not import `numpy` for the earlier detector test because model dependencies are installed in the private VM environment. This does not affect the new routing/layout scripts. The latest host rule requires Rust gates per touched crate; the orchestrator runs full-workspace gates at merge time. No API was changed, so the adversarial API round did not apply. `cargo clean` removed 214 files, 105.1 MiB; web build output was removed. No push, deployment, or merge into `dev` was made.
Author
Owner

Orchestrator decision on round 2: accept pdf-inspector routing plus text layer first. Do not default to Docling.rs: calternal.cloud has 7.7 GB of RAM shared with the server, and the owner's rule is minimal memory. Default OCR for textless pages: PaddleOCR mobile ONNX (1.5 GB peak, 4 s per page, best text recovery 0.97), in one worker loaded only while the queue has work and unloaded when idle, with a named RSS ceiling of 2 GB. Docling.rs becomes an Admin option (quality over memory, needs 6 GB or more), off by default. AnyDoc may extract office-file text if it runs fully local. Round 3 (queued) measures the tag classifier on the owner's paperless labels (604/614 matched, 36 tags with 10 or more documents).

Orchestrator decision on round 2: accept pdf-inspector routing plus text layer first. **Do not** default to Docling.rs: calternal.cloud has 7.7 GB of RAM shared with the server, and the owner's rule is minimal memory. Default OCR for textless pages: PaddleOCR mobile ONNX (1.5 GB peak, 4 s per page, best text recovery 0.97), in one worker loaded only while the queue has work and unloaded when idle, with a named RSS ceiling of 2 GB. Docling.rs becomes an Admin option (quality over memory, needs 6 GB or more), off by default. AnyDoc may extract office-file text if it runs fully local. Round 3 (queued) measures the tag classifier on the owner's paperless labels (604/614 matched, 36 tags with 10 or more documents).
Author
Owner

Starting recog-417 round 3 on branch job/recog-417. Starting HEAD: 31a523176c; merge base with dev: 09e238d527; dev HEAD: 28ac39f917. I read the orchestrator decision: PaddleOCR mobile ONNX is default OCR; Docling.rs is an Admin option. I am measuring tag classification from the private Paperless export with aggregate-only reporting.

Starting recog-417 round 3 on branch job/recog-417. Starting HEAD: 31a523176c2b91bbd1bac006732b36e1fb16a359; merge base with dev: 09e238d5276b4f7a1eaf457eef07677cf9e3e9cc; dev HEAD: 28ac39f917f3dedd44dec2e863ee2854d64c2d5a. I read the orchestrator decision: PaddleOCR mobile ONNX is default OCR; Docling.rs is an Admin option. I am measuring tag classification from the private Paperless export with aggregate-only reporting.
Author
Owner

Round 3 data preparation findings (aggregate only): exact matching joined 604/614 label records. The matched subset has 573 tagged records and 35 tags with 10 or more documents; the full export has 582 tagged records and 36 such tags, so nine tagged records did not match. Local PDF text and Office XML extraction produced 544 usable rows; 60 matched rows had no usable text. The round-2 OCR text artifacts were not present in the local private work directory, so this run uses embedded PDF text and Office XML only. The local Photos CLIP text encoder produced 512-dimensional features for all 544 rows using two context windows per row; inference took 115.506 seconds.

Round 3 data preparation findings (aggregate only): exact matching joined 604/614 label records. The matched subset has 573 tagged records and 35 tags with 10 or more documents; the full export has 582 tagged records and 36 such tags, so nine tagged records did not match. Local PDF text and Office XML extraction produced 544 usable rows; 60 matched rows had no usable text. The round-2 OCR text artifacts were not present in the local private work directory, so this run uses embedded PDF text and Office XML only. The local Photos CLIP text encoder produced 512-dimensional features for all 544 rows using two context windows per row; inference took 115.506 seconds.
Author
Owner

The prepared, date-ordered classifier input contains 544 text rows and 96 distinct ranked tags. Six tags from the matched label set have no usable text row after the local text-layer and Office XML pass. They cannot be scored by a text model in this run; the report will show the 96 measured ranks and call out this coverage limit. No raw tag values or document identifiers are included.

The prepared, date-ordered classifier input contains 544 text rows and 96 distinct ranked tags. Six tags from the matched label set have no usable text row after the local text-layer and Office XML pass. They cannot be scored by a text model in this run; the report will show the 96 measured ranks and call out this coverage limit. No raw tag values or document identifiers are included.
Author
Owner

The Rust candidate's five-fold result is available. On the 96 tags with usable text, micro precision was 87.6%, micro recall was 33.7%, and micro ECE was 0.52%. One of 96 tags had a training-only 95% bar. Its held-out auto precision was 95.7% at 4.4% coverage of true tag instances. The date-ordered cold-start run found no tag with a qualifying bar through 200 corrections. These are aggregate corpus results; tag ranks and source labels are not included in this progress comment.

The Rust candidate's five-fold result is available. On the 96 tags with usable text, micro precision was 87.6%, micro recall was 33.7%, and micro ECE was 0.52%. One of 96 tags had a training-only 95% bar. Its held-out auto precision was 95.7% at 4.4% coverage of true tag instances. The date-ordered cold-start run found no tag with a qualifying bar through 200 corrections. These are aggregate corpus results; tag ranks and source labels are not included in this progress comment.
Author
Owner

Method review found that public rank ids must be based on tag frequency in the full 614-record export, so unmatched documents cannot reorder T01–T36. I corrected the join, verified aggregate input counts (604 matched, 544 usable text rows, 102 tags total, 36 tags with at least 10 records in the full export), and am rerunning the models with the corrected ranks. No original tag values or document identifiers are included here.

Method review found that public rank ids must be based on tag frequency in the full 614-record export, so unmatched documents cannot reorder T01–T36. I corrected the join, verified aggregate input counts (604 matched, 544 usable text rows, 102 tags total, 36 tags with at least 10 records in the full export), and am rerunning the models with the corrected ranks. No original tag values or document identifiers are included here.
Author
Owner

Document tag classification, round 3 (#417)

This run measures tag classification on the private Paperless export. It uses
aggregate results only. It does not include document text, titles, file names,
or original tag values.

Input and text

The export has 614 records. Exact basename matching found 604 records in the
corpus. Of those, 573 have tags. The matched set has 102 distinct tags. Thirty
five tags occur on at least 10 matched records. Nine tagged records did not
match. Six tags have no usable text row in this run.

The private input contains 544 text rows and 96 distinct ranked tags. It sorts
rows by the export's creation date. The tool reads embedded PDF text and text
from DOCX and XLSX files. It omits 60 matched records that have no usable text.
The round-2 OCR text output was not present in the private work directory. This
run does not score OCR. It does not include the textless rows in the classifier.

The rank order follows tag frequency in the full 614-record export. This keeps
public ids stable when a record has no corpus match. Public per-tag tables show
only T01–T36. The other measured tags contribute to the all-tag micro and macro
results.

Models and scoring

  • Paperless-style: word count unigrams and bigrams with one MLP that has a
    shared hidden layer and one binary output per tag.
  • Rust: word TF-IDF unigrams and bigrams with one logistic model per tag.
    The implementation uses sparse stochastic gradient descent with L2 decay.
  • TF-IDF + CLIP: word TF-IDF with a 512-value vector from the local Photos
    CLIP text model, followed by one logistic model per tag. The encoder uses at
    most two evenly spaced context windows. Feature extraction took 115.506 s
    across 544 rows.

Each fold fits its vocabulary, inverse document frequency, model, Platt
calibration, and confidence bar from training rows only. The Python models use
the same iterative multilabel folds. The Rust model stratifies each tag's
positive and negative rows separately. Each run uses five outer folds.

The confidence bar selects the widest calibration set with at least 20
predictions and at least 95% observed precision. Precision and recall use a
calibrated probability of 0.5. Coverage at 95% precision is the share of true
tag instances that the bar correctly auto-applies. ECE uses ten equal-width
probability bins. Micro ECE uses the combined tag-document probability bins.
Macro ECE is the mean of per-tag ECE values.

The cold-start run sorts usable text rows by creation date. It trains on the
first 10, 25, 50, 100, or 200 rows. The first 80% train the model. The last 20%
set its calibration bar. Later rows are held out. The run does not update the
model as it scores those later rows.

Results

The Paperless-style model has the highest micro recall (59.7%) and micro
precision of 74.7%. TF-IDF plus CLIP has 76.6% micro precision and 53.7% recall.
The Rust TF-IDF candidate has the highest micro precision (87.6%) and the
lowest recall (33.7%). Its macro precision and recall are low because many
infrequent tags receive no positive prediction. Of the 60 other measured tags,
18 receive no positive prediction from the Paperless-style model, 24 from Rust,
and 22 from TF-IDF plus CLIP. The CLIP feature does not improve micro recall
over the Paperless-style model in this corpus.

The Rust model is the only one with a held-out auto-apply precision at or above
95% (95.7%). Its coverage is 4.4% of true tag instances. The Paperless-style
model reaches 6.8% coverage, but its held-out auto-apply precision is 90.9%.
The CLIP model selects no held-out tags at its training-only bars. These results
show that a bar meeting 95% precision on calibration rows may not meet that
target on later documents. C95 below measures the share of true tags passed by
bars selected on training-only calibration rows; it does not promise 95%
precision on future documents.

No frequent tag has an eligible bar in any cold-start run through 200
corrections. The run cannot give a correction count at which most frequent
tags become eligible. The threshold is unknown and beyond the tested 200-row
limit if it is reached. Each N uses the newest fifth of those N rows for
calibration, so 10, 25, and 50 corrections do not provide the minimum 20
calibration candidates used by this benchmark.

Five-fold results

Precision and recall use calibrated probability ≥0.5. C95 is the share of true tags auto-applied by a bar selected on training-only calibration data; ECE uses ten probability bins.

Model Micro P Micro R Macro P Macro R Micro ECE Macro ECE Micro C95 Macro C95 Held-out auto P
Paperless MLP 74.7% 59.7% 48.0% 36.1% 0.004 0.009 6.8% 0.5% 90.9%
Rust TF-IDF logistic 87.6% 33.7% 14.6% 5.5% 0.005 0.007 4.4% 0.3% 95.7%
TF-IDF + CLIP 76.6% 53.7% 52.6% 31.6% 0.003 0.009 0.0% 0.0% —

Per-tag results

This table shows the 36 frequent tags only. Ranks follow frequency in the full export. Other measured tags contribute to the all-tag micro and macro results. The positive count is the number of usable text rows carrying the tag.

Rank Positive docs MLP P MLP R MLP C95 MLP ECE Rust P Rust R Rust C95 Rust ECE CLIP P CLIP R CLIP C95 CLIP ECE
T01 136 91.5% 86.8% 51.5% 0.039 82.2% 77.9% 33.1% 0.048 81.0% 75.0% 0.0% 0.059
T02 120 92.3% 90.0% 0.0% 0.031 89.1% 75.0% 0.0% 0.054 88.0% 79.2% 0.0% 0.070
T03 85 92.5% 72.9% 0.0% 0.014 98.1% 60.0% 0.0% 0.019 93.0% 62.4% 0.0% 0.015
T04 73 88.3% 72.6% 0.0% 0.031 88.7% 64.4% 0.0% 0.021 88.2% 61.6% 0.0% 0.024
T05 40 65.6% 52.5% 0.0% 0.022 80.0% 30.0% 0.0% 0.033 70.3% 65.0% 0.0% 0.028
T06 35 69.0% 82.9% 0.0% 0.025 100.0% 51.4% 0.0% 0.035 79.4% 77.1% 0.0% 0.023
T07 31 73.7% 45.2% 0.0% 0.021 80.0% 25.8% 0.0% 0.038 70.4% 61.3% 0.0% 0.029
T08 12 83.3% 41.7% 0.0% 0.009 0.0% 0.0% 0.0% 0.014 50.0% 33.3% 0.0% 0.010
T09 14 28.6% 14.3% 0.0% 0.016 33.3% 7.1% 0.0% 0.006 50.0% 14.3% 0.0% 0.015
T10 9 28.6% 22.2% 0.0% 0.017 0.0% 0.0% 0.0% 0.005 25.0% 22.2% 0.0% 0.014
T11 12 70.0% 58.3% 0.0% 0.010 100.0% 33.3% 0.0% 0.011 75.0% 50.0% 0.0% 0.016
T12 12 81.8% 75.0% 0.0% 0.007 0.0% 0.0% 0.0% 0.010 87.5% 58.3% 0.0% 0.005
T13 11 66.7% 36.4% 0.0% 0.009 100.0% 9.1% 0.0% 0.009 80.0% 36.4% 0.0% 0.007
T14 9 25.0% 11.1% 0.0% 0.016 0.0% 0.0% 0.0% 0.009 25.0% 22.2% 0.0% 0.014
T15 10 50.0% 10.0% 0.0% 0.004 0.0% 0.0% 0.0% 0.004 40.0% 20.0% 0.0% 0.013
T16 11 63.6% 63.6% 0.0% 0.011 100.0% 9.1% 0.0% 0.010 80.0% 72.7% 0.0% 0.008
T17 10 85.7% 60.0% 0.0% 0.007 100.0% 10.0% 0.0% 0.019 100.0% 60.0% 0.0% 0.004
T18 11 75.0% 54.5% 0.0% 0.006 0.0% 0.0% 0.0% 0.007 100.0% 63.6% 0.0% 0.005
T19 11 40.0% 18.2% 0.0% 0.011 0.0% 0.0% 0.0% 0.008 50.0% 18.2% 0.0% 0.011
T20 7 0.0% 0.0% 0.0% 0.012 0.0% 0.0% 0.0% 0.005 16.7% 14.3% 0.0% 0.011
T21 6 50.0% 33.3% 0.0% 0.008 0.0% 0.0% 0.0% 0.002 40.0% 33.3% 0.0% 0.006
T22 10 71.4% 50.0% 0.0% 0.010 0.0% 0.0% 0.0% 0.017 100.0% 50.0% 0.0% 0.009
T23 10 70.0% 70.0% 0.0% 0.008 0.0% 0.0% 0.0% 0.014 100.0% 40.0% 0.0% 0.008
T24 10 45.5% 50.0% 0.0% 0.017 0.0% 0.0% 0.0% 0.015 62.5% 50.0% 0.0% 0.011
T25 10 44.4% 40.0% 0.0% 0.020 0.0% 0.0% 0.0% 0.009 75.0% 60.0% 0.0% 0.008
T26 10 71.4% 50.0% 0.0% 0.008 0.0% 0.0% 0.0% 0.017 100.0% 40.0% 0.0% 0.012
T27 10 100.0% 50.0% 0.0% 0.012 0.0% 0.0% 0.0% 0.013 100.0% 40.0% 0.0% 0.008
T28 10 75.0% 60.0% 0.0% 0.009 0.0% 0.0% 0.0% 0.003 85.7% 60.0% 0.0% 0.008
T29 10 100.0% 40.0% 0.0% 0.005 0.0% 0.0% 0.0% 0.003 100.0% 30.0% 0.0% 0.003
T30 10 55.6% 50.0% 0.0% 0.013 0.0% 0.0% 0.0% 0.013 100.0% 30.0% 0.0% 0.010
T31 10 100.0% 40.0% 0.0% 0.011 0.0% 0.0% 0.0% 0.008 55.6% 50.0% 0.0% 0.009
T32 10 57.1% 80.0% 0.0% 0.015 0.0% 0.0% 0.0% 0.011 80.0% 40.0% 0.0% 0.009
T33 10 100.0% 30.0% 0.0% 0.012 0.0% 0.0% 0.0% 0.022 25.0% 30.0% 0.0% 0.018
T34 10 54.5% 60.0% 0.0% 0.015 0.0% 0.0% 0.0% 0.017 80.0% 40.0% 0.0% 0.010
T35 10 57.1% 40.0% 0.0% 0.012 0.0% 0.0% 0.0% 0.007 50.0% 10.0% 0.0% 0.010
T36 10 100.0% 60.0% 0.0% 0.007 0.0% 0.0% 0.0% 0.015 80.0% 40.0% 0.0% 0.009

Cold-start curve

The model trains on the first N dated rows, calibrates on the newest fifth of those rows, and scores all later rows. ‘Barred frequent tags’ counts T01–T36 with a usable calibration bar.

Model Corrections Barred frequent tags Micro P Micro R Micro C95 Held-out auto P Micro ECE
Paperless MLP 10 0/36 10.6% 3.4% 0.0% — 0.050
Paperless MLP 25 0/36 9.3% 5.0% 0.0% — 0.032
Paperless MLP 50 0/36 23.7% 4.0% 0.0% — 0.035
Paperless MLP 100 0/36 23.4% 22.4% 0.0% — 0.032
Paperless MLP 200 0/36 53.6% 32.9% 0.0% — 0.012
Rust TF-IDF logistic 10 0/36 36.6% 2.6% 0.0% — 0.022
Rust TF-IDF logistic 25 0/36 8.0% 0.6% 0.0% — 0.010
Rust TF-IDF logistic 50 0/36 0.0% 0.0% 0.0% — 0.015
Rust TF-IDF logistic 100 0/36 29.7% 17.7% 0.0% — 0.013
Rust TF-IDF logistic 200 0/36 96.2% 12.0% 0.0% — 0.007
TF-IDF + CLIP 10 0/36 33.8% 2.4% 0.0% — 0.094
TF-IDF + CLIP 25 0/36 66.7% 0.2% 0.0% — 0.042
TF-IDF + CLIP 50 0/36 28.9% 4.0% 0.0% — 0.035
TF-IDF + CLIP 100 0/36 50.9% 13.9% 0.0% — 0.024
TF-IDF + CLIP 200 0/36 74.0% 17.6% 0.0% — 0.011

First correction count with a bar for each frequent tag

A dash means no training-only calibration bar was available through 200 corrections.

Rank Paperless MLP Rust TF-IDF logistic TF-IDF + CLIP
T01 — — —
T02 — — —
T03 — — —
T04 — — —
T05 — — —
T06 — — —
T07 — — —
T08 — — —
T09 — — —
T10 — — —
T11 — — —
T12 — — —
T13 — — —
T14 — — —
T15 — — —
T16 — — —
T17 — — —
T18 — — —
T19 — — —
T20 — — —
T21 — — —
T22 — — —
T23 — — —
T24 — — —
T25 — — —
T26 — — —
T27 — — —
T28 — — —
T29 — — —
T30 — — —
T31 — — —
T32 — — —
T33 — — —
T34 — — —
T35 — — —
T36 — — —

At least half of frequent tags

Paperless MLP: fewer than 18 of the 36 frequent tags had a bar by 200 corrections.
Rust TF-IDF logistic: fewer than 18 of the 36 frequent tags had a bar by 200 corrections.
TF-IDF + CLIP: fewer than 18 of the 36 frequent tags had a bar by 200 corrections.

Decision for #341 Q6

Use a separate calibrated bar for each tag in the User's own correction
history. Auto-apply only when the bar has at least 95% precision on that
history. Require at least 20 calibration predictions before setting a bar.
Keep a rejection floor above the confidence of any rejected suggestion when
the model is updated.

The experiment does not establish that a small history will keep 95% precision
on future documents. The cold-start table shows how many frequent tags have a
training-only bar at each correction count.

Correspondents and document types

The export has no correspondent or document-type labels. It cannot support a
supervised precision or recall result for those fields. Document text may
contain sender strings or letterhead cues. These are possible zero-shot
signals, not measured evidence. Keep them as quiet suggestions until the User
confirms them. Do not auto-apply them from this corpus.

Scope and next step

The test uses PDF text and Office XML, not OCR output. The round-2 decision
sets PaddleOCR mobile ONNX as default OCR and Docling.rs as an Admin option.
Apply that OCR route in a later run to recover text from scans, then measure
whether tag coverage changes. No product recognition path is built here.

# Document tag classification, round 3 (#417) This run measures tag classification on the private Paperless export. It uses aggregate results only. It does not include document text, titles, file names, or original tag values. ## Input and text The export has 614 records. Exact basename matching found 604 records in the corpus. Of those, 573 have tags. The matched set has 102 distinct tags. Thirty five tags occur on at least 10 matched records. Nine tagged records did not match. Six tags have no usable text row in this run. The private input contains 544 text rows and 96 distinct ranked tags. It sorts rows by the export's creation date. The tool reads embedded PDF text and text from DOCX and XLSX files. It omits 60 matched records that have no usable text. The round-2 OCR text output was not present in the private work directory. This run does not score OCR. It does not include the textless rows in the classifier. The rank order follows tag frequency in the full 614-record export. This keeps public ids stable when a record has no corpus match. Public per-tag tables show only T01–T36. The other measured tags contribute to the all-tag micro and macro results. ## Models and scoring - **Paperless-style:** word count unigrams and bigrams with one MLP that has a shared hidden layer and one binary output per tag. - **Rust:** word TF-IDF unigrams and bigrams with one logistic model per tag. The implementation uses sparse stochastic gradient descent with L2 decay. - **TF-IDF + CLIP:** word TF-IDF with a 512-value vector from the local Photos CLIP text model, followed by one logistic model per tag. The encoder uses at most two evenly spaced context windows. Feature extraction took 115.506 s across 544 rows. Each fold fits its vocabulary, inverse document frequency, model, Platt calibration, and confidence bar from training rows only. The Python models use the same iterative multilabel folds. The Rust model stratifies each tag's positive and negative rows separately. Each run uses five outer folds. The confidence bar selects the widest calibration set with at least 20 predictions and at least 95% observed precision. Precision and recall use a calibrated probability of 0.5. Coverage at 95% precision is the share of true tag instances that the bar correctly auto-applies. ECE uses ten equal-width probability bins. Micro ECE uses the combined tag-document probability bins. Macro ECE is the mean of per-tag ECE values. The cold-start run sorts usable text rows by creation date. It trains on the first 10, 25, 50, 100, or 200 rows. The first 80% train the model. The last 20% set its calibration bar. Later rows are held out. The run does not update the model as it scores those later rows. ## Results The Paperless-style model has the highest micro recall (59.7%) and micro precision of 74.7%. TF-IDF plus CLIP has 76.6% micro precision and 53.7% recall. The Rust TF-IDF candidate has the highest micro precision (87.6%) and the lowest recall (33.7%). Its macro precision and recall are low because many infrequent tags receive no positive prediction. Of the 60 other measured tags, 18 receive no positive prediction from the Paperless-style model, 24 from Rust, and 22 from TF-IDF plus CLIP. The CLIP feature does not improve micro recall over the Paperless-style model in this corpus. The Rust model is the only one with a held-out auto-apply precision at or above 95% (95.7%). Its coverage is 4.4% of true tag instances. The Paperless-style model reaches 6.8% coverage, but its held-out auto-apply precision is 90.9%. The CLIP model selects no held-out tags at its training-only bars. These results show that a bar meeting 95% precision on calibration rows may not meet that target on later documents. C95 below measures the share of true tags passed by bars selected on training-only calibration rows; it does not promise 95% precision on future documents. No frequent tag has an eligible bar in any cold-start run through 200 corrections. The run cannot give a correction count at which most frequent tags become eligible. The threshold is unknown and beyond the tested 200-row limit if it is reached. Each N uses the newest fifth of those N rows for calibration, so 10, 25, and 50 corrections do not provide the minimum 20 calibration candidates used by this benchmark. ### Five-fold results Precision and recall use calibrated probability ≥0.5. C95 is the share of true tags auto-applied by a bar selected on training-only calibration data; ECE uses ten probability bins. | Model | Micro P | Micro R | Macro P | Macro R | Micro ECE | Macro ECE | Micro C95 | Macro C95 | Held-out auto P | |---|---:|---:|---:|---:|---:|---:|---:|---:|---:| | Paperless MLP | 74.7% | 59.7% | 48.0% | 36.1% | 0.004 | 0.009 | 6.8% | 0.5% | 90.9% | | Rust TF-IDF logistic | 87.6% | 33.7% | 14.6% | 5.5% | 0.005 | 0.007 | 4.4% | 0.3% | 95.7% | | TF-IDF + CLIP | 76.6% | 53.7% | 52.6% | 31.6% | 0.003 | 0.009 | 0.0% | 0.0% | — | ### Per-tag results This table shows the 36 frequent tags only. Ranks follow frequency in the full export. Other measured tags contribute to the all-tag micro and macro results. The positive count is the number of usable text rows carrying the tag. | Rank | Positive docs | MLP P | MLP R | MLP C95 | MLP ECE | Rust P | Rust R | Rust C95 | Rust ECE | CLIP P | CLIP R | CLIP C95 | CLIP ECE | |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:| | T01 | 136 | 91.5% | 86.8% | 51.5% | 0.039 | 82.2% | 77.9% | 33.1% | 0.048 | 81.0% | 75.0% | 0.0% | 0.059 | | T02 | 120 | 92.3% | 90.0% | 0.0% | 0.031 | 89.1% | 75.0% | 0.0% | 0.054 | 88.0% | 79.2% | 0.0% | 0.070 | | T03 | 85 | 92.5% | 72.9% | 0.0% | 0.014 | 98.1% | 60.0% | 0.0% | 0.019 | 93.0% | 62.4% | 0.0% | 0.015 | | T04 | 73 | 88.3% | 72.6% | 0.0% | 0.031 | 88.7% | 64.4% | 0.0% | 0.021 | 88.2% | 61.6% | 0.0% | 0.024 | | T05 | 40 | 65.6% | 52.5% | 0.0% | 0.022 | 80.0% | 30.0% | 0.0% | 0.033 | 70.3% | 65.0% | 0.0% | 0.028 | | T06 | 35 | 69.0% | 82.9% | 0.0% | 0.025 | 100.0% | 51.4% | 0.0% | 0.035 | 79.4% | 77.1% | 0.0% | 0.023 | | T07 | 31 | 73.7% | 45.2% | 0.0% | 0.021 | 80.0% | 25.8% | 0.0% | 0.038 | 70.4% | 61.3% | 0.0% | 0.029 | | T08 | 12 | 83.3% | 41.7% | 0.0% | 0.009 | 0.0% | 0.0% | 0.0% | 0.014 | 50.0% | 33.3% | 0.0% | 0.010 | | T09 | 14 | 28.6% | 14.3% | 0.0% | 0.016 | 33.3% | 7.1% | 0.0% | 0.006 | 50.0% | 14.3% | 0.0% | 0.015 | | T10 | 9 | 28.6% | 22.2% | 0.0% | 0.017 | 0.0% | 0.0% | 0.0% | 0.005 | 25.0% | 22.2% | 0.0% | 0.014 | | T11 | 12 | 70.0% | 58.3% | 0.0% | 0.010 | 100.0% | 33.3% | 0.0% | 0.011 | 75.0% | 50.0% | 0.0% | 0.016 | | T12 | 12 | 81.8% | 75.0% | 0.0% | 0.007 | 0.0% | 0.0% | 0.0% | 0.010 | 87.5% | 58.3% | 0.0% | 0.005 | | T13 | 11 | 66.7% | 36.4% | 0.0% | 0.009 | 100.0% | 9.1% | 0.0% | 0.009 | 80.0% | 36.4% | 0.0% | 0.007 | | T14 | 9 | 25.0% | 11.1% | 0.0% | 0.016 | 0.0% | 0.0% | 0.0% | 0.009 | 25.0% | 22.2% | 0.0% | 0.014 | | T15 | 10 | 50.0% | 10.0% | 0.0% | 0.004 | 0.0% | 0.0% | 0.0% | 0.004 | 40.0% | 20.0% | 0.0% | 0.013 | | T16 | 11 | 63.6% | 63.6% | 0.0% | 0.011 | 100.0% | 9.1% | 0.0% | 0.010 | 80.0% | 72.7% | 0.0% | 0.008 | | T17 | 10 | 85.7% | 60.0% | 0.0% | 0.007 | 100.0% | 10.0% | 0.0% | 0.019 | 100.0% | 60.0% | 0.0% | 0.004 | | T18 | 11 | 75.0% | 54.5% | 0.0% | 0.006 | 0.0% | 0.0% | 0.0% | 0.007 | 100.0% | 63.6% | 0.0% | 0.005 | | T19 | 11 | 40.0% | 18.2% | 0.0% | 0.011 | 0.0% | 0.0% | 0.0% | 0.008 | 50.0% | 18.2% | 0.0% | 0.011 | | T20 | 7 | 0.0% | 0.0% | 0.0% | 0.012 | 0.0% | 0.0% | 0.0% | 0.005 | 16.7% | 14.3% | 0.0% | 0.011 | | T21 | 6 | 50.0% | 33.3% | 0.0% | 0.008 | 0.0% | 0.0% | 0.0% | 0.002 | 40.0% | 33.3% | 0.0% | 0.006 | | T22 | 10 | 71.4% | 50.0% | 0.0% | 0.010 | 0.0% | 0.0% | 0.0% | 0.017 | 100.0% | 50.0% | 0.0% | 0.009 | | T23 | 10 | 70.0% | 70.0% | 0.0% | 0.008 | 0.0% | 0.0% | 0.0% | 0.014 | 100.0% | 40.0% | 0.0% | 0.008 | | T24 | 10 | 45.5% | 50.0% | 0.0% | 0.017 | 0.0% | 0.0% | 0.0% | 0.015 | 62.5% | 50.0% | 0.0% | 0.011 | | T25 | 10 | 44.4% | 40.0% | 0.0% | 0.020 | 0.0% | 0.0% | 0.0% | 0.009 | 75.0% | 60.0% | 0.0% | 0.008 | | T26 | 10 | 71.4% | 50.0% | 0.0% | 0.008 | 0.0% | 0.0% | 0.0% | 0.017 | 100.0% | 40.0% | 0.0% | 0.012 | | T27 | 10 | 100.0% | 50.0% | 0.0% | 0.012 | 0.0% | 0.0% | 0.0% | 0.013 | 100.0% | 40.0% | 0.0% | 0.008 | | T28 | 10 | 75.0% | 60.0% | 0.0% | 0.009 | 0.0% | 0.0% | 0.0% | 0.003 | 85.7% | 60.0% | 0.0% | 0.008 | | T29 | 10 | 100.0% | 40.0% | 0.0% | 0.005 | 0.0% | 0.0% | 0.0% | 0.003 | 100.0% | 30.0% | 0.0% | 0.003 | | T30 | 10 | 55.6% | 50.0% | 0.0% | 0.013 | 0.0% | 0.0% | 0.0% | 0.013 | 100.0% | 30.0% | 0.0% | 0.010 | | T31 | 10 | 100.0% | 40.0% | 0.0% | 0.011 | 0.0% | 0.0% | 0.0% | 0.008 | 55.6% | 50.0% | 0.0% | 0.009 | | T32 | 10 | 57.1% | 80.0% | 0.0% | 0.015 | 0.0% | 0.0% | 0.0% | 0.011 | 80.0% | 40.0% | 0.0% | 0.009 | | T33 | 10 | 100.0% | 30.0% | 0.0% | 0.012 | 0.0% | 0.0% | 0.0% | 0.022 | 25.0% | 30.0% | 0.0% | 0.018 | | T34 | 10 | 54.5% | 60.0% | 0.0% | 0.015 | 0.0% | 0.0% | 0.0% | 0.017 | 80.0% | 40.0% | 0.0% | 0.010 | | T35 | 10 | 57.1% | 40.0% | 0.0% | 0.012 | 0.0% | 0.0% | 0.0% | 0.007 | 50.0% | 10.0% | 0.0% | 0.010 | | T36 | 10 | 100.0% | 60.0% | 0.0% | 0.007 | 0.0% | 0.0% | 0.0% | 0.015 | 80.0% | 40.0% | 0.0% | 0.009 | ### Cold-start curve The model trains on the first N dated rows, calibrates on the newest fifth of those rows, and scores all later rows. ‘Barred frequent tags’ counts T01–T36 with a usable calibration bar. | Model | Corrections | Barred frequent tags | Micro P | Micro R | Micro C95 | Held-out auto P | Micro ECE | |---|---:|---:|---:|---:|---:|---:|---:| | Paperless MLP | 10 | 0/36 | 10.6% | 3.4% | 0.0% | — | 0.050 | | Paperless MLP | 25 | 0/36 | 9.3% | 5.0% | 0.0% | — | 0.032 | | Paperless MLP | 50 | 0/36 | 23.7% | 4.0% | 0.0% | — | 0.035 | | Paperless MLP | 100 | 0/36 | 23.4% | 22.4% | 0.0% | — | 0.032 | | Paperless MLP | 200 | 0/36 | 53.6% | 32.9% | 0.0% | — | 0.012 | | Rust TF-IDF logistic | 10 | 0/36 | 36.6% | 2.6% | 0.0% | — | 0.022 | | Rust TF-IDF logistic | 25 | 0/36 | 8.0% | 0.6% | 0.0% | — | 0.010 | | Rust TF-IDF logistic | 50 | 0/36 | 0.0% | 0.0% | 0.0% | — | 0.015 | | Rust TF-IDF logistic | 100 | 0/36 | 29.7% | 17.7% | 0.0% | — | 0.013 | | Rust TF-IDF logistic | 200 | 0/36 | 96.2% | 12.0% | 0.0% | — | 0.007 | | TF-IDF + CLIP | 10 | 0/36 | 33.8% | 2.4% | 0.0% | — | 0.094 | | TF-IDF + CLIP | 25 | 0/36 | 66.7% | 0.2% | 0.0% | — | 0.042 | | TF-IDF + CLIP | 50 | 0/36 | 28.9% | 4.0% | 0.0% | — | 0.035 | | TF-IDF + CLIP | 100 | 0/36 | 50.9% | 13.9% | 0.0% | — | 0.024 | | TF-IDF + CLIP | 200 | 0/36 | 74.0% | 17.6% | 0.0% | — | 0.011 | ### First correction count with a bar for each frequent tag A dash means no training-only calibration bar was available through 200 corrections. | Rank | Paperless MLP | Rust TF-IDF logistic | TF-IDF + CLIP | |---|---:|---:|---:| | T01 | — | — | — | | T02 | — | — | — | | T03 | — | — | — | | T04 | — | — | — | | T05 | — | — | — | | T06 | — | — | — | | T07 | — | — | — | | T08 | — | — | — | | T09 | — | — | — | | T10 | — | — | — | | T11 | — | — | — | | T12 | — | — | — | | T13 | — | — | — | | T14 | — | — | — | | T15 | — | — | — | | T16 | — | — | — | | T17 | — | — | — | | T18 | — | — | — | | T19 | — | — | — | | T20 | — | — | — | | T21 | — | — | — | | T22 | — | — | — | | T23 | — | — | — | | T24 | — | — | — | | T25 | — | — | — | | T26 | — | — | — | | T27 | — | — | — | | T28 | — | — | — | | T29 | — | — | — | | T30 | — | — | — | | T31 | — | — | — | | T32 | — | — | — | | T33 | — | — | — | | T34 | — | — | — | | T35 | — | — | — | | T36 | — | — | — | ### At least half of frequent tags Paperless MLP: fewer than 18 of the 36 frequent tags had a bar by 200 corrections. Rust TF-IDF logistic: fewer than 18 of the 36 frequent tags had a bar by 200 corrections. TF-IDF + CLIP: fewer than 18 of the 36 frequent tags had a bar by 200 corrections. ## Decision for #341 Q6 Use a separate calibrated bar for each tag in the User's own correction history. Auto-apply only when the bar has at least 95% precision on that history. Require at least 20 calibration predictions before setting a bar. Keep a rejection floor above the confidence of any rejected suggestion when the model is updated. The experiment does not establish that a small history will keep 95% precision on future documents. The cold-start table shows how many frequent tags have a training-only bar at each correction count. ## Correspondents and document types The export has no correspondent or document-type labels. It cannot support a supervised precision or recall result for those fields. Document text may contain sender strings or letterhead cues. These are possible zero-shot signals, not measured evidence. Keep them as quiet suggestions until the User confirms them. Do not auto-apply them from this corpus. ## Scope and next step The test uses PDF text and Office XML, not OCR output. The round-2 decision sets PaddleOCR mobile ONNX as default OCR and Docling.rs as an Admin option. Apply that OCR route in a later run to recover text from scans, then measure whether tag coverage changes. No product recognition path is built here.
Author
Owner

The required local API adversarial pass was attempted with ADVERSARIAL_API_ONLY=1 and the web build skipped. The runner completed dependency setup but did not start its server or send probes: its cargo build process remained in futex_wait_queue for 14:48 with no server binary and no compiler output, while the shared host had many concurrent Cargo jobs. I sent SIGINT to that build and the runner exited with status 1. This is a setup block, not an API finding; no adversarial API assertions ran. I did not retry the round.

The required local API adversarial pass was attempted with `ADVERSARIAL_API_ONLY=1` and the web build skipped. The runner completed dependency setup but did not start its server or send probes: its `cargo build` process remained in `futex_wait_queue` for 14:48 with no server binary and no compiler output, while the shared host had many concurrent Cargo jobs. I sent SIGINT to that build and the runner exited with status 1. This is a setup block, not an API finding; no adversarial API assertions ran. I did not retry the round.
Author
Owner

Completed

Built the rank-only multi-label tag benchmark for #417. It compares a Paperless-style shared MLP, Rust TF-IDF one-vs-rest logistic regression, and TF-IDF plus the local CLIP text feature. The five-fold tables and cold-start results are in docs/research/document-recognition-417-round3.md; this comment thread also has the complete tables. Public output uses only T-rank ids.

The input contained 614 export records, 604 corpus matches, and 544 usable text rows. There are 102 tags in the full export; 36 occur on at least 10 records. Six tags had no usable text row. The Rust candidate had 87.6% micro precision, 33.7% micro recall, and 4.4% true-tag coverage at training-calibrated 95% bars; held-out auto precision was 95.7%. No frequent tag had an eligible bar in the cold-start runs through 200 corrections, so the count for most tags remains unknown beyond the tested range.

Decisions

  • Rank tags across the full export before matching corpus text, so unmatched records cannot reorder T-ranks.
  • Use iterative multi-label stratification for the Python models and per-tag binary stratification for Rust.
  • Fit per-tag Platt calibration from training-only predictions. A bar needs at least 20 calibration candidates and at least 95% observed precision; choose the widest eligible bar.
  • For #341 Q6, require a separate 95% bar from the User's own history and keep each rejection as a floor above its rejected confidence when recalibrating.
  • Cold-start uses the first N dated usable-text rows, fits on the first 80%, calibrates on the last 20%, and scores later rows. This run does not update the model after each score.
  • Use at most two evenly spaced CLIP context windows. This local feature pass took 115.506 seconds across 544 rows.

Known gaps

  • The local R2 OCR artifacts were absent. This run uses embedded PDF text and DOCX/XLSX XML and omits 60 matched rows without usable text. It does not measure PaddleOCR or Docling.rs.
  • Correspondent and document-type labels are null in the export. Letterhead or sender text may support quiet suggestions, but this data provides no supervised evidence or accuracy numbers for those fields.
  • The cold-start test found no frequent-tag calibration bar through 200 corrections. It cannot estimate when most frequent tags become eligible.
  • The local API adversarial run was attempted with ADVERSARIAL_API_ONLY=1 and the web build skipped. Setup installed dependencies but its server cargo build stayed in futex_wait_queue for 14:48 with no binary and no compiler output while many shared Cargo jobs were active. I stopped it with SIGINT; the runner exited 1 before starting the server or sending probes. No API assertions ran. This is a setup block, not an API finding. I did not retry.

Files

  • bench/recog/: private label preparation, CLIP features, shared Python models, Rust candidate, rank-only report generator, tests, and run documentation.
  • docs/research/document-recognition-417.md: link to the round-three report.
  • docs/research/document-recognition-417-round3.md: method, aggregate and per-tag tables, cold-start curve, and recommendation.

Merge and gates

Merged local dev once before final gates. Head SHA: a0821e5d392c3c0038832d63c1364da1b14039fc.

cargo fmt --manifest-path bench/recog/Cargo.toml --check (exit 0; output was empty).

cargo clippy --manifest-path bench/recog/Cargo.toml --all-targets -- -D warnings (exit 0), output:

    Finished `dev` profile [unoptimized + debuginfo] target(s) in 1.21s

cargo test --manifest-path bench/recog/Cargo.toml (exit 0), output:

    Finished `test` profile [unoptimized + debuginfo] target(s) in 0.49s
     Running unittests src/main.rs (/mnt/hdd/targets/jobs/recog-417/debug/deps/recog_bench-c785a733d7ec9970)

running 3 tests
test tests::stratified_folds_balance_each_class ... ok
test tests::coverage_is_a_share_of_true_tags ... ok
test tests::confidence_bar_requires_twenty_predictions ... ok

test result: ok. 3 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s

Python benchmark tests (exit 0), output:

.................
----------------------------------------------------------------------
Ran 17 tests in 0.050s

OK

py_compile succeeded with no output. cargo clean --manifest-path bench/recog/Cargo.toml output:

     Removed 1230 files, 584.7MiB total
## Completed Built the rank-only multi-label tag benchmark for #417. It compares a Paperless-style shared MLP, Rust TF-IDF one-vs-rest logistic regression, and TF-IDF plus the local CLIP text feature. The five-fold tables and cold-start results are in `docs/research/document-recognition-417-round3.md`; this comment thread also has the complete tables. Public output uses only T-rank ids. The input contained 614 export records, 604 corpus matches, and 544 usable text rows. There are 102 tags in the full export; 36 occur on at least 10 records. Six tags had no usable text row. The Rust candidate had 87.6% micro precision, 33.7% micro recall, and 4.4% true-tag coverage at training-calibrated 95% bars; held-out auto precision was 95.7%. No frequent tag had an eligible bar in the cold-start runs through 200 corrections, so the count for most tags remains unknown beyond the tested range. ## Decisions - Rank tags across the full export before matching corpus text, so unmatched records cannot reorder T-ranks. - Use iterative multi-label stratification for the Python models and per-tag binary stratification for Rust. - Fit per-tag Platt calibration from training-only predictions. A bar needs at least 20 calibration candidates and at least 95% observed precision; choose the widest eligible bar. - For #341 Q6, require a separate 95% bar from the User's own history and keep each rejection as a floor above its rejected confidence when recalibrating. - Cold-start uses the first N dated usable-text rows, fits on the first 80%, calibrates on the last 20%, and scores later rows. This run does not update the model after each score. - Use at most two evenly spaced CLIP context windows. This local feature pass took 115.506 seconds across 544 rows. ## Known gaps - The local R2 OCR artifacts were absent. This run uses embedded PDF text and DOCX/XLSX XML and omits 60 matched rows without usable text. It does not measure PaddleOCR or Docling.rs. - Correspondent and document-type labels are null in the export. Letterhead or sender text may support quiet suggestions, but this data provides no supervised evidence or accuracy numbers for those fields. - The cold-start test found no frequent-tag calibration bar through 200 corrections. It cannot estimate when most frequent tags become eligible. - The local API adversarial run was attempted with `ADVERSARIAL_API_ONLY=1` and the web build skipped. Setup installed dependencies but its server `cargo build` stayed in `futex_wait_queue` for 14:48 with no binary and no compiler output while many shared Cargo jobs were active. I stopped it with SIGINT; the runner exited 1 before starting the server or sending probes. No API assertions ran. This is a setup block, not an API finding. I did not retry. ## Files - `bench/recog/`: private label preparation, CLIP features, shared Python models, Rust candidate, rank-only report generator, tests, and run documentation. - `docs/research/document-recognition-417.md`: link to the round-three report. - `docs/research/document-recognition-417-round3.md`: method, aggregate and per-tag tables, cold-start curve, and recommendation. ## Merge and gates Merged local `dev` once before final gates. Head SHA: `a0821e5d392c3c0038832d63c1364da1b14039fc`. `cargo fmt --manifest-path bench/recog/Cargo.toml --check` (exit 0; output was empty). `cargo clippy --manifest-path bench/recog/Cargo.toml --all-targets -- -D warnings` (exit 0), output: ```text Finished `dev` profile [unoptimized + debuginfo] target(s) in 1.21s ``` `cargo test --manifest-path bench/recog/Cargo.toml` (exit 0), output: ```text Finished `test` profile [unoptimized + debuginfo] target(s) in 0.49s Running unittests src/main.rs (/mnt/hdd/targets/jobs/recog-417/debug/deps/recog_bench-c785a733d7ec9970) running 3 tests test tests::stratified_folds_balance_each_class ... ok test tests::coverage_is_a_share_of_true_tags ... ok test tests::confidence_bar_requires_twenty_predictions ... ok test result: ok. 3 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s ``` Python benchmark tests (exit 0), output: ```text ................. ---------------------------------------------------------------------- Ran 17 tests in 0.050s OK ``` `py_compile` succeeded with no output. `cargo clean --manifest-path bench/recog/Cargo.toml` output: ```text Removed 1230 files, 584.7MiB total ```
Author
Owner

Round 4 resumed on job/recog-417 from base 43c1377c8fd3b1b495841446989a64156f8ae1e4; current head a0821e5d392c3c0038832d63c1364da1b14039fc. I am extending the private, rank-only benchmark for the owner’s 95%/20-confirmation 90% rule, semantic nearest neighbours, extra signals, and OCR recovery. No corpus content or names will enter the public repository or issue.

Round 4 resumed on `job/recog-417` from base `43c1377c8fd3b1b495841446989a64156f8ae1e4`; current head `a0821e5d392c3c0038832d63c1364da1b14039fc`. I am extending the private, rank-only benchmark for the owner’s 95%/20-confirmation 90% rule, semantic nearest neighbours, extra signals, and OCR recovery. No corpus content or names will enter the public repository or issue.
Author
Owner

Round-4 OCR finding: Of the 60 matched rows without usable round-3 text, the source set is 56 PDFs, three legacy spreadsheets and one image. PDFium reports password protection for 16 PDFs, so their pages cannot be rendered or OCRed without keys. The first OCR attempt also found that this host has no pdftoppm; the benchmark now renders with the already used PDFium library. OCR is running on the readable PDFs and image. The locked PDFs and legacy spreadsheets will remain an explicit measured gap; no private names or content are reported.

Round-4 OCR finding: Of the 60 matched rows without usable round-3 text, the source set is 56 PDFs, three legacy spreadsheets and one image. PDFium reports password protection for 16 PDFs, so their pages cannot be rendered or OCRed without keys. The first OCR attempt also found that this host has no `pdftoppm`; the benchmark now renders with the already used PDFium library. OCR is running on the readable PDFs and image. The locked PDFs and legacy spreadsheets will remain an explicit measured gap; no private names or content are reported.
Author
Owner

Locked four-vCPU VM measurements (load averages recorded inside flock /root/perf.lock): Search's pinned MiniLM ONNX encoder on 50 private documents took 1.621 s total, 32.4 ms/document, with 121,552 KB peak process RSS; starting load 0.04, 0.20, 0.50. A separate run of TF-IDF feature extraction, tag heads and cosine voting on 109 held-out private documents measured p50 36.035 ms, p95 42.106 ms, 187,688 KB peak RSS; starting load 0.02, 0.08, 0.36. These are separate processes, so their RSS values cannot simply be added as a measured combined peak. OCR cost is separate.

Locked four-vCPU VM measurements (load averages recorded inside `flock /root/perf.lock`): Search's pinned MiniLM ONNX encoder on 50 private documents took 1.621 s total, 32.4 ms/document, with 121,552 KB peak process RSS; starting load 0.04, 0.20, 0.50. A separate run of TF-IDF feature extraction, tag heads and cosine voting on 109 held-out private documents measured p50 36.035 ms, p95 42.106 ms, 187,688 KB peak RSS; starting load 0.02, 0.08, 0.36. These are separate processes, so their RSS values cannot simply be added as a measured combined peak. OCR cost is separate.
Author
Owner

OCR recovery completed. The matched corpus has 604 records. Round 3 had 544 usable text rows and omitted 60. PaddleOCR mobile ONNX recovered nonempty text for 41 of those 60, giving 585 usable rows, 555 with at least one Tag and 100 measured Tag ranks. Sixteen PDFs remain inaccessible because PDFium reports password protection; three legacy spreadsheets are unsupported. First-page source lines were recovered for 537 usable rows. The other 48 use a bounded text prefix as a header proxy. All text, names and vectors remain in the private work directory.

OCR recovery completed. The matched corpus has 604 records. Round 3 had 544 usable text rows and omitted 60. PaddleOCR mobile ONNX recovered nonempty text for 41 of those 60, giving 585 usable rows, 555 with at least one Tag and 100 measured Tag ranks. Sixteen PDFs remain inaccessible because PDFium reports password protection; three legacy spreadsheets are unsupported. First-page source lines were recovered for 537 usable rows. The other 48 use a bounded text prefix as a header proxy. All text, names and vectors remain in the private work directory.
Author
Owner

Final round-4 aggregate finding (585 usable rows, 100 Tag ranks): MiniLM k=3/5/10 gave 0% held-out C95; filename lexical voting gave 1.9% C95 at 95.5% held-out auto precision, and its 20-confirmation 90% slice gave 3.8% coverage at 91.3% precision. The calibrated stack gave 8.9% C95 and 10.7% mixed-policy coverage, but only 81.8% held-out auto precision, below both owner bars. In the chronological simulation, no model made all top 5, 10 or 20 Tags eligible through 500 corrections, and automatic precision was unstable where any bar existed. This supports quiet combined suggestions and cautious filename-only automation after sufficient User-specific evidence; the stack must not auto-apply from these calibration bars.

Updated locked VM inference for the final lexical-vote pipeline: on 117 held-out rows and 100 Tag heads, feature extraction plus voting measured p50 11.239 ms, p95 14.307 ms, peak RSS 177,008 KB; starting load 0.64, 0.53, 0.32. Search encoding was measured separately at 32.4 ms/document and 121,552 KB peak RSS on 50 rows. The prior 36.035 ms classifier result used the slower all-logistic candidate before lexical votes replaced its four signal heads; it is superseded.

Final round-4 aggregate finding (585 usable rows, 100 Tag ranks): MiniLM k=3/5/10 gave 0% held-out C95; filename lexical voting gave 1.9% C95 at 95.5% held-out auto precision, and its 20-confirmation 90% slice gave 3.8% coverage at 91.3% precision. The calibrated stack gave 8.9% C95 and 10.7% mixed-policy coverage, but only 81.8% held-out auto precision, below both owner bars. In the chronological simulation, no model made all top 5, 10 or 20 Tags eligible through 500 corrections, and automatic precision was unstable where any bar existed. This supports quiet combined suggestions and cautious filename-only automation after sufficient User-specific evidence; the stack must not auto-apply from these calibration bars. Updated locked VM inference for the final lexical-vote pipeline: on 117 held-out rows and 100 Tag heads, feature extraction plus voting measured p50 11.239 ms, p95 14.307 ms, peak RSS 177,008 KB; starting load 0.64, 0.53, 0.32. Search encoding was measured separately at 32.4 ms/document and 121,552 KB peak RSS on 50 rows. The prior 36.035 ms classifier result used the slower all-logistic candidate before lexical votes replaced its four signal heads; it is superseded.
Author
Owner

The final combined resident-process VM measurement supersedes the sum of separate timings. Under flock /root/perf.lock with starting load 0.20, 0.42, 0.29, the pinned Search encoder plus TF-IDF heads, lexical and semantic votes, and stack inputs processed 117 held-out rows (100 Tag heads) at p50 48.686 ms and p95 59.395 ms per new record. Peak RSS was 252,156 KB; the process reported 258,576,384 bytes after inference. OCR remains a separate scan-only stage; round 2 measured PaddleOCR mobile ONNX at 4.02 wall s/page and 1.52 GB peak RSS on the same VM.

The final combined resident-process VM measurement supersedes the sum of separate timings. Under `flock /root/perf.lock` with starting load 0.20, 0.42, 0.29, the pinned Search encoder plus TF-IDF heads, lexical and semantic votes, and stack inputs processed 117 held-out rows (100 Tag heads) at p50 48.686 ms and p95 59.395 ms per new record. Peak RSS was 252,156 KB; the process reported 258,576,384 bytes after inference. OCR remains a separate scan-only stage; round 2 measured PaddleOCR mobile ONNX at 4.02 wall s/page and 1.52 GB peak RSS on the same VM.
Author
Owner

Production-parity correction: calternal-embed disables the pinned tokenizer's fixed 128-token padding, uses two ONNX intra-op threads and disables thread spinning. I changed the prototype to those settings, regenerated all 585 Search vectors, and reran all five folds and cold-start points. The recommendation and coverage are unchanged to one decimal place (the small MiniLM precision/recall values move by at most 0.2 percentage points). Under flock /root/perf.lock, the corrected combined resident-process run on 117 records measured p50 75.214 ms, p95 85.331 ms, peak RSS 252,276 KB, with starting load 0.72, 0.50, 0.35. These numbers supersede the earlier 48.686/59.395 ms run, which used the tokenizer's fixed padding and four ONNX threads. The public round-4 report now contains the corrected values.

Production-parity correction: `calternal-embed` disables the pinned tokenizer's fixed 128-token padding, uses two ONNX intra-op threads and disables thread spinning. I changed the prototype to those settings, regenerated all 585 Search vectors, and reran all five folds and cold-start points. The recommendation and coverage are unchanged to one decimal place (the small MiniLM precision/recall values move by at most 0.2 percentage points). Under `flock /root/perf.lock`, the corrected combined resident-process run on 117 records measured p50 75.214 ms, p95 85.331 ms, peak RSS 252,276 KB, with starting load 0.72, 0.50, 0.35. These numbers supersede the earlier 48.686/59.395 ms run, which used the tokenizer's fixed padding and four ONNX threads. The public round-4 report now contains the corrected values.
Author
Owner

Final workspace Clippy exposed one pre-existing Rust 1.98 lint in calternal-fs: its one-byte Search readiness marker compared a byte array with [b'1']. Clippy with -D warnings requests the equivalent *b"1". The marker writer already uses b"1"; I made that one-line, behavior-preserving change outside the benchmark-owned files so the required gate can continue. The initial full gate exited 101 at that lint. I am verifying calternal-fs before the follow-up workspace run.

Final workspace Clippy exposed one pre-existing Rust 1.98 lint in `calternal-fs`: its one-byte Search readiness marker compared a byte array with `[b'1']`. Clippy with `-D warnings` requests the equivalent `*b"1"`. The marker writer already uses `b"1"`; I made that one-line, behavior-preserving change outside the benchmark-owned files so the required gate can continue. The initial full gate exited 101 at that lint. I am verifying `calternal-fs` before the follow-up workspace run.
Author
Owner

Gate update: the first full-workspace Clippy run found one byte_char_slices warning in calternal-fs on the Search readiness marker. Commit c8736bbf8 changes the read-side comparison from [b'1'] to *b"1", matching the writer. cargo clippy -p calternal-fs --all-targets -- -D warnings and cargo test -p calternal-fs pass (42 tests). Full-workspace Clippy rerun is still compiling dependencies on the shared build host; no second diagnostic has appeared yet. The benchmark's 23 Python tests pass.

Gate update: the first full-workspace Clippy run found one `byte_char_slices` warning in `calternal-fs` on the Search readiness marker. Commit `c8736bbf8` changes the read-side comparison from `[b'1']` to `*b"1"`, matching the writer. `cargo clippy -p calternal-fs --all-targets -- -D warnings` and `cargo test -p calternal-fs` pass (42 tests). Full-workspace Clippy rerun is still compiling dependencies on the shared build host; no second diagnostic has appeared yet. The benchmark's 23 Python tests pass.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#417
No description provided.