OCR: evaluate pdf-inspector + PP-OCRv6 Small (Rust/ONNX, Apache/MIT), AnyDoc and PaddleOCR-VL on the owner's corpus vs #417 #584

Open
opened 2026-10-01 07:07:47 +00:00 by kayg · 24 comments
Owner

Owner question (2026-10-01): "are those the only libraries available? nothing more modern? like the firecrawl libraries?"

#417 compared Tesseract 5, ocrs 0.13 and PaddleOCR v5 mobile. Newer candidates, found with web research on 2026-10-01:

  • firecrawl/pdf-inspector (Rust, MIT, about 13k stars): reads PDF structure without rendering; classifies each page as text vs. needs-OCR in about 20 ms; extracts native text with reading order. Its OCR route uses PP-OCRv6 Small through oar-ocr (Apache-2.0) on ONNX Runtime + PDFium, CPU, about 31 MB of models, downloaded and SHA-256 pinned on first use. Clean text PDFs never load OCR. Sources: https://github.com/firecrawl/pdf-inspector/blob/main/docs/ocr-runtime.md , https://www.firecrawl.dev/blog/anydoc-and-pdf-inspector
  • firecrawl/anydoc (Rust, MIT): docx/doc/xlsx/xls/pptx/ppt/rtf/odt/ods/odp/epub/csv → markdown at about 4 ms per document. It would let Search index Office documents too.
  • PaddleOCR-VL 1.6 (0.9B, Apache-2.0, 109 languages; the top OmniDocBench score): for hard pages (tables, handwriting). Probably GPU-class cost: an optional "deep" pass only.
  • Ruled out by the licence rules: anything with non-commercial weights. DeepSeek-OCR (MIT, 3B) and olmOCR (Apache-2.0, 7B, 12 GB VRAM) are too heavy for the default.
    Job (measurement; the measurement-job protocol: Luna Max, a fixed decision rule up front, interleaved runs): evaluate on the owner's #417 corpus (aggregates only; private data stays in ~/calternal-private/; never commit or print document text):
  1. pdf-inspector's classification vs. the #417 labels (the 43 known textless PDFs + 16 uninspected): precision/recall of "needs OCR".
  2. PP-OCRv6 Small (oar-ocr, Rust/ONNX) vs. the #417 winners: proxy recall on the same 12 pages + 50 more textless pages; CPU seconds per page; peak RSS; on the perf VM (flock /root/perf.lock, 4 vCPU / 7 GB).
  3. AnyDoc on any Office files in the corpus: conversion success rate and time.
  4. PaddleOCR-VL 1.6 on CPU, 10 pages only, to see whether it is even feasible as an opt-in deep pass.
    Decision rule (fixed now): the default = the engine with the best proxy recall among those with ≤ 5 s CPU per page and ≤ 400 MB peak RSS on the perf VM; ties go to the smaller model and Rust-native integration. Report a table and a recommendation on this issue. No product code.
## Owner question (2026-10-01): "are those the only libraries available? nothing more modern? like the firecrawl libraries?" #417 compared Tesseract 5, ocrs 0.13 and PaddleOCR v5 mobile. Newer candidates, found with web research on 2026-10-01: - **firecrawl/pdf-inspector** (Rust, **MIT**, about 13k stars): reads PDF structure without rendering; classifies each page as text vs. needs-OCR in about 20 ms; extracts native text with reading order. Its OCR route uses **PP-OCRv6 Small** through `oar-ocr` (**Apache-2.0**) on ONNX Runtime + PDFium, **CPU**, about **31 MB** of models, downloaded and SHA-256 pinned on first use. Clean text PDFs never load OCR. Sources: https://github.com/firecrawl/pdf-inspector/blob/main/docs/ocr-runtime.md , https://www.firecrawl.dev/blog/anydoc-and-pdf-inspector - **firecrawl/anydoc** (Rust, **MIT**): docx/doc/xlsx/xls/pptx/ppt/rtf/odt/ods/odp/epub/csv → markdown at about 4 ms per document. It would let Search index Office documents too. - **PaddleOCR-VL 1.6** (0.9B, **Apache-2.0**, 109 languages; the top OmniDocBench score): for hard pages (tables, handwriting). Probably GPU-class cost: an optional "deep" pass only. - Ruled out by the licence rules: anything with non-commercial weights. DeepSeek-OCR (MIT, 3B) and olmOCR (Apache-2.0, 7B, 12 GB VRAM) are too heavy for the default. **Job (measurement; the measurement-job protocol: Luna Max, a fixed decision rule up front, interleaved runs):** evaluate on the owner's #417 corpus (aggregates only; private data stays in `~/calternal-private/`; never commit or print document text): 1. pdf-inspector's classification vs. the #417 labels (the 43 known textless PDFs + 16 uninspected): precision/recall of "needs OCR". 2. PP-OCRv6 Small (oar-ocr, Rust/ONNX) vs. the #417 winners: proxy recall on the same 12 pages + 50 more textless pages; CPU seconds per page; peak RSS; on the perf VM (`flock /root/perf.lock`, 4 vCPU / 7 GB). 3. AnyDoc on any Office files in the corpus: conversion success rate and time. 4. PaddleOCR-VL 1.6 on CPU, 10 pages only, to see whether it is even feasible as an opt-in deep pass. **Decision rule (fixed now):** the default = the engine with the best proxy recall among those with ≤ 5 s CPU per page and ≤ 400 MB peak RSS on the perf VM; ties go to the smaller model and Rust-native integration. Report a table and a recommendation on this issue. **No product code.**
Author
Owner

Starting OCR evaluation on job/ocr-eval-584, based on origin/dev at cc25c441b7. Worktree is clean. I will report aggregate measurements only; private document contents remain in ~/calternal-private.

Starting OCR evaluation on job/ocr-eval-584, based on origin/dev at cc25c441b7a974185622a1dee853cf38686d2b67. Worktree is clean. I will report aggregate measurements only; private document contents remain in ~/calternal-private.
Author
Owner

Finding: current crates.io metadata resolves pdf-inspector 1.25.2 (MIT), oar-ocr 0.9.2 (Apache-2.0), and anydoc 0.2.4 (MIT). The current pdf-inspector release includes PP-OCRv6 Small; its documented reproducible OCR runtime pins ONNX Runtime 1.27.0 and PP-OCRv6 artifact revision oar-ocr-v0.7.0. I will use these published pins for the local CPU comparison and retain #417’s PP-OCRv5 mobile ONNX numbers as the prior-run baseline.

Finding: current crates.io metadata resolves pdf-inspector 1.25.2 (MIT), oar-ocr 0.9.2 (Apache-2.0), and anydoc 0.2.4 (MIT). The current pdf-inspector release includes PP-OCRv6 Small; its documented reproducible OCR runtime pins ONNX Runtime 1.27.0 and PP-OCRv6 artifact revision oar-ocr-v0.7.0. I will use these published pins for the local CPU comparison and retain #417’s PP-OCRv5 mobile ONNX numbers as the prior-run baseline.
Author
Owner

Finding from the existing #417 locked VM run: pdf-inspector 1.25.2 inspected 439/455 PDFs; the other 16 remain uninspected. Of the inspected set, 43 were known textless and 396 had a text layer. It routed all 43 textless PDFs to OCR and retained 383/396 text-bearing PDFs as text, so document-level needs-OCR precision is 43/(43+13) = 76.8% and recall is 43/43 = 100%; the 16 uninspected PDFs are excluded. Page-level routing was 102 true OCR routes, 24 text pages routed unnecessarily and 5 textless pages missed: 81.0% precision and 95.3% recall. I will reuse these same-release #417 figures rather than duplicate that full-corpus pass.

Finding from the existing #417 locked VM run: pdf-inspector 1.25.2 inspected 439/455 PDFs; the other 16 remain uninspected. Of the inspected set, 43 were known textless and 396 had a text layer. It routed all 43 textless PDFs to OCR and retained 383/396 text-bearing PDFs as text, so document-level needs-OCR precision is 43/(43+13) = 76.8% and recall is 43/43 = 100%; the 16 uninspected PDFs are excluded. Page-level routing was 102 true OCR routes, 24 text pages routed unnecessarily and 5 textless pages missed: 81.0% precision and 95.3% recall. I will reuse these same-release #417 figures rather than duplicate that full-corpus pass.
Author
Owner

Finding during the benchmark build: pdf-inspector 1.25.2 does not compile with only its ocr-oar feature because vision::fusion imports complete_table_markdown_from_items, which is gated behind the bundled ocr feature. I am using the crate's documented ocr feature set for the rerun; the measurement feeds PNGs directly to OAR and does not load PDFium during inference.

Finding during the benchmark build: pdf-inspector 1.25.2 does not compile with only its `ocr-oar` feature because `vision::fusion` imports `complete_table_markdown_from_items`, which is gated behind the bundled `ocr` feature. I am using the crate's documented `ocr` feature set for the rerun; the measurement feeds PNGs directly to OAR and does not load PDFium during inference.
Author
Owner

Finding: private staging produced the exact 12-page mixed sample (6 text-layer, 6 textless), plus 50 distinct extra textless pages. It also staged all 158 Office inputs (140 DOCX, 15 XLSX, 3 XLS) under generated numeric names. PDFium could not inspect 5 source PDFs during scan selection; it still found the full 50-page sample. No source file names or document text were written to the repo or printed.

Finding: private staging produced the exact 12-page mixed sample (6 text-layer, 6 textless), plus 50 distinct extra textless pages. It also staged all 158 Office inputs (140 DOCX, 15 XLSX, 3 XLS) under generated numeric names. PDFium could not inspect 5 source PDFs during scan selection; it still found the full 50-page sample. No source file names or document text were written to the repo or printed.
Author
Owner

Finding: AnyDoc converted 157 of 158 staged Office documents (99.4%) on the local fallback host. Mean conversion CPU was 0.0183 s/document, wall p50/p95 was 0.0211/0.0334 s, and peak RSS was 476 MiB. I do not identify the failed input.

The perf VM lock was busy at the single availability check, so these timings are local. PaddleOCR-VL 1.6 failed before reading pages because its required paddlex[ocr] extra was absent from the benchmark environment. I am adding the pinned extra and will retry the ten-page CPU run. No corpus text or filenames are included here.

Finding: AnyDoc converted 157 of 158 staged Office documents (99.4%) on the local fallback host. Mean conversion CPU was 0.0183 s/document, wall p50/p95 was 0.0211/0.0334 s, and peak RSS was 476 MiB. I do not identify the failed input. The perf VM lock was busy at the single availability check, so these timings are local. PaddleOCR-VL 1.6 failed before reading pages because its required `paddlex[ocr]` extra was absent from the benchmark environment. I am adding the pinned extra and will retry the ten-page CPU run. No corpus text or filenames are included here.
Author
Owner

Finding: The #417 pdf-inspector 1.25.2 file-level run covered 439 PDFs: 43 true positive textless routes, 13 false positive OCR routes, and 0 false negatives. That is 76.8% precision and 100% recall; 16 uninspected PDFs are excluded rather than treated as negatives. The #417 page-level labels gave 102 TP, 24 FP and 5 FN (81.0% precision, 95.3% recall).

I interleaved PP-OCRv6 Small (pdf-inspector 1.25.2 / oar-ocr 0.9.2, ONNX Runtime 1.27.0) and the #417 PP-OCRv5 Mobile ONNX winner (PaddleOCR 3.7.0, ONNX Runtime 1.23.2) on 62 identical pages: the original 12 and 50 additional textless pages. Both completed 62/62. On the six text-layer reference pages, mean token recall was 95.4% for v6 and 96.1% for v5. Both emitted at least three token hashes on 46/50 added scans; this is nonempty-output coverage, not ground-truth accuracy.

The single-run local averages were v6 4.31 CPU s/page, wall p50/p95 1.86/5.84 s, peak RSS 822 MiB; v5 35.22 CPU s/page, wall p50/p95 5.96/20.91 s, peak RSS 1,654 MiB. The perf VM lock was busy at the availability check, so these numbers are local and the host was shared. Under the fixed <=5 CPU s/page and <=400 MiB rule, neither qualifies because both exceed the memory ceiling. I do not recommend promoting v6 as the default from this run; the measured candidates do not produce a qualifying default.

Finding: The #417 pdf-inspector 1.25.2 file-level run covered 439 PDFs: 43 true positive textless routes, 13 false positive OCR routes, and 0 false negatives. That is 76.8% precision and 100% recall; 16 uninspected PDFs are excluded rather than treated as negatives. The #417 page-level labels gave 102 TP, 24 FP and 5 FN (81.0% precision, 95.3% recall). I interleaved PP-OCRv6 Small (pdf-inspector 1.25.2 / oar-ocr 0.9.2, ONNX Runtime 1.27.0) and the #417 PP-OCRv5 Mobile ONNX winner (PaddleOCR 3.7.0, ONNX Runtime 1.23.2) on 62 identical pages: the original 12 and 50 additional textless pages. Both completed 62/62. On the six text-layer reference pages, mean token recall was 95.4% for v6 and 96.1% for v5. Both emitted at least three token hashes on 46/50 added scans; this is nonempty-output coverage, not ground-truth accuracy. The single-run local averages were v6 4.31 CPU s/page, wall p50/p95 1.86/5.84 s, peak RSS 822 MiB; v5 35.22 CPU s/page, wall p50/p95 5.96/20.91 s, peak RSS 1,654 MiB. The perf VM lock was busy at the availability check, so these numbers are local and the host was shared. Under the fixed <=5 CPU s/page and <=400 MiB rule, neither qualifies because both exceed the memory ceiling. I do not recommend promoting v6 as the default from this run; the measured candidates do not produce a qualifying default.
Author
Owner

Finding: After adding the pinned paddlex[ocr]==3.7.2 runtime extra, PaddleOCR-VL 1.6 initialized and entered its CPU run. The ten-page attempt did not produce its final aggregate within the one-hour local cap. I stopped it at 60 minutes; the process had used 38m35s CPU, and sampled RSS reached 3.9 GiB. The harness does not emit per-page progress, so I cannot claim how many pages completed or report a page latency/recall. The run is incomplete and does not qualify as a feasible CPU deep pass under the 400 MiB memory limit. The perf VM lock was unavailable; these figures are from the shared local host.

Finding: After adding the pinned `paddlex[ocr]==3.7.2` runtime extra, PaddleOCR-VL 1.6 initialized and entered its CPU run. The ten-page attempt did not produce its final aggregate within the one-hour local cap. I stopped it at 60 minutes; the process had used 38m35s CPU, and sampled RSS reached 3.9 GiB. The harness does not emit per-page progress, so I cannot claim how many pages completed or report a page latency/recall. The run is incomplete and does not qualify as a feasible CPU deep pass under the 400 MiB memory limit. The perf VM lock was unavailable; these figures are from the shared local host.
Author
Owner

#584 evaluation report

Head: 0c3235106d8dcaf369b58a6f2e272c3be6090723 (bench: add private OCR v6 evaluation harness; bench: pin PaddleOCR-VL OCR extra). No product code or private corpus data was committed.

The perf VM was occupied at the one availability check. Per the job rule, measurements ran on the local shared host and are labeled local. Inputs and OCR output stayed under ~/calternal-private/; issue comments contain aggregates only.

Candidate / check Aggregate result Decision
pdf-inspector 1.25.2 classification from #417 439 inspected PDFs: 43 TP, 13 FP, 0 FN; file-level precision 76.8%, recall 100%. The 16 uninspected PDFs are excluded. Available page labels: 102 TP, 24 FP, 5 FN; precision 81.0%, recall 95.3%. Reused the same-release #417 run; no duplicate corpus classification pass.
PP-OCRv6 Small, 62 pages 62/62 succeeded; mean token recall 95.4% on the six text-layer reference pages; at least three token hashes on 46/50 additional scans; 4.31 CPU s/page, wall p50/p95 1.86/5.84 s, peak RSS 822 MiB. Meets the 5 CPU s/page limit locally, but exceeds 400 MiB.
PP-OCRv5 Mobile ONNX, same 62 pages 62/62 succeeded; mean token recall 96.1% on the same six reference pages; nonempty on 46/50 additional scans; 35.22 CPU s/page, wall p50/p95 5.96/20.91 s, peak RSS 1,654 MiB. Exceeds both limits.
AnyDoc 0.2.4, 158 Office documents 157 succeeded (99.4%), 1 failed; 0.0183 CPU s/document; wall p50/p95 0.0211/0.0334 s; peak RSS 476 MiB. Failed input is not identified. Conversion result only.
PaddleOCR-VL 1.6 on CPU, ten-page attempt Did not finish within a 60-minute cap. At termination it had used 38m35s process CPU; sampled RSS reached 3.9 GiB. No page-count or page-timing aggregate was emitted. Incomplete and far above the memory ceiling; not feasible as a routine CPU pass in this run.

Token recall is the hashed-token overlap proxy. “Nonempty” means at least three unique token hashes; it is coverage, not ground-truth accuracy. The six-page recall sample is small. The 62-page timing and RSS figures come from the local shared host, not the 4-vCPU perf VM.

Decision

No measured OCR candidate meets both fixed limits (<=5 CPU s/page and <=400 MiB peak RSS). Do not promote PP-OCRv6 Small as the default from this run. PP-OCRv5 has slightly higher proxy recall, but it also misses both limits. PaddleOCR-VL did not complete its sample.

Decisions not covered by DESIGN.md

  • I reused pdf-inspector’s #417 same-release classification data rather than scan the corpus a second time.
  • I used the 20-byte extractable-text cutoff and 1600-pixel render size from #417, and compared both engines in alternating order on the same page images.
  • The perf lock was busy, so I used the permitted local fallback and labeled its results. I capped the PaddleOCR-VL attempt at one hour after it made no aggregate available; its exact completed-page count is unknown.
  • The PP-OCRv5 Python worker uses ONNX Runtime 1.23.2 as in #417. The Rust PP-OCRv6 worker uses the current pdf-inspector-pinned ONNX Runtime 1.27.0 library. PaddleOCR-VL required paddlex[ocr]==3.7.2, now pinned in the private benchmark requirements.

Files

  • bench/ocr-eval/src/main.rs, Cargo.toml, Cargo.lock: aggregate-only Rust OCR and AnyDoc runner.
  • bench/ocr-eval/stage.py, interleave.py, paddle_worker.py: private input staging and measurement workers.
  • bench/ocr-eval/requirements.txt, README.md: pinned evaluation environment and method.

Gates (verbatim output)

cargo fmt --manifest-path bench/ocr-eval/Cargo.toml -- --check: exit 0, no stdout.

cargo clippy --manifest-path bench/ocr-eval/Cargo.toml --all-targets -- -D warnings:

    Checking unicode-casefold v0.2.0
    Checking sha2 v0.10.9
    Checking anydoc v0.2.4
    Checking calternal-ocr-eval v0.1.0 (/home/kayg/Developer/calternal-wt/ocr-eval-584/bench/ocr-eval)
    Finished `dev` profile [unoptimized + debuginfo] target(s) in 46.05s

cargo test --manifest-path bench/ocr-eval/Cargo.toml:

   Compiling anydoc v0.2.4
   Compiling calternal-ocr-eval v0.1.0 (/home/kayg/Developer/calternal-wt/ocr-eval-584/bench/ocr-eval)
    Finished `test` profile [unoptimized + debuginfo] target(s) in 50.74s
     Running unittests src/main.rs (/mnt/hdd/targets/jobs/ocr-eval-584/debug/deps/calternal_ocr_eval-0a890e72c51184a3)

running 0 tests

test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s

Final cargo clean output:

     Removed 4166 files, 1.8GiB total

Python syntax check: Parsed 3 Python modules. git diff --check exited 0. The working tree is clean. I fetched and merged origin/dev once before gates; origin/dev advanced afterward, so this branch currently shows 2 commits ahead and 1 behind. I did not fetch or merge again under the one-time merge rule.

## #584 evaluation report Head: `0c3235106d8dcaf369b58a6f2e272c3be6090723` (`bench: add private OCR v6 evaluation harness`; `bench: pin PaddleOCR-VL OCR extra`). No product code or private corpus data was committed. The perf VM was occupied at the one availability check. Per the job rule, measurements ran on the local shared host and are labeled local. Inputs and OCR output stayed under `~/calternal-private/`; issue comments contain aggregates only. | Candidate / check | Aggregate result | Decision | | --- | --- | --- | | pdf-inspector 1.25.2 classification from #417 | 439 inspected PDFs: 43 TP, 13 FP, 0 FN; file-level precision 76.8%, recall 100%. The 16 uninspected PDFs are excluded. Available page labels: 102 TP, 24 FP, 5 FN; precision 81.0%, recall 95.3%. | Reused the same-release #417 run; no duplicate corpus classification pass. | | PP-OCRv6 Small, 62 pages | 62/62 succeeded; mean token recall 95.4% on the six text-layer reference pages; at least three token hashes on 46/50 additional scans; 4.31 CPU s/page, wall p50/p95 1.86/5.84 s, peak RSS 822 MiB. | Meets the 5 CPU s/page limit locally, but exceeds 400 MiB. | | PP-OCRv5 Mobile ONNX, same 62 pages | 62/62 succeeded; mean token recall 96.1% on the same six reference pages; nonempty on 46/50 additional scans; 35.22 CPU s/page, wall p50/p95 5.96/20.91 s, peak RSS 1,654 MiB. | Exceeds both limits. | | AnyDoc 0.2.4, 158 Office documents | 157 succeeded (99.4%), 1 failed; 0.0183 CPU s/document; wall p50/p95 0.0211/0.0334 s; peak RSS 476 MiB. Failed input is not identified. | Conversion result only. | | PaddleOCR-VL 1.6 on CPU, ten-page attempt | Did not finish within a 60-minute cap. At termination it had used 38m35s process CPU; sampled RSS reached 3.9 GiB. No page-count or page-timing aggregate was emitted. | Incomplete and far above the memory ceiling; not feasible as a routine CPU pass in this run. | Token recall is the hashed-token overlap proxy. “Nonempty” means at least three unique token hashes; it is coverage, not ground-truth accuracy. The six-page recall sample is small. The 62-page timing and RSS figures come from the local shared host, not the 4-vCPU perf VM. ### Decision No measured OCR candidate meets both fixed limits (<=5 CPU s/page and <=400 MiB peak RSS). Do not promote PP-OCRv6 Small as the default from this run. PP-OCRv5 has slightly higher proxy recall, but it also misses both limits. PaddleOCR-VL did not complete its sample. ### Decisions not covered by DESIGN.md - I reused pdf-inspector’s #417 same-release classification data rather than scan the corpus a second time. - I used the 20-byte extractable-text cutoff and 1600-pixel render size from #417, and compared both engines in alternating order on the same page images. - The perf lock was busy, so I used the permitted local fallback and labeled its results. I capped the PaddleOCR-VL attempt at one hour after it made no aggregate available; its exact completed-page count is unknown. - The PP-OCRv5 Python worker uses ONNX Runtime 1.23.2 as in #417. The Rust PP-OCRv6 worker uses the current pdf-inspector-pinned ONNX Runtime 1.27.0 library. PaddleOCR-VL required `paddlex[ocr]==3.7.2`, now pinned in the private benchmark requirements. ### Files - `bench/ocr-eval/src/main.rs`, `Cargo.toml`, `Cargo.lock`: aggregate-only Rust OCR and AnyDoc runner. - `bench/ocr-eval/stage.py`, `interleave.py`, `paddle_worker.py`: private input staging and measurement workers. - `bench/ocr-eval/requirements.txt`, `README.md`: pinned evaluation environment and method. ### Gates (verbatim output) `cargo fmt --manifest-path bench/ocr-eval/Cargo.toml -- --check`: exit 0, no stdout. `cargo clippy --manifest-path bench/ocr-eval/Cargo.toml --all-targets -- -D warnings`: ```text Checking unicode-casefold v0.2.0 Checking sha2 v0.10.9 Checking anydoc v0.2.4 Checking calternal-ocr-eval v0.1.0 (/home/kayg/Developer/calternal-wt/ocr-eval-584/bench/ocr-eval) Finished `dev` profile [unoptimized + debuginfo] target(s) in 46.05s ``` `cargo test --manifest-path bench/ocr-eval/Cargo.toml`: ```text Compiling anydoc v0.2.4 Compiling calternal-ocr-eval v0.1.0 (/home/kayg/Developer/calternal-wt/ocr-eval-584/bench/ocr-eval) Finished `test` profile [unoptimized + debuginfo] target(s) in 50.74s Running unittests src/main.rs (/mnt/hdd/targets/jobs/ocr-eval-584/debug/deps/calternal_ocr_eval-0a890e72c51184a3) running 0 tests test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s ``` Final `cargo clean` output: ```text Removed 4166 files, 1.8GiB total ``` Python syntax check: `Parsed 3 Python modules.` `git diff --check` exited 0. The working tree is clean. I fetched and merged `origin/dev` once before gates; `origin/dev` advanced afterward, so this branch currently shows 2 commits ahead and 1 behind. I did not fetch or merge again under the one-time merge rule.
Author
Owner

Owner decision (2026-10-01): adopt PP-OCRv6 Small through pdf-inspector (its ocr feature, SHA-256-pinned model set) for scanned PDF pages and images, and AnyDoc for Office documents. PaddleOCR-VL is out (CPU too slow).
Build notes: run OCR in a bounded background worker (one page at a time, memory cap on the worker, not the server) so the 822 MiB peak never lands in the server process; reuse the #417 text-layer/scan classifier; low priority behind interactive work.

**Owner decision (2026-10-01):** adopt **PP-OCRv6 Small through pdf-inspector** (its `ocr` feature, SHA-256-pinned model set) for scanned PDF pages and images, and **AnyDoc** for Office documents. PaddleOCR-VL is out (CPU too slow). Build notes: run OCR in a bounded background worker (one page at a time, memory cap on the worker, not the server) so the 822 MiB peak never lands in the server process; reuse the #417 text-layer/scan classifier; low priority behind interactive work.
Author
Owner

Starting #584 on job/ocr-build. Initial branch/base SHA: 687ff70313. Working tree is clean; I will merge origin/dev once before final gates as required.

Starting #584 on job/ocr-build. Initial branch/base SHA: 687ff703136e71e89f8dfba139e93cd0788b25c1. Working tree is clean; I will merge origin/dev once before final gates as required.
Author
Owner

Finding: the pinned product dependencies have permissive licenses with no non-commercial restriction. cargo info resolves pdf-inspector 1.25.2 (MIT) and AnyDoc 0.2.4 (MIT). The upstream ONNX Runtime license is MIT. pdf-inspector's OCR runtime docs pin ONNX Runtime 1.27.0 and the checksum-verified PP-OCRv6 Small artifact revision oar-ocr-v0.7.0; the model source license is Apache-2.0. These licenses allow commercial use and are compatible with this AGPL-3.0-only project. Sources: https://github.com/firecrawl/pdf-inspector/blob/main/LICENSE ; https://github.com/firecrawl/anydoc/blob/main/LICENSE ; https://github.com/microsoft/onnxruntime/blob/main/LICENSE ; https://github.com/firecrawl/pdf-inspector/blob/main/docs/ocr-runtime.md ; https://github.com/GreatV/oar-ocr/blob/main/LICENSE

Finding: the pinned product dependencies have permissive licenses with no non-commercial restriction. `cargo info` resolves pdf-inspector 1.25.2 (MIT) and AnyDoc 0.2.4 (MIT). The upstream ONNX Runtime license is MIT. pdf-inspector's OCR runtime docs pin ONNX Runtime 1.27.0 and the checksum-verified PP-OCRv6 Small artifact revision `oar-ocr-v0.7.0`; the model source license is Apache-2.0. These licenses allow commercial use and are compatible with this AGPL-3.0-only project. Sources: https://github.com/firecrawl/pdf-inspector/blob/main/LICENSE ; https://github.com/firecrawl/anydoc/blob/main/LICENSE ; https://github.com/microsoft/onnxruntime/blob/main/LICENSE ; https://github.com/firecrawl/pdf-inspector/blob/main/docs/ocr-runtime.md ; https://github.com/GreatV/oar-ocr/blob/main/LICENSE
Author
Owner

Finding: the Search indexer accepts the uploaded DOCX path and indexes its file name, but indexed_file() only calls the extractor for .pdf. I added a regression test using a test-only extractor; it fails because Search cannot find the emitted DOCX body word. The test output was: search did not find users/alice/Files/report.docx for "officebodyneedle". I am extending the existing indexing path so PDF, image, and AnyDoc Office text enter the same per-User Search generation.

Finding: the Search indexer accepts the uploaded DOCX path and indexes its file name, but `indexed_file()` only calls the extractor for `.pdf`. I added a regression test using a test-only extractor; it fails because Search cannot find the emitted DOCX body word. The test output was: `search did not find users/alice/Files/report.docx for "officebodyneedle"`. I am extending the existing indexing path so PDF, image, and AnyDoc Office text enter the same per-User Search generation.
Author
Owner

Finding: pdf-inspector's ocr feature uses dynamically loaded PDFium and ONNX Runtime libraries for scanned pages. The validated Linux runtime assets documented by pdf-inspector are Firecrawl PDFium native-v7988 and ONNX Runtime 1.27.0; the PP-OCRv6 Small model artifacts remain SHA-pinned by pdf-inspector. I am adding checksum-verified amd64/arm64 assets to the server image recipes and setting the model cache below calternal-fs's reserved .system directory. This keeps clean native-text PDFs and Office conversion off the OCR runtime path.

Finding: pdf-inspector's `ocr` feature uses dynamically loaded PDFium and ONNX Runtime libraries for scanned pages. The validated Linux runtime assets documented by pdf-inspector are Firecrawl PDFium `native-v7988` and ONNX Runtime `1.27.0`; the PP-OCRv6 Small model artifacts remain SHA-pinned by pdf-inspector. I am adding checksum-verified amd64/arm64 assets to the server image recipes and setting the model cache below calternal-fs's reserved `.system` directory. This keeps clean native-text PDFs and Office conversion off the OCR runtime path.
Author
Owner

Finding: the full cargo test -p calternal-search run had one timeout in the existing rebuild_matches_incremental_results test: search did not find users/alice/Files/first.txt for "phrase" after its existing 5-second wait. The file is .txt, so it stays on the unchanged text indexing branch. An isolated rerun passed verbatim in 5.27s (1 passed; 0 failed; 0 ignored). The first full run followed an 18m25s dependency build while the shared host was running other Rust builds. This points to host contention, not a stable product failure. I kept the existing expectation unchanged.

Finding: the full `cargo test -p calternal-search` run had one timeout in the existing `rebuild_matches_incremental_results` test: `search did not find users/alice/Files/first.txt for "phrase"` after its existing 5-second wait. The file is `.txt`, so it stays on the unchanged text indexing branch. An isolated rerun passed verbatim in 5.27s (`1 passed; 0 failed; 0 ignored`). The first full run followed an 18m25s dependency build while the shared host was running other Rust builds. This points to host contention, not a stable product failure. I kept the existing expectation unchanged.
Author
Owner

Finding: pdf-inspector 1.25.2 documents that its ocr feature requires Rust 1.95 or newer, while this workspace declares Rust 1.94. Its pinned PP-OCRv6 Small artifact manifest also identifies OAR 0.9.1. I will pin OAR 0.9.1 and raise only calternal-server's declared MSRV to 1.95, leaving the other workspace crates at 1.94. This keeps the selected OCR feature and model manifest aligned without changing the workspace-wide MSRV.

Finding: pdf-inspector 1.25.2 documents that its `ocr` feature requires Rust 1.95 or newer, while this workspace declares Rust 1.94. Its pinned PP-OCRv6 Small artifact manifest also identifies OAR 0.9.1. I will pin OAR 0.9.1 and raise only `calternal-server`'s declared MSRV to 1.95, leaving the other workspace crates at 1.94. This keeps the selected OCR feature and model manifest aligned without changing the workspace-wide MSRV.
Author
Owner

Finding: the initial cargo check -p calternal-server reached the application crate and exposed that the OCR module needs the Search crate's existing MAX_PDF_PAGES constant re-exported. The server also embeds apps/web/build, which was absent in this worktree. I stopped this failed check to avoid keeping a high-memory rustc process alive, and will build the production web assets, expose the constant, and rerun the check.

Finding: the initial `cargo check -p calternal-server` reached the application crate and exposed that the OCR module needs the Search crate's existing `MAX_PDF_PAGES` constant re-exported. The server also embeds `apps/web/build`, which was absent in this worktree. I stopped this failed check to avoid keeping a high-memory rustc process alive, and will build the production web assets, expose the constant, and rerun the check.
Author
Owner

Finding: pinning only oar-ocr = 0.9.1 still lets Cargo resolve oar-ocr-core = 0.9.2 because OAR declares a compatible-version range. pdf-inspector 1.25.2 states its PP-OCRv6 Small artifact hashes match OAR core 0.9.1. I will pin both OAR crates to 0.9.1, with SIMD enabled and runtime downloading disabled, so the model manifest and engine registry stay aligned.

Finding: pinning only `oar-ocr = 0.9.1` still lets Cargo resolve `oar-ocr-core = 0.9.2` because OAR declares a compatible-version range. pdf-inspector 1.25.2 states its PP-OCRv6 Small artifact hashes match OAR core 0.9.1. I will pin both OAR crates to 0.9.1, with SIMD enabled and runtime downloading disabled, so the model manifest and engine registry stay aligned.
Author
Owner

Finding: bun run test exited 1 with 1,048 passed and five existing component tests timing out at their unchanged 5,000 ms limit: OverlaySurface, agenda, StatRow, MailSection and Composer. Vitest reported 2,130.75 seconds of transform time across the run. The failures are outside the Jobs API files changed here, and several Cargo/Vitest jobs were active on the shared host. I am keeping all existing expectations unchanged and will run the changed Jobs API test file alone.

Finding: `bun run test` exited 1 with 1,048 passed and five existing component tests timing out at their unchanged 5,000 ms limit: OverlaySurface, agenda, StatRow, MailSection and Composer. Vitest reported 2,130.75 seconds of transform time across the run. The failures are outside the Jobs API files changed here, and several Cargo/Vitest jobs were active on the shared host. I am keeping all existing expectations unchanged and will run the changed Jobs API test file alone.
Author
Owner

Finding: the first cargo check -p calternal-server compile of the worker reported two type errors: ocr_model_cache_directory() returns calternal_fs::Error, which does not convert to io::Error, and the worker Drop implementation cannot move a JoinHandle out of &mut self. Clippy would also reject one unused ChildStdout import. I will map the filesystem error to a generic I/O error, store the join handle in an Option so Drop can take it, and remove the import.

Finding: the first `cargo check -p calternal-server` compile of the worker reported two type errors: `ocr_model_cache_directory()` returns `calternal_fs::Error`, which does not convert to `io::Error`, and the worker Drop implementation cannot move a `JoinHandle` out of `&mut self`. Clippy would also reject one unused `ChildStdout` import. I will map the filesystem error to a generic I/O error, store the join handle in an Option so Drop can take it, and remove the import.
Author
Owner

Pipeline integration committed as 2e14b21a0 after wiring the bounded OCR worker, Search backfill version marker, runtime libraries, migration, and PDF/DOCX privacy e2e. The final cargo check -p calternal-server is still compiling with no diagnostics; process inspection shows rustc waiting on host resources. At the time of the check, uptime reported load average 116.51, 131.20, 133.43; /proc/pressure/io reported some avg10=95.53 avg60=93.57 avg300=93.98 and full avg10=8.17 avg60=15.33 avg300=21.02. I am waiting for this running check before starting another Cargo command.

Pipeline integration committed as `2e14b21a0` after wiring the bounded OCR worker, Search backfill version marker, runtime libraries, migration, and PDF/DOCX privacy e2e. The final `cargo check -p calternal-server` is still compiling with no diagnostics; process inspection shows rustc waiting on host resources. At the time of the check, `uptime` reported load average `116.51, 131.20, 133.43`; `/proc/pressure/io` reported `some avg10=95.53 avg60=93.57 avg300=93.98` and `full avg10=8.17 avg60=15.33 avg300=21.02`. I am waiting for this running check before starting another Cargo command.
Author
Owner

calternal-search test gate is still running. It reached the Search test target but has not emitted test results. The Cargo process has been active for about 25 minutes; uptime reported load average 189.08, 156.10, 141.13, and the target linker was in D state waiting on disk I/O for about 9 minutes. The gate also reported sccache: warning: The server looks like it shut down unexpectedly, compiling locally instead. No code or test expectations were changed in response.

`calternal-search` test gate is still running. It reached the Search test target but has not emitted test results. The Cargo process has been active for about 25 minutes; `uptime` reported load average `189.08, 156.10, 141.13`, and the target linker was in `D` state waiting on disk I/O for about 9 minutes. The gate also reported `sccache: warning: The server looks like it shut down unexpectedly, compiling locally instead`. No code or test expectations were changed in response.
Author
Owner

The required cargo test -p calternal-search run completed: 36 passed; 1 failed; 1 ignored in 133.08s. The only failure was the unchanged indexer::tests::staged_publication_waits_for_search_readers_before_swapping_directories assertion: staged publication did not wait for the active Search reader. Its code has a 5-second wait for the rebuild task to queue behind a read lock. This run had host load 146.80, 150.59, 148.22 and severe I/O pressure; no indexer.rs changes are in this branch. An earlier isolated run of the same pre-existing test passed, as reported above. I kept its assertion unchanged and did not rerun the gate.

The required `cargo test -p calternal-search` run completed: `36 passed; 1 failed; 1 ignored` in `133.08s`. The only failure was the unchanged `indexer::tests::staged_publication_waits_for_search_readers_before_swapping_directories` assertion: `staged publication did not wait for the active Search reader`. Its code has a 5-second wait for the rebuild task to queue behind a read lock. This run had host load `146.80, 150.59, 148.22` and severe I/O pressure; no `indexer.rs` changes are in this branch. An earlier isolated run of the same pre-existing test passed, as reported above. I kept its assertion unchanged and did not rerun the gate.
Author
Owner

#584 implementation report — stopped at the shared four-hour limit

Branch: job/ocr-build
Head: d6ee2d44da893d6264f3f58e9fd2e752b17cbcd7

Commits:

  • dc54c0885 fs: expose private OCR model cache handle
  • 8c2d22a83 jobs: group OCR backfill under Search
  • 2e14b21a0 server: add isolated OCR search pipeline
  • d6ee2d44d bench: add OCR backfill profile
  • The OCR evaluation runner was cherry-picked from job/ocr-eval-584 (18fc59529, dfef85d03).

Built

  • Added bounded PDF/image OCR and Office extraction in a persistent child worker, using the pinned PP-OCRv6 Small models and AnyDoc. Search file reads remain through calternal-fs; the worker receives bytes and an extension, never a Home path. Search document text uses the existing private index and preview path.
  • Added an OCR version marker to Search migration 0004. The existing per-User low-priority rebuild job reports OCR progress and marks completion only after a full private rebuild. New uploads use the existing Search indexer.
  • Grouped system.search.* jobs under Search and added a Jobs API regression assertion.
  • Added checksum-pinned PDFium and ONNX Runtime assets to both runtime image recipes, plus a synthetic PDF/DOCX privacy E2E and an aggregate-only 12/128-page performance profile.

Files: Cargo.toml, Cargo.lock, Containerfile, apps/web/e2e/ocr-584.mjs, apps/web/src/lib/jobs/api.ts, apps/web/src/lib/jobs/api.test.ts, bench/ocr-eval/*, crates/calternal-fs/src/{lib.rs,root.rs}, crates/calternal-search/migrations/0004_search_ocr_versions.sql, crates/calternal-search/src/{indexer.rs,lib.rs,pdf.rs}, crates/calternal-search/tests/indexer.rs, crates/calternal-server/Cargo.toml, crates/calternal-server/src/{main.rs,ocr.rs,wire.rs}, deploy/Containerfile.runtime, and deploy/fetch-ocr-runtime.

License checks were recorded earlier on this issue: pdf-inspector 1.25.2 and AnyDoc 0.2.4 are MIT; ONNX Runtime 1.27.0 is MIT; the pinned PP-OCRv6 Small model revision is Apache-2.0. These permit commercial use and are compatible with AGPL-3.0-only.

Gates and verification

  • cargo fmt --check: exit 0, no output.
  • cargo check -p calternal-server: exit 0. Output: Finished `dev` profile [unoptimized + debuginfo] target(s) in 27m 44s.
  • cargo clippy -p calternal-search --all-targets -- -D warnings: exit 0. Output: Finished `dev` profile [unoptimized + debuginfo] target(s) in 6m 49s.
  • cargo test -p calternal-search: test result: FAILED. 36 passed; 1 failed; 1 ignored; 0 measured; 0 filtered out; finished in 133.08s. The unchanged staged_publication_waits_for_search_readers_before_swapping_directories test hit its existing 5-second writer-queue deadline under high shared-host load. Its same test passed in an earlier isolated run; no expectation was changed and this gate was not rerun.
  • cargo clippy -p calternal-server --all-targets -- -D warnings: exit 0. Output: Finished `dev` profile [unoptimized + debuginfo] target(s) in 30m 29s.
  • cargo test -p calternal-server: interrupted with exit code 130 at the shared four-hour stop point while compiling calternal_server; the test runner did not launch and Cargo emitted no final summary.
  • Earlier calternal-fs fmt/clippy/test gates passed before that slice was committed. cargo clean later removed 17,973 files / 8.7 GiB.
  • bun run check: svelte-check found 0 errors and 0 warnings (exit 0).
  • bun run test: Test Files 5 failed | 148 passed (153); Tests 5 failed | 1048 passed (1053); Duration 496.98s. The five unchanged failures were existing 5-second component timeouts in OverlaySurface, Agenda, Composer, MailSection, and StatRow.
  • bun run test --maxWorkers=1 src/lib/jobs/api.test.ts: Test Files 1 passed (1); Tests 5 passed (5); Duration 112.49s.
  • The production web build, node --check apps/web/e2e/ocr-584.mjs, sh -n deploy/fetch-ocr-runtime, and Python profile syntax/help checks passed.

UX gaps closed

  • Search rebuild jobs now appear in the Search group with the Search icon, and the job reports when OCR text starts and completes.

UX gaps left

  • The real-server scanned-PDF/DOCX upload and cross-User Search E2E was not run. Protocol unit tests exist, but no real local-server adversarial round was run.
  • No Jobs screenshots were captured at 390/820/1440 in light/dark with macOS Playwright emulation. The Jobs interaction and accessibility walk is therefore incomplete.
  • The OCR profile was not run, and docs/perf/baseline.json has no OCR measurement.
  • Direct Mac GUI checks remain pending because the Mac VM is offline under the current owner rule.

Decisions not specified in DESIGN.md

  • Set only calternal-server to Rust 1.95 because pinned OAR 0.9.1 requires it; other workspace crates remain at their existing MSRV.
  • Chose one serialized warm OCR worker per server, a 2 GiB address-space cap, a 20-minute request deadline, and nice/idle-I/O priority.
  • Chose the existing 16 MiB document input limit, 128-page PDF limit, 150 DPI rendering, 16 MP / 8192-pixel image decoder bounds, 64 MiB rendered-page cap, and 256 KiB extracted-text cap.
  • The image path processes the first GIF frame. The OCR model cache is rooted through an inherited calternal-fs directory handle under .system.
  • Backfill enqueue reads Users in pages of 256 and uses the existing Search rebuild queue.

Migration audit: origin/dev contains Search migrations 0001–0003; 0004 is free.
Screenshots are not attached because the E2E did not run. No push, deploy, merge, or issue close was performed.

#584 implementation report — stopped at the shared four-hour limit Branch: `job/ocr-build` Head: `d6ee2d44da893d6264f3f58e9fd2e752b17cbcd7` Commits: - `dc54c0885` fs: expose private OCR model cache handle - `8c2d22a83` jobs: group OCR backfill under Search - `2e14b21a0` server: add isolated OCR search pipeline - `d6ee2d44d` bench: add OCR backfill profile - The OCR evaluation runner was cherry-picked from `job/ocr-eval-584` (`18fc59529`, `dfef85d03`). ## Built - Added bounded PDF/image OCR and Office extraction in a persistent child worker, using the pinned PP-OCRv6 Small models and AnyDoc. Search file reads remain through calternal-fs; the worker receives bytes and an extension, never a Home path. Search document text uses the existing private index and preview path. - Added an OCR version marker to Search migration 0004. The existing per-User low-priority rebuild job reports OCR progress and marks completion only after a full private rebuild. New uploads use the existing Search indexer. - Grouped `system.search.*` jobs under Search and added a Jobs API regression assertion. - Added checksum-pinned PDFium and ONNX Runtime assets to both runtime image recipes, plus a synthetic PDF/DOCX privacy E2E and an aggregate-only 12/128-page performance profile. Files: `Cargo.toml`, `Cargo.lock`, `Containerfile`, `apps/web/e2e/ocr-584.mjs`, `apps/web/src/lib/jobs/api.ts`, `apps/web/src/lib/jobs/api.test.ts`, `bench/ocr-eval/*`, `crates/calternal-fs/src/{lib.rs,root.rs}`, `crates/calternal-search/migrations/0004_search_ocr_versions.sql`, `crates/calternal-search/src/{indexer.rs,lib.rs,pdf.rs}`, `crates/calternal-search/tests/indexer.rs`, `crates/calternal-server/Cargo.toml`, `crates/calternal-server/src/{main.rs,ocr.rs,wire.rs}`, `deploy/Containerfile.runtime`, and `deploy/fetch-ocr-runtime`. License checks were recorded earlier on this issue: pdf-inspector 1.25.2 and AnyDoc 0.2.4 are MIT; ONNX Runtime 1.27.0 is MIT; the pinned PP-OCRv6 Small model revision is Apache-2.0. These permit commercial use and are compatible with AGPL-3.0-only. ## Gates and verification - `cargo fmt --check`: exit 0, no output. - `cargo check -p calternal-server`: exit 0. Output: ``Finished `dev` profile [unoptimized + debuginfo] target(s) in 27m 44s``. - `cargo clippy -p calternal-search --all-targets -- -D warnings`: exit 0. Output: ``Finished `dev` profile [unoptimized + debuginfo] target(s) in 6m 49s``. - `cargo test -p calternal-search`: `test result: FAILED. 36 passed; 1 failed; 1 ignored; 0 measured; 0 filtered out; finished in 133.08s`. The unchanged `staged_publication_waits_for_search_readers_before_swapping_directories` test hit its existing 5-second writer-queue deadline under high shared-host load. Its same test passed in an earlier isolated run; no expectation was changed and this gate was not rerun. - `cargo clippy -p calternal-server --all-targets -- -D warnings`: exit 0. Output: ``Finished `dev` profile [unoptimized + debuginfo] target(s) in 30m 29s``. - `cargo test -p calternal-server`: interrupted with exit code 130 at the shared four-hour stop point while compiling `calternal_server`; the test runner did not launch and Cargo emitted no final summary. - Earlier `calternal-fs` fmt/clippy/test gates passed before that slice was committed. `cargo clean` later removed 17,973 files / 8.7 GiB. - `bun run check`: `svelte-check found 0 errors and 0 warnings` (exit 0). - `bun run test`: `Test Files 5 failed | 148 passed (153)`; `Tests 5 failed | 1048 passed (1053)`; `Duration 496.98s`. The five unchanged failures were existing 5-second component timeouts in OverlaySurface, Agenda, Composer, MailSection, and StatRow. - `bun run test --maxWorkers=1 src/lib/jobs/api.test.ts`: `Test Files 1 passed (1)`; `Tests 5 passed (5)`; `Duration 112.49s`. - The production web build, `node --check apps/web/e2e/ocr-584.mjs`, `sh -n deploy/fetch-ocr-runtime`, and Python profile syntax/help checks passed. ## UX gaps closed - Search rebuild jobs now appear in the Search group with the Search icon, and the job reports when OCR text starts and completes. ## UX gaps left - The real-server scanned-PDF/DOCX upload and cross-User Search E2E was not run. Protocol unit tests exist, but no real local-server adversarial round was run. - No Jobs screenshots were captured at 390/820/1440 in light/dark with macOS Playwright emulation. The Jobs interaction and accessibility walk is therefore incomplete. - The OCR profile was not run, and `docs/perf/baseline.json` has no OCR measurement. - Direct Mac GUI checks remain pending because the Mac VM is offline under the current owner rule. ## Decisions not specified in DESIGN.md - Set only `calternal-server` to Rust 1.95 because pinned OAR 0.9.1 requires it; other workspace crates remain at their existing MSRV. - Chose one serialized warm OCR worker per server, a 2 GiB address-space cap, a 20-minute request deadline, and nice/idle-I/O priority. - Chose the existing 16 MiB document input limit, 128-page PDF limit, 150 DPI rendering, 16 MP / 8192-pixel image decoder bounds, 64 MiB rendered-page cap, and 256 KiB extracted-text cap. - The image path processes the first GIF frame. The OCR model cache is rooted through an inherited calternal-fs directory handle under `.system`. - Backfill enqueue reads Users in pages of 256 and uses the existing Search rebuild queue. Migration audit: `origin/dev` contains Search migrations 0001–0003; 0004 is free. Screenshots are not attached because the E2E did not run. No push, deploy, merge, or issue close was performed.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#584
No description provided.