DOCUMENTS: a Documents plugin with paperless-ngx features — files and plain text, not a database (grill first) #341

Open
opened 2026-09-28 13:26:13 +00:00 by kayg · 10 comments
Owner

Owner (2026-09-28): 'we need to have a "Documents" plugin for paperless-ngx like features. i have all my documents there but since it's a database, i barely use it. I vastly prefer filesystems and plaintext over a database!'

Principle: file over app. Documents live as ordinary files in Home (Documents/); any metadata (title, correspondent, document type, tags, dates, amounts, notes) is plain text that survives without calternal (e.g. a sidecar .md per document, or front matter in a companion note, or tags via the existing tag system); the Index is only a rebuildable cache.
Paperless-ngx features to study (research first, then grill the owner): OCR of scans/photos (local only, like voice; Tesseract/ocrmypdf licences, AGPL-compatible), full-text search of OCR text in the global palette (#258), auto-classification (correspondent, type, tags, date detection; local, learn from the owner's corrections, like Money's payee memory), consumption (a watched inbox folder, email attachments from Mail #313, the mobile scanner / share sheet, WebDAV upload #339), archive PDF/A with the text layer next to the original, workflows (→ Automations #296), saved views (→ smart folders/saved searches), permissions/sharing (Files Shares), dates on the Calendar (bills, warranties, expiry reminders → Tasks), links to Money transactions (receipts, #316).
Deliver docs/research/documents-plugin.md with the feature map, a plain-text metadata format proposal (readable, extensible, like the Money format spike #335), and numbered grill questions. No implementation before the grill.
Related: the owner's paperless-ngx library is imported to ~/Documents on calternal.cloud (import issue), with the paperless manifest kept so metadata can be migrated into the chosen format later.

Owner (2026-09-28): 'we need to have a "Documents" plugin for paperless-ngx like features. i have all my documents there but since it's a database, i barely use it. I vastly prefer filesystems and plaintext over a database!' Principle: **file over app**. Documents live as ordinary files in Home (`Documents/`); any metadata (title, correspondent, document type, tags, dates, amounts, notes) is plain text that survives without calternal (e.g. a sidecar `.md` per document, or front matter in a companion note, or tags via the existing tag system); the Index is only a rebuildable cache. Paperless-ngx features to study (research first, then grill the owner): **OCR** of scans/photos (local only, like voice; Tesseract/ocrmypdf licences, AGPL-compatible), full-text search of OCR text in the global palette (#258), **auto-classification** (correspondent, type, tags, date detection; local, learn from the owner's corrections, like Money's payee memory), **consumption** (a watched inbox folder, email attachments from Mail #313, the mobile scanner / share sheet, WebDAV upload #339), archive PDF/A with the text layer next to the original, **workflows** (→ Automations #296), saved views (→ smart folders/saved searches), permissions/sharing (Files Shares), dates on the Calendar (bills, warranties, expiry reminders → Tasks), links to Money transactions (receipts, #316). Deliver docs/research/documents-plugin.md with the feature map, a plain-text metadata format proposal (readable, extensible, like the Money format spike #335), and numbered grill questions. No implementation before the grill. Related: the owner's paperless-ngx library is imported to `~/Documents` on calternal.cloud (import issue), with the paperless manifest kept so metadata can be migrated into the chosen format later.
Author
Owner

Starting research on job/documents-research, based on c46515046871936ad681c4d00e8b76de69dde3b7 (the current dev base). I will keep the paperless inspection read-only and report only aggregate metadata counts, with no document names or content.

Starting research on `job/documents-research`, based on `c46515046871936ad681c4d00e8b76de69dde3b7` (the current `dev` base). I will keep the paperless inspection read-only and report only aggregate metadata counts, with no document names or content.
Author
Owner

Finding: the configured VPN host resolves to its configured address, but the read-only Podman listing fails with Permission denied (publickey). I have not queried paperless or read any document data. I am checking for a local read-only mount and continuing the independent research and OCR benchmark.

Finding: the configured VPN host resolves to its configured address, but the read-only Podman listing fails with `Permission denied (publickey)`. I have not queried paperless or read any document data. I am checking for a local read-only mount and continuing the independent research and OCR benchmark.
Author
Owner

Owner (2026-09-29): D1 file over app like every mode; documents can live anywhere in Home, default ~/Documents. D2 core agreed; research of the best paperless-ngx features is running (job documents-research). D3 asked why ingestion matters (to refine). D4 all changes are metadata only — calternal never moves or renames files for classification; folders do not matter to Documents. The owner asked Claude to refine the feature before more questions.

Owner (2026-09-29): D1 file over app like every mode; documents can live anywhere in Home, default ~/Documents. D2 core agreed; research of the best paperless-ngx features is running (job documents-research). D3 asked why ingestion matters (to refine). D4 all changes are metadata only — calternal never moves or renames files for classification; folders do not matter to Documents. The owner asked Claude to refine the feature before more questions.
Author
Owner

Finding: OCRmyPDF is MPL-2.0, and its own contributor guide says GPLv3/AGPLv3 integration is likely incompatible. The host measurement used an ephemeral local CLI only; the report will not recommend bundling OCRmyPDF under calternal's current AGPL-3.0-only license. Source: https://github.com/ocrmypdf/OCRmyPDF/blob/main/docs/contributing.md

Finding: OCRmyPDF is MPL-2.0, and its own contributor guide says GPLv3/AGPLv3 integration is likely incompatible. The host measurement used an ephemeral local CLI only; the report will not recommend bundling OCRmyPDF under calternal's current AGPL-3.0-only license. Source: https://github.com/ocrmypdf/OCRmyPDF/blob/main/docs/contributing.md
Author
Owner

Finding: Paperless parses its initial document date separately from the Auto classifier. Its configuration says it reads dates from filenames and content and uses the first one; the usage guide warns this can be wrong when OCR is poor or the document has several dates. The report treats parsed dates as suggestions. Paperless storage paths also move archive files, which conflicts with calternal DESIGN §40's rule that folders do not define meaning and calternal does not move files on its own. Source: https://docs.paperless-ngx.com/configuration/

Finding: Paperless parses its initial document date separately from the Auto classifier. Its configuration says it reads dates from filenames and content and uses the first one; the usage guide warns this can be wrong when OCR is poor or the document has several dates. The report treats parsed dates as suggestions. Paperless storage paths also move archive files, which conflicts with calternal DESIGN §40's rule that folders do not define meaning and calternal does not move files on its own. Source: https://docs.paperless-ngx.com/configuration/
Author
Owner

Owner decisions (2026-09-29):

  • R1 (agreed): documents are chosen by their content, not their folder. PDFs and office files anywhere in Home count. Images count only when they are clearly paper (OCR finds a page of text, such as a photographed receipt or a scanned letter) or when they sit in ~/Documents. Ordinary photos stay in Photos.
  • R2: the owner wants Immich-style automatic enrichment: auto-tagging and auto-classification. Folder structure does not matter. Metadata such as faces, memories and location is tagged automatically.
  • Open (grill in progress): which metadata stays app-over-file (derived, in the index) and which becomes file-over-app (written next to the document). The current rule is DESIGN §"Sidecars": derived data is never written into the user tree; human decisions are files.
Owner decisions (2026-09-29): - R1 (agreed): documents are chosen by their content, not their folder. PDFs and office files anywhere in Home count. Images count only when they are clearly paper (OCR finds a page of text, such as a photographed receipt or a scanned letter) or when they sit in `~/Documents`. Ordinary photos stay in Photos. - R2: the owner wants Immich-style automatic enrichment: auto-tagging and auto-classification. Folder structure does not matter. Metadata such as faces, memories and location is tagged automatically. - Open (grill in progress): which metadata stays app-over-file (derived, in the index) and which becomes file-over-app (written next to the document). The current rule is DESIGN §"Sidecars": derived data is never written into the user tree; human decisions are files.
Author
Owner

Owner decisions (2026-09-29, round 2):

  • Q1 agreed: the app works out metadata (OCR text, type, sender, dates, amounts, faces, place, tags, memories), keeps it in the index only, and can rebuild it. A confirm, correction or rejection by the User is a human decision and goes to a file. Rejections are stored, so the app does not add a rejected tag again.
  • Q3: automatic tags are applied at once, but only when the app is confident. Below the confidence bar a tag is not applied (details in the next grill round).
  • Q4: a photo of paper shows in Documents only (Photos hides it and has a filter for it). How the app finds these photos is being grilled next.
  • Q5: recognition runs locally only. The docs site to record these conventions is #415.
Owner decisions (2026-09-29, round 2): - Q1 agreed: the app works out metadata (OCR text, type, sender, dates, amounts, faces, place, tags, memories), keeps it in the index only, and can rebuild it. A confirm, correction or rejection by the User is a human decision and goes to a file. Rejections are stored, so the app does not add a rejected tag again. - Q3: automatic tags are applied at once, but only when the app is confident. Below the confidence bar a tag is not applied (details in the next grill round). - Q4: a photo of paper shows in Documents only (Photos hides it and has a filter for it). How the app finds these photos is being grilled next. - Q5: recognition runs locally only. The docs site to record these conventions is #415.
Author
Owner

Owner decision Q11 (2026-09-29): sidecars keep their standard names, so other apps find them. calternal hides them by default and carries each one with its parent file. Settings → Files → Hidden files lets the User show them: #420. Q10 (what goes into the XMP) is still open: the owner asked whether documents have their own metadata standard.

Owner decision Q11 (2026-09-29): sidecars keep their standard names, so other apps find them. calternal hides them by default and carries each one with its parent file. Settings → Files → Hidden files lets the User show them: #420. Q10 (what goes into the XMP) is still open: the owner asked whether documents have their own metadata standard.
Author
Owner

Research complete for #341. No product code changed.

Artifact: docs/research/documents-plugin.md (335 lines), committed as 28303e55 and merged once with local dev at 1ad7b728. Branch head: 4f3d7162aeef4d163472d817cd197f1fb9c06c3c; pushed job/documents-research (remote was already current). The note covers the Paperless feature map, local-only OCR and PDF extraction, local classification, two readable Markdown metadata examples, rename behavior, fit with DESIGN §§4/17/32/33/40/45/48, sources, recommendations, and 12 numbered grill questions.

Evidence and limits:

  • Synthetic image-only A4 scan at 300 dpi, English, one worker; Tesseract 5.5.3 through OCRmyPDF 17.11.0. First page: 30.97 s, 198 MiB max RSS; warm repeat: 10.79 s, 199 MiB; 20 pages: 2m45.33s, 186 MiB, 8.27 s/page. This is a single host run, not an accuracy or concurrency benchmark.
  • The configured Paperless host denied SSH with Permission denied (publickey). No Paperless database/API query was made. Actual imported-library field names and populated counts remain unavailable; the note clearly separates standard Paperless fields from observed library facts. No document names, filenames, titles, OCR text, or metadata values were accessed or reported.
  • License note: Tesseract and its best trained data are Apache-2.0. OCRmyPDF is MPL-2.0; its upstream cautions that GPLv3/AGPLv3 integration is likely incompatible, so the note uses it only for the local CLI benchmark and does not recommend bundling it. PaddleOCR/docTR are candidates pending per-model/dependency license and offline checks. Apple Vision is proprietary and Apple-only.

Decisions recommended where DESIGN is silent: make a readable Markdown Sidecar per document the canonical metadata source, retain a folder catalog Note as an alternative, keep OCR/classifier output in Derived data, save PDF/A only by explicit action, use explicit rules before considering a local classifier, and avoid automatic storage-path moves under §40. These are research recommendations for the grill, not product decisions.

Validation (verbatim gate output):

cargo fmt --check produced no stdout or stderr; exit status 0.

cargo clippy --all-targets -- -D warnings:

Finished `dev` profile [unoptimized + debuginfo] target(s) in 149m 26s

Exit status 0.

cargo test was capped at 45 minutes. It exited 124 while compiling and did not reach test execution. Last captured Cargo output:

   Compiling mail-parser v0.11.9
   Compiling elliptic-curve v0.13.8
   Compiling synstructure v0.13.2

bun run check and bun run test were not run because the job passed the approximately four-hour ceiling; remaining gates are open.

One scoped adversarial attempt (ADVERSARIAL_API_ONLY=1 tests/adversarial/run.sh) also reached its 20-minute setup cap while compiling the isolated server. It was stopped before the HTTP probe; this is not a clean probe result.

cargo clean ran. Removed apps/web/build, target/tmp, and target/e2e-media-runtime.

Research complete for #341. No product code changed. Artifact: `docs/research/documents-plugin.md` (335 lines), committed as `28303e55` and merged once with local `dev` at `1ad7b728`. Branch head: `4f3d7162aeef4d163472d817cd197f1fb9c06c3c`; pushed `job/documents-research` (remote was already current). The note covers the Paperless feature map, local-only OCR and PDF extraction, local classification, two readable Markdown metadata examples, rename behavior, fit with DESIGN §§4/17/32/33/40/45/48, sources, recommendations, and 12 numbered grill questions. Evidence and limits: - Synthetic image-only A4 scan at 300 dpi, English, one worker; Tesseract 5.5.3 through OCRmyPDF 17.11.0. First page: 30.97 s, 198 MiB max RSS; warm repeat: 10.79 s, 199 MiB; 20 pages: 2m45.33s, 186 MiB, 8.27 s/page. This is a single host run, not an accuracy or concurrency benchmark. - The configured Paperless host denied SSH with `Permission denied (publickey)`. No Paperless database/API query was made. Actual imported-library field names and populated counts remain unavailable; the note clearly separates standard Paperless fields from observed library facts. No document names, filenames, titles, OCR text, or metadata values were accessed or reported. - License note: Tesseract and its best trained data are Apache-2.0. OCRmyPDF is MPL-2.0; its upstream cautions that GPLv3/AGPLv3 integration is likely incompatible, so the note uses it only for the local CLI benchmark and does not recommend bundling it. PaddleOCR/docTR are candidates pending per-model/dependency license and offline checks. Apple Vision is proprietary and Apple-only. Decisions recommended where DESIGN is silent: make a readable Markdown Sidecar per document the canonical metadata source, retain a folder catalog Note as an alternative, keep OCR/classifier output in Derived data, save PDF/A only by explicit action, use explicit rules before considering a local classifier, and avoid automatic storage-path moves under §40. These are research recommendations for the grill, not product decisions. Validation (verbatim gate output): `cargo fmt --check` produced no stdout or stderr; exit status 0. `cargo clippy --all-targets -- -D warnings`: ```text Finished `dev` profile [unoptimized + debuginfo] target(s) in 149m 26s ``` Exit status 0. `cargo test` was capped at 45 minutes. It exited 124 while compiling and did not reach test execution. Last captured Cargo output: ```text Compiling mail-parser v0.11.9 Compiling elliptic-curve v0.13.8 Compiling synstructure v0.13.2 ``` `bun run check` and `bun run test` were not run because the job passed the approximately four-hour ceiling; remaining gates are open. One scoped adversarial attempt (`ADVERSARIAL_API_ONLY=1 tests/adversarial/run.sh`) also reached its 20-minute setup cap while compiling the isolated server. It was stopped before the HTTP probe; this is not a clean probe result. `cargo clean` ran. Removed `apps/web/build`, `target/tmp`, and `target/e2e-media-runtime`.
Author
Owner

Owner decision Q6b (2026-09-29): both. The default auto-apply bar is 95% calibrated precision per tag on the User's own history. A tag the User has confirmed 20 or more times may use a 90% bar. Each rejection still raises that tag's bar. Below the bar, a value is a suggestion in the inspector (Q7). #417 round 4 measures nearest-neighbour tagging over calternal's existing text embeddings, plus file-name, sender and first-page-header features, to raise coverage.

Owner decision Q6b (2026-09-29): **both.** The default auto-apply bar is 95% calibrated precision per tag on the User's own history. A tag the User has confirmed 20 or more times may use a 90% bar. Each rejection still raises that tag's bar. Below the bar, a value is a suggestion in the inspector (Q7). #417 round 4 measures nearest-neighbour tagging over calternal's existing text embeddings, plus file-name, sender and first-page-header features, to raise coverage.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#341
No description provided.