Documents: local recognition research on the owner's corpus (OCR, calibrated classifier, Laya, doc-vs-photo detector) #417
Open
opened 2026-09-29 09:09:23 +00:00 by kayg
·
42 comments
No Branch/Tag specified
dev
wip/merge-round-7c
wip/mchrome-1084
wip/mailghost2-1094
wip/mailghost-1094
wip/kbpreview2-1118
wip/kbpreview-1118
wip/kanban-1092
wip/importhang-1121
wip/hiderev-1153
wip/hide4-1153
wip/hide3-1153
wip/hide2-1153
wip/hide-1153
wip/editreg-1132
wip/editorrail3-1113
wip/editorrail2-1113
wip/editorrail-1113
wip/e2e-b2-1071
wip/e2e-b-1071
wip/draw4-1101
wip/draw3-1101
wip/draw2-1101
wip/draw-1101
wip/directory-1199
wip/delete-1119
wip/collabrev-1197
wip/collabloss-1197
wip/cards2-1083
wip/cards-1083
wip/canvas-visual
wip/canvasvis2-976
wip/calhdr-1112
wip/calcards-1115
wip/browserfix
wip/blocks-1125
wip/allday-1107
wip/agenda-decks
wip/agenda-1086
wip/adv7c-1105
wip/txentry-1198
wip/trayicons2-1095
wip/trayicons-1095
wip/sidebar3-1094
wip/rev2-webperf
wip/rev2-money-ident
wip/restyle-settings
wip/previewcard-1098
wip/palette2-1123
wip/palette-1093
job/segmented-1200
wip/onboard2-1141
job/tocrail-1191
wip/onboard-1141.aborted-early
wip/onboard-1141
job/restyle-settings
wip/notifloop-1194
wip/nlpchip-1127
wip/morph-1104
job/tagperf-1186
wip/merge-round-7c5
wip/merge-round-7c4
wip/merge-round-7c3
wip/merge-round-7c2
job/onboard-1141
job/notifloop-1194
wip/tagperf-1186
wip/segmented-1200
job/collabloss-1197
job/restyle-files
job/tagdnd-1187
job/merge30
job/perf-1124
job/cards-1179
wip/cards2-1179
wip/cards-1179
wip/tocrail-1191
job/hide-1153
wip/tagdnd-1187
wip/restyle-files
wip/perf-1124
wip/merge30j
job/restyle-notes
wip/restyle-notes
job/wizchoices-1140
job/adv-1202
wip/wizchoices-1140
wip/restyle-1190
job/moneyfmt-1180
job/txentry-1198
wip/moneyfmt2-1180
wip/moneyfmt-1180-r
wip/moneyfmt-1180
job/pillglass-1189
job/flags-1181
wip/flags-1181
job/restyle-1190
job/restyle-mailmoney
job/restyle-search
job/settingsreg-1195
job/wizard-1140
site/website
wip/wizardrev2-1140
wip/wizardrev-1140
wip/wizard5-1140
wip/wizard4-1140
wip/wizard3-1140
wip/wizard2-1140
wip/wizard-1140
wip/pillglass-1189
wip/settingsreg-1195
job/merge29
job/fu-1171
wip/merge29j
wip/fu-1171
job/fu-1166
job/directory-1199
job/txresearch-1188
wip/fu-1166
job/merge28
job/search-1066
wip/search-1066
wip/merge28j
job/gateslot-1182
job/bulkimport-1157
job/mailnet-1160
wip/mailnetrev-1160
wip/mailnet-1160
wip/bulkrev-1157
wip/bulkimport-1157
job/startup-1161
wip/startup-1161
job/merge27
job/linkcards-1151
wip/linkcards3-1151
wip/linkcards2-1151
wip/linkcards-1151
job/traydate-1144
wip/traydate3-1144
wip/traydate2-1144
wip/traydate-1144
job/draw-1101
wip/merge27j
job/blockpill-1152
wip/blockpill3-1152
wip/blockpill2-1152
wip/blockpill-1152
job/minihover-1149
wip/minihover2-1149
wip/minihover-1149
job/merge25
wip/merge25-r
wip/merge25b
wip/merge25
job/inspector-1129
job/tags-1110
wip/inspector3-1129
wip/inspector2-1129
wip/inspector-1129
wip/tagsrev-1110
wip/tags2-1110
wip/tags-1110
job/dates-1148
wip/datesrev-1148
wip/dates2-1148
wip/dates-1148
job/licence-1145
wip/licence2-1145
wip/licence-1145
job/selfhost-1156
job/merge23
wip/merge23
job/tagfilter-1109
wip/tagfilter2-1109
wip/tagfilter-1109
job/kbd-1134
wip/kbd2-1134
wip/kbd-1134
job/palfoot-1137
wip/selfhost-1156
wip/palfoot2-1137
wip/palfoot-1137
job/toggle-1158
wip/toggle-1158
job/kbpreview-1118
job/docratchet-1155
job/perflint-1133
job/devtests-1159
wip/docratchet-1155
wip/devtests-1159
job/segv-1136
wip/toast-1142
wip/segv-1136
job/toast-1142
job/blockreload-1147
wip/blockreload-1147
job/font-1150
wip/font-1150
job/importui-1120
job/minimonth-1149
wip/importui-1120
wip/minimonth-1149
job/depcheck-1146
wip/perflint-1133
wip/depcheck-1146
job/calcards-1115
job/blocks-1125
job/plus-1128
job/shift-1138
wip/plus2-1128
wip/plus-1128
wip/shift-1138
job/moneyfid-1130
job/editorrail-1113
wip/moneyrev-1130
wip/moneyfid-1130
job/noext-851
wip/noext-851
wip/noext3-851
wip/noext2-851
job/week-1135
wip/week-1135
job/editreg-1132
job/smoke-1122
wip/smoke-1122
job/docs-1143
job/palette2-1123
job/calhdr-1112
job/nlpchip-1127
job/mailghost-1094
job/reconnect-1131
wip/reconnect-1131
job/trayicons-1095
job/delete-1119
job/importhang-1121
job/cards-1083
job/palette-1093
job/mchrome-1084
job/e2e-a-1071
job/canvas-visual
job/previewcard-1098
job/allday-1107
wip/e2e-a2-1071
wip/e2e-a-1071
job/e2e-b-1071
job/adv7c-1105
job/kanban-1092
job/agenda-1086
job/merge-round-7c
job/morph-1104
wip/surfaces-p2
job/merge-round-9
wip/merge-round-9
job/7cfix-small
wip/7cfix-small
job/mailui-1078
job/merge-round-8
wip/merge-round-8
wip/mailui-1078
job/mailround-1038
job/applemail-accept
wip/settitle-1068
wip/mailround2-1038
wip/mailround-1038
wip/e2e-7b
job/crash-1069
wip/crash-1069
job/searchlost-1066
wip/searchlost-1066
job/7b-reconcile
job/flake-1065
wip/flake-1065
wip/merge-round-7b7
wip/merge-round-7b6
wip/merge-round-7b5
wip/merge-round-7b4
wip/7b-reconcile
job/appupdate-1059
job/nfd-1044
wip/appupdate-1059
job/e2e-7b
job/loop-1062
wip/loop-1062
job/pdfprev-1045
job/invtoggle-1053
wip/pdfprev-1045
wip/nfd-1044
wip/invtoggle-1053
job/7bfix-e2e
job/mailstress-b
wip/7bfix-e2e
wip/mailstress-b
job/7bfix-adv
wip/7bfix-adv
job/mailstress-a
job/stack-1054
wip/stack-1054
wip/mailstress-a
job/mailstress-1038
wip/mailstress-1038
job/upload500-1051
wip/upload500-1051
job/share-1034
wip/share-1034
job/syncerr-1037
job/7bfix-photos
wip/7bfix-photos
job/paste-1036
job/setside-1039
wip/setside-1039
wip/paste-1036
job/lease-1042
wip/syncerr-1037
wip/lease-1042
job/7bfix-data
job/passkeybind-1043
wip/apprevoke-1041
job/invite-1035
wip/invite-1035
job/merge-round-7b2
wip/merge-round-7b2
job/mailproxy-486
job/apprevoke-1041
job/rebuild-1033
job/pillborder-1029
wip/pillborder-1029
wip/mailproxy-486
wip/applemail-486
job/headless-998
wip/headless-998
job/groups-1028
wip/groups-1028
job/rebuildwarn-1016
wip/rebuildwarn-1016
job/startup-1011
wip/startup-1011
job/monthpill-1009
job/bgthumb-1025
job/sharetitle-1012
wip/monthpill-1009
wip/bgthumb-1025
wip/sharetitle-1012
job/canvas-cards-977
wip/canvas-cards-977
job/canvas-pencil-978
job/canvas-sketch-990
wip/canvas-sketch-990
wip/canvas-pencil-978
job/canvas-files-989
wip/canvas-files-989
job/canvas-collab-991
wip/canvas-collab-991
job/weekscroll-1018
wip/weekscroll-1018
wip/canvas-core-976
job/canvas-core-976
job/round-drag
wip/round-drag
job/round-settings
job/browserfix
wip/oapi-974
job/oapi-974
job/hist2-integrate
job/mailhtml-726
wip/mailhtml-726
wip/hist2-integrate
job/moneyfu-984
job/drag-1015
wip/drag-1015
job/rename-1017
wip/rename-1017
job/hist2-api
wip/hist2-api
job/oneacct-1014
wip/oneacct-1014
wip/moneyfu-984
job/hist2-bench
job/hist2-restore
wip/hist2-bench
job/hist2-write
job/hotfix-724
wip/hotfix-724
wip/hist2-write
wip/hist2-restore
job/hist2-store
job/hist2-ui
wip/hist2-ui
wip/hist2-store
job/searchstarve-965
job/shutdown-963
wip/shutdown-963
wip/pubedit-981
job/pubedit-981
job/analytics-973
wip/searchstarve-965
job/authflash-850
job/weeklane-969
job/pvtitle-1004
job/hist-975
wip/authflash-850
job/voicepill-617
wip/pvtitle-1004
job/headring-1003
wip/weeklane-969
wip/voicepill-617
wip/headring-1003
wip/analytics-973
job/agentscope-980
wip/thumbsandbox-988
job/thumbsandbox-988
wip/hist-975
job/links-856
wip/links-856
job/davetag-966
wip/davetag-966
job/filesstorm-1000
job/hoverpad-725
wip/filesstorm-1000
job/ffmpegblas-993
job/merge-round-7a
wip/hoverpad-725
wip/ffmpegblas-993
job/nowdot-1002
wip/verify-7a
job/noteid-857
wip/nowdot-1002
wip/noteid-857
wip/merge-round-7a
wip/agentscope-980
job/imapedge
job/a11yfix2
wip/imapedge-941
wip/imapedge
wip/a11yfix2
job/notetask-986
job/logheading
wip/logheading-998
job/textthumb-652
job/photolive-987
wip/photolive-987
job/davactive-983
job/savefix-985
job/tabicons-607
wip/davactive-983
wip/tabicons-607
wip/notetask-986
wip/savefix-985
job/dirid-627
job/buildspeed-1007
wip/dirid-627
job/agenda-decks
job/perfguards-impl
job/undo-a11y
wip/undo-a11y
job/mailperf
job/wal-824
wip/settings-50
job/settings-50
job/notesfilter-606
wip/notesfilter-606
job/surfaces-p2
wip/wal-824
job/maillayouts
wip/mailperf
wip/maillayouts
job/taskmeta-659
job/money-ident
wip/money-ident
wip/taskmeta-659
job/errstates
wip/perfguards-impl
job/headings-881
wip/headings-881
wip/errstates
job/voice-619
job/gaps-827
job/notesperf
wip/notesperf
wip/voice-619
job/hddsql-549
job/perf-stream-668
wip/perf-stream-668
wip/deeplinks-fix
job/deeplinks-fix
job/authfix
job/docsfix-rust
wip/docsfix-rust
job/webperf
job/docsfix-web
job/datafix2
job/webdav-lock-476
job/copyfix
wip/copyfix
wip/webperf
job/focus-658
wip/protofix
job/mediafix
job/protofix
wip/mediafix
job/agentfix
job/hhmm-724
wip/agentfix
job/undo-722
job/reuse
wip/webdav-lock-476
wip/reuse
job/scopefix
job/datafix
wip/hhmm-724
wip/undo-722
job/surfaces-p1
wip/hddsql-549
job/voicememos-618
wip/datafix2
wip/surfaces-p1
job/fix-940
wip/fix-940
job/blaze-surfaces
wip/datafix
wip/blaze-surfaces
job/taskday-655
job/linknav-639
wip/linknav-639
wip/gaps-827
job/isolation-707
job/audiophotos-720
wip/audiophotos-720
job/advfind-664
wip/voicememos-618
wip/taskday-655
wip/isolation-707
wip/advfind-664
wip/scopefix
wip/focus-658
job/testgaps
wip/testgaps
job/overscroll-718
wip/authfix
job/deps
wip/overscroll-718
job/rev2-agentfix
job/rev2-money-ident
job/rev2-mailperf
wip/deps
job/hardening-728
wip/hardening-728
job/searchgen-832
wip/searchgen-832
job/photopw-849
job/mailsql-825
wip/photopw-849
job/sharefix
wip/sharefix
job/rev2-mailhtml-726
job/rev2-perfguards
job/copyval-723
job/lightglass-r2
wip/lightglass-r2
wip/docsfix-web
job/copy-audit
job/macinterop-staging-r2
job/design-sync
job/rev2-taskmeta-659
job/rev2-webperf
job/docs-audit
job/rev2-advfind-664
job/rev2-mailproxy-486
job/states-audit
job/rev2-datafix
job/design-drift
job/test-gaps
job/rev2-voicememos-618
job/rev2-mediafix
job/rev2-deps
job/rev2-datafix2
job/licence-audit
job/issue-hygiene
job/rev2-protofix
job/rev2-voice-619
job/rev2-isolation-707
job/rev2-surfaces-p1
job/deeplink-audit2
job/rev2-audiophotos-720
wip/test-gaps
job/rev2-overscroll-718
job/rev2-undo-722
wip/states-audit
job/rev2-dropmd-719
job/rev2-linknav-639
job/merge-7b-plan
wip/merge-7b-plan
job/rev2-taskday-655
wip/mailsql-825
job/rev2-webdav-lock-476
job/rev2-browserfix
wip/design-drift
job/rev2-hddsql-549
wip/deeplink-audit2
job/rev2-scopefix
job/rev2-authfix
job/rev2-hardening-728
job/rev2-wal-824
job/rev2-sharefix
job/calsidebar-638
job/chrome-audit
job/ioperf
wip/ioperf
wip/chrome-audit
wip/calsidebar-638
job/dropmd-719
wip/dropmd-719
job/ocr-build
wip/ocr-build
job/blaze-settings
wip/copyval-723
job/toastring-721
wip/toastring-721
job/deployfix-732
wip/deployfix-732
wip/blaze-settings
job/money-import-recheck
job/rev-a11y
job/perf-arch-db
job/rev-7b-data
wip/textthumb-652
wip/perf-arch-db
job/sec-protocols
job/sidehdr-660
job/rev-7b-security
job/research-surfaces
job/rev-design-gaps
job/rev-mcp-api
wip/sidehdr-660
job/perf-arch-memory
wip/sec-protocols
job/perf-arch-bundle
job/snapedge-714
wip/rev-mcp-api
job/sec-supplychain
wip/research-surfaces
job/perf-arch-sync
job/rev-consistency
job/perf-arch-server
wip/perf-arch-server
wip/perf-arch-memory
job/perf-arch-io
job/perf-arch-client
job/sec-fs
job/sec-mcp-scopes
job/sec-sharing
job/perf-guards
job/sec-browser
job/sec-admin-deploy
job/sec-auth
wip/snapedge-714
job/bgpicker-717
wip/perf-arch-bundle
wip/money-import-recheck
job/advsetup-654
wip/bgpicker-717
wip/advsetup-654
job/burst-709
job/kbdcaps-710
job/app-pw-chooser
wip/burst-709
wip/app-pw-chooser
job/imaptest-625
wip/kbdcaps-710
job/fix-499
wip/fix-499
job/perf-mut-667
job/calimg-589
job/perf-snap-666
wip/calimg-589
wip/perf-snap-666
wip/perf-mut-667
job/perf-cache-665
wip/perf-cache-665
job/voicefiles-620
wip/voicefiles-620
job/admin-burst-705
wip/admin-burst-705
job/voicememos-review
wip/voicememos-review
wip/ryw-653
job/ryw-653
job/writeonopen-661
job/instant-663
wip/writeonopen-661
job/money-import-review
wip/money-import-review
wip/importjs-610
review/integrations-407-round6
wip/integrations-review
job/dragghost-612
wip/dragghost-612
job/integrations
wip/integrations
job/decider-656
job/merge-round-6
job/perf-rerun
wip/merge-round-6
job/integrations-review-round5
job/selalign-576
wip/selalign-576
job/mcp-events-491
job/files-631
job/cal-e2e-569
wip/cal-e2e-569
job/reload-423
wip/reload-423
wip/mcp-events-491
wip/files-631
job/notesbridge-644
wip/notesbridge-644
job/editor-series
job/calcard-series
wip/calcard-series
job/mcp-events-review-491
wip/mcp-events-review
wip/editor-series
job/quirks-546
job/integrations-recheck
job/tocrail-636
wip/tocrail-636
wip/quirks-546
wip/reminders-643
job/reminders-643
wip/davscale-573
job/davscale-573
job/integrations-review
wip/ocr-eval-584
job/ocr-eval-584
job/esc-537
wip/esc-537
job/toastname-586
wip/toastname-586
job/submenu-579
wip/submenu-579
job/tasks-mode
wip/tasks-mode
job/agentdocs-630
job/dupwrite-634
wip/agentdocs-630
wip/dupwrite-634
job/lightglass-588
wip/lightglass-588
job/tabswitch-549
job/ghosttask-623
wip/ghosttask-623
job/toaststack-616
job/weekstate-609
job/mailsync-613
wip/mailsync-613
wip/weekstate-609
job/maildup-626
wip/tabswitch-549
wip/maildup-626
wip/toaststack-616
job/motion-611
wip/motion-611
job/tlstest-601
wip/tlstest-601
job/perf-495
job/floating-sheet
wip/floating-sheet
job/remdup-585
wip/remdup-585
job/fix-502
wip/fix-502
job/attachplay-622
job/perf-batch
wip/perf-batch-563
wip/perf-495
hotfix/mail-sync-diag
job/mail-m3
wip/mail-m3
job/attach-poof-603
job/calhover-608
job/editorbar-604
job/mentions-605
job/merge-round-4
job/allday-514
wip/merge-round-4
wip/allday-514
job/merge-round-4a
wip/merge-round-4a
job/sharestack-580
job/fix-501
wip/sharestack-580
wip/fix-501
job/perf-batch-563
job/apw-cache-review
wip/apw-cache-review
job/probe-520
wip/probe-520
job/mac-393
wip/mac-393
job/header-571
job/flake-513
wip/flake-513
job/docs-thumb-547
wip/header-571
job/webcal-572
wip/webcal-572
wip/shortcuts-542
job/shortcuts-542
wip/docs-thumb-547
job/caldav-stress
wip/caldav-stress
wip/sweep-478
job/apw-cache-512
wip/apw-cache-512
job/money-empty-540
wip/restart-505
wip/money-empty-540
wip/fix-510
job/restart-505
job/fix-503
job/perf-496
wip/perf-496
job/fix-498
wip/fix-498
job/info-inspector-465
wip/info-inspector-465
job/fix-510
job/fix-507
wip/fix-507
wip/fix-503
job/fix-493
job/money-kinds
wip/money-kinds
job/hygiene-548
job/merge-round-3
wip/fix-493
job/drag-snap-536
wip/merge-round-3
wip/merge-round-0930
wip/drag-snap-536
job/align-538
wip/align-538
job/bg-flash
wip/bg-flash
job/money-import
job/search-count-544
wip/search-count-544
wip/money-import
job/settings-key-541
wip/settings-key-541
job/toast-539
job/preview-421
wip/preview-421
wip/toast-539
job/tasks-500-531
job/title-plain-526
wip/title-plain-526
wip/tasks-500-531
job/notes-bridge
wip/parity-484
job/parity-484
job/files-slow
job/crash-525
wip/notes-bridge
wip/files-slow
wip/crash-525
job/kbd-motion-527
wip/bg-422
job/analytics-504
wip/analytics-504
wip/kbd-motion-527
job/upload-pill-523
wip/upload-pill-523
wip/tray-order
job/tray-order
wip/overflow-mid
wip/merge-round-2
job/perf-494
wip/perf-494
wip/mcp-fast-492
wip/motion-477
wip/asr-ab-489
wip/theme-variants-506
wip/overflow-511
wip/week-header-508
wip/attach-427
job/dav-delete-471
job/iso-435
wip/iso-435
wip/files-sel-keys
wip/dav-delete-471
job/align-253
job/siwc-490
wip/siwc-490
job/money-kinds-review
wip/align-253
wip/money-kinds-review
job/small-bugs-3
wip/overlay-title-487
wip/multiget-500
wip/hidden-420
wip/webcal-ui
wip/webcal-431
job/perf-367
job/location
wip/small-bugs-3
wip/location
wip/perf-367
wip/admin-deny-483
job/tag-unicode-473
wip/tag-unicode-473
job/blur-436
wip/photos-470
wip/blur-436
wip/small-bugs-4
wip/hunt-20260930
wip/settings-hdr-482
wip/chips-416
job/dedup-375
wip/dedup-375
job/doc-stack
wip/doc-stack
job/tokens-literals
wip/tokens-literals
job/jobs-leftovers
wip/send-fast
wip/paste-467
wip/money-numbers
job/money-plugin
wip/money-plugin
job/break-dav
wip/merge-batch
wip/crossday-469
wip/mac-verify
wip/mail-m2
wip/break-dav
wip/money-review2
job/money-md
job/modes-424
wip/money-md
wip/jobs-leftovers
job/agenda-413
wip/agenda-413
wip/modes-424
job/recog-417
wip/recog-417
wip/bounce-425
wip/ab-384-luna
job/webdav-perf
wip/webdav-perf
job/toast-ring
wip/toast-ring
job/money-review
wip/money-review
wip/micro-motion
wip/settings-card
wip/minical
job/notes-imap-428
job/least-priv
wip/ui-small-2
wip/flaky-426
wip/drag-end-418
job/jank
wip/jank
wip/least-priv
wip/docs-site
job/agenda
job/sec-batch
wip/sec-batch
wip/per-user-index
job/area-calendars
wip/area-calendars
job/parity
wip/parity
job/documents-research
wip/documents-research
job/test-infra
job/reminders-sync
wip/small-bugs-2
wip/reminders-sync
wip/gestures
job/google-oauth
wip/tags-merge
wip/tags
job/e2e-theme
wip/e2e-theme
job/icon-align
wip/test-infra
wip/select-align
wip/editor-385
job/voice
wip/webdav
job/webdav
job/app-pw-ui
job/editor-integrity
wip/editor-integrity
wip/voice
wip/quota
wip/cal-followups
wip/icon-align
job/composer-scale
wip/composer-scale
job/jobs-page
wip/jobs-page
job/hig-type
wip/hig-type
wip/app-pw-ui
job/motion-spring
job/mcp
wip/motion-spring
wip/mcp
job/small-bugs
wip/push-hosts
job/profile-sign
wip/touch-369
wip/profile-sign
job/mobile-focus
wip/mobile-focus
wip/ui-polish-354
wip/small-bugs
wip/dup-task
job/toast-polish
job/app-pw-scopes
wip/toast-polish
wip/app-pw-scopes
wip/cli-agent
wip/selection-pills
job/preview-attach
wip/preview-attach
job/dav-proppatch
wip/dav-proppatch
wip/cal-switcher
job/atomic-race
wip/atomic-race
job/photos-shared
wip/photos-shared
wip/cal-grid
wip/note-rewrite
wip/search-rebuild
job/mail-m1
job/paperless-import
wip/paperless-import
wip/mail-m1
wip/hidden-activity
wip/search-d
wip/pricing-research
wip/cursors
wip/auto-scheme
job/single-pills
wip/single-pills
wip/xuser-matrix
wip/money-format
wip/app-pw-setup
wip/purge-dos
wip/vault-health
wip/caldav-apple
wip/xuser-audit
wip/e2e-green
wip/tabbar
wip/adv-harness
wip/maple-mono
job/search-fix
wip/search-fix
wip/search-perf-c
job/adv-harness
wip/sidebar-headers
job/glass
wip/temp-index
job/polish
wip/polish
wip/file-protocols
wip/money-research
wip/glass
wip/voice-models
wip/collab-redo
job/voice-research
wip/hunt-20260928
wip/notes-actions-research
wip/search-pad
wip/search-perf
wip/search-sticky
wip/editor-undo
wip/chrome-rules
wip/motion
wip/appearance-research
wip/appearance
wip/audit-bugs
wip/cal-glass
wip/block-actions
wip/authz-order
wip/event-stripes
wip/chrome-sidebar
wip/auth-flaky
wip/robust-2
wip/gate-fix
wip/menu-blur
wip/import-calternaljs
wip/tray-fix
job/import-calternaljs
wip/index-order
wip/audit-fixes
wip/search-chevrons
research/mail
wip/phone-chrome
wip/dedup-break
wip/csp
wip/ui-audit
wip/select-toast
wip/perf
wip/flat-layout
wip/fonts
wip/event-tint
wip/sync-converge
wip/data-split
wip/glass-audit
wip/robustness
wip/sync-chaos
wip/search-thumbs
wip/fuzz
wip/menu-icons
wip/search-pill
wip/sync-changing
wip/heading-links
wip/date-formats
wip/a11y
wip/break-editor
wip/e2e-fix
wip/settings-sections
wip/sync-root-guard
wip/search-palette
wip/share-edit
job/toasts
wip/toasts
wip/cont-analytics
wip/authz-review
wip/popovers
wip/overlay-glass
wip/change-feed
wip/editor-modes
wip/composer-align
wip/cont-agenda
wip/agenda-merge
job/agent-conventions
wip/agent-conventions
wip/backend-misc
job/route-audit
wip/route-audit
wip/ui-batch
wip/heif-hardening
wip/grid-resize
wip/ask-page
wip/webmcp
job/deeplink-audit
wip/deeplinks
wip/shortcuts
wip/cont-tz-days
main
No results found.
Labels
Clear labels
No items
No labels
Milestone
Clear milestone
No items
No milestone
Projects
Clear projects
No items
No project
Assignees
Clear assignees
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".
No due date set.
Dependencies
No dependencies set
Reference
kayg/calternal#417
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Part of #341 (Documents). Owner decisions, 2026-09-29:
Research and prototype (evidence, not opinions)
Corpus: the owner's real documents at
~/calternal-private/docs-corpus/(614 files, 192 MB, read-only). Privacy is a hard rule. This repo and its issue tracker are public. Never commit, post or log any document content, file name, name, amount or number from the corpus. Report aggregate metrics only. Keep all intermediate files (OCR text, features, models) under~/calternal-private/recog-work/(mode 700), never in the worktree. Labels: the folder path of each file. The paperless import turned tags into folders; derive the type and sender labels from it and describe the mapping in aggregate.Measure on the corpus with k-fold cross-validation, per field (type, sender, tags):
ocrs(pure Rust, MIT) and PaddleOCR via ONNX (Apache-2.0). For each: text quality on a sample you check by eye (report counts only), CPU seconds per page, and peak RSS on a 4-core machine. Also measure how many PDFs already have a text layer.Report: accuracy, coverage at 95% precision (the share of values the app could auto-apply), expected calibration error, training time, per-document inference time, and memory. Also simulate the cold start: the first 10, 25, 50 and 100 corrections.
netbird ssh --no-browser --user root 10.69.69.63: 4 vCPU, 7 GB). You may copy the corpus there under/root/private/(mode 700) and delete it when done.Deliverable
bench/recog/(code only, no corpus data), so it can be rerun. Do not build the product feature yet.Starting on
job/recog-417, based on99ee1fc9bbb0b3dd1f2d5b35e15529d1d0b3633c. I will keep all corpus reads and intermediate results private, commit only benchmark code, and report aggregate measurements. The #341 comments currently do not confirm an XMP Sidecar decision, so I will leave DESIGN §Sidecars unchanged unless that changes.Finding: the provided private corpus has 614 files (455 PDF, 140 DOCX, 15 XLSX, 3 XLS, 1 PNG), all directly in its root. The stated folder-path label source therefore contains zero type, sender, or tag classes. I am checking for a separate private label export while building the rerunnable benchmark. I will not fabricate supervised accuracy or 95% precision coverage from unlabeled data.
Finding: the
ocrscode is MIT/Apache-2.0, but its published.rtenweight files have no explicit license in the model repository. The upstream licensing question is still open (https://github.com/robertknight/ocrs-models/issues/34). Therefore I can benchmark the CLI, but I cannot recommend bundling those weights in an AGPL-3.0-only release until the owner grants compatible terms. Paperless-ngx's current classifier source uses CountVectorizer word 1–2 grams and separate scikit-learn MLPClassifier heads for type, correspondent and tags (https://github.com/paperless-ngx/paperless-ngx/blob/dev/src/documents/classifier.py).Finding: I ran the same deterministic 12-page PDF sample with Tesseract 5 on the 4-vCPU, 7-GB perf-test VM. All 12 pages produced nonempty text. OCR used 1.806 CPU seconds/page and 0.834 wall seconds/page; peak child-process RSS upper bound was 64,560 KiB. On the loaded build host the same sample measured 24.534 CPU seconds/page and 20.708 wall seconds/page, so VM figures are the useful runtime estimate. The PDFs were copied privately under
/root/private/recog-sampleand will be removed after measurements.Measurement correction: the initial 12-page OCR sample all had a first-page text layer. I have changed the benchmark to select equal numbers of text-layer and textless first pages. I will use the mixed sample for OCR recommendations; the earlier Tesseract numbers remain a text-layer-only reference.
Finding on the initial text-layer-only sample: PaddleOCR 3.7 with PP-OCRv5 mobile ONNX models on the 4-vCPU VM found text on all 12 pages, but used 20.361 CPU seconds/page, 7.324 wall seconds/page, and 1.802 GB observed RSS. Its model load took 0.586 seconds after weights were cached. Tesseract 5 used 1.806 CPU seconds/page, 0.834 wall seconds/page, and at most 64,560 KiB child-process RSS on those same pages. I am repeating both on a 6 text-layer / 6 textless-page sample before choosing an OCR backend.
Finding: the four-signal prototype combines OCR word-box coverage, a large quadrilateral, the existing MIT-licensed CLIP model's zero-shot image/text embeddings, and actual source provenance. On the available labeled images (1 corpus document image plus 4 CC0 Photos fixtures), with the corpus item's real document-import source hint, it classified 1/1 document and 4/4 negatives, with no errors. The rectangle signal fired on 0/5 images; the single document required OCR + CLIP + source. This sample is far too small to estimate field precision or screenshot routing, and it contains no screenshot. Without the source hint, the document stayed in Photos, as the conservative rule requires.
Finding on the mixed 6 text-layer / 6 textless-page sample, on the 4-vCPU VM: Tesseract 5 processed 12/12 pages with nonempty text. It used 25.512 CPU seconds/page and 11.790 wall seconds/page; peak child-process RSS upper bound was 64,392 KiB. Mean token recall against the six extractable PDF text layers was 0.887. Token recall is a proxy, not hand-checked OCR accuracy; the six textless pages have no ground-truth transcript. I am running the same sample through
ocrsand PaddleOCR ONNX.Finding on the same mixed 12-page VM sample:
ocrs0.13.1 found text on 12/12, used 4.166 CPU seconds/page and 1.852 wall seconds/page, with a 528,744 KiB peak child-process RSS upper bound. Mean token recall against the six text layers was 0.881, close to Tesseract's 0.887 on this proxy. The CLI loads its models per page, so a persistent library worker may improve throughput. The publishedocrsweights still lack explicit redistribution terms, which blocks bundling them in the AGPL release even if their benchmark is favorable.I found the new owner Q11 comment on #341. I recorded its confirmed Sidecar rules in DESIGN §4: standard names, hidden by default, the Hidden files setting, and carrying Sidecars with a moved or renamed parent. Q10 (which XMP fields Documents should use) remains OPEN, so I did not specify a Document XMP schema or build that behavior. Commit:
0d64b2ba.Finding on the mixed 12-page 4-vCPU VM sample: PaddleOCR 3.7 / PP-OCRv5 mobile ONNX processed 12/12 pages with nonempty text. It used 13.934 CPU seconds/page and 4.022 wall seconds/page, with 1.524 GB observed RSS and 0.753 seconds model load from cached weights. Mean token recall against the six existing PDF text layers was 0.968. The matched values are Tesseract: 25.512 CPU s/page, 11.790 wall s/page, ≤64,392 KiB RSS, recall 0.887;
ocrs: 4.166 CPU s/page, 1.852 wall s/page, ≤528,744 KiB RSS, recall 0.881. This proxy says nothing about transcription correctness on the six scans without text layers. Paddle offers the strongest measured text recovery with an Apache-2.0 model, but needs one bounded background worker because its memory is much larger.ocrsis fastest but its weight license remains unresolved.Finding from the private corpus using the rule extractor on first-page PDF text and DOCX text (DMY dates, comma decimals): 579/614 files were readable by this path; 35 were unsupported or unreadable. It emitted document-date candidates for 205, due-date candidates for 4, and amount candidates for 20. These are candidate counts, not precision or recall. A hand-labeled sample is still needed before any field can auto-apply. The conservative amount rule requires a currency marker; it misses unmarked amounts by design.
Perf-test VM detector round: the same 1 document image and 4 CC0 Photos fixtures gave 1 true positive, 4 true negatives, 0 errors, at 7.369 wall seconds/image after CLIP model initialization. The rectangle signal fired on none of these five images. This is a tiny functional check, not a precision estimate. The detector must run as bounded background work and should short-circuit cheap negative signals before CLIP in a product build. The benchmark keeps all four signals visible for comparison.
Merged
devonce at853fcc51before final gates. The one conflict was an unrelated Svelte selector inAccountRow.svelte. I retaineddev's:global(:nth-child(2))fix for the nested action count and made no other web edits. The recognition benchmark and confirmed Sidecar text remain intact.#417 report — local document recognition prototype
Branch
job/recog-417; final gate output will follow in a separate comment. The code is underbench/recog/. It reads the corpus in place, puts model and OCR work under a mode-0700 private directory, and prints aggregate results only. The copies on the perf-test VM have been removed.Corpus and labels
The private corpus has 614 files: 455 PDF, 140 DOCX, 15 XLSX, 3 XLS, and 1 PNG. All 614 files are at the corpus root. Thus the specified folder-path mapping has one constant class and yields no type, sender, or tag labels. There is no usable k-fold ground truth or correction order. I did not infer labels from names or invent accuracy, calibration, or 95%-precision coverage. A private Paperless export or hand-labeled manifest is required to complete those comparisons.
Of 455 PDFs, 396 have at least one extractable-text page. The PDF inventory reports 1,557 pages, of which 1,450 meet the text threshold. Sixteen PDFs failed inspection; the remaining 107 pages are a mix of textless and uninspected pages. The product should use existing text before OCR and queue OCR only for pages that need it.
OCR: 4-vCPU, 7-GB VM
The matched sample has 12 first pages: six with text layers and six without. Every engine returned nonempty text on all 12. Token recall compares unique OCR tokens with the six existing text layers. It is a proxy, not hand-checked transcription quality on scans.
ocrs0.13.1The
ocrsCLI loads models for each page, while the Paddle process stays loaded. These RSS measures use different methods. Tesseract and PaddleOCR use Apache-2.0 and Apache-2.0 Paddle detection weights and recognition weights. Theocrscode is MIT/Apache-2.0, but the published weight license remains unanswered. Do not bundle those weights yet.OCR recommendation: use the PDF text layer first. For missing text, provisionally use PaddleOCR mobile ONNX in one bounded background worker and unload it when idle; keep Tesseract as the low-memory path. The 1.5-GB active Paddle footprint is much larger than the 148-MB server idle RSS baseline, so it requires an Instance limit and a separate worker process. Before a product choice, hand-check scan transcriptions on a private labeled sample.
ocrsis the fastest measured option but cannot ship with the present weight-license evidence.Classification, confidence, and cold start
The benchmark includes a Paperless-style CountVectorizer 1–2 gram + MLP baseline and a small Rust TF-IDF softmax regression model with Platt scaling. The Rust tool fits vocabulary, model, calibrator, and confidence bar inside training folds. It reports type and sender separately, each tag as a private binary task, and first-10/25/50/100-correction simulations. Neither model has been measured on the owner's fields because the labels are absent.
The Laya authors report 0.362 zero-shot accuracy for the base English checkpoint and 0.766 for a tuned checkpoint on their typed-decision benchmark; these are not corpus results. On that same benchmark, they report ECE 0.175 for the base checkpoint, 0.213 for the tuned checkpoint, and 0.144 for hosted Jev. These are not corpus calibration values; Jev does not meet the local-only rule. They report roughly 4–5 hours on two T4 GPUs for 30,000 training questions. I did not run Laya on the corpus because the folder labels contain no target values. Cold-start accuracy and coverage at 10, 25, 50, and 100 corrections are also unmeasurable on this corpus.
Classifier recommendation: start with the local Rust linear model once a private label export exists. Calibrate each field for each User from their corrections. Use a separate held-out set to measure achieved precision and coverage; raise a tag's bar after a rejection. On a new User, apply only deterministic high-certainty values. Put other values in the inspector as quiet suggestions. Compare the Paperless baseline and Laya on the same folds before selecting a product model.
Extraction and image routing
The locale-aware rule prototype reads first-page PDF text and DOCX text. It read 579/614 files; 35 were unsupported or unreadable. There is no private hand-labeled date/amount sample. Laya is a typed decision model and would need rule-generated candidates before it could choose a date or amount.
Candidate counts are not precision or recall.
The detector uses OCR word-box coverage, a large quadrilateral, CLIP zero-shot labels, and source hints. With the real document-import hint, it routed the one corpus document image to Documents and four CC0 Photos fixtures to Photos: 1 true positive, 4 true negatives, 0 errors. It took 7.369 wall seconds/image on the VM after CLIP initialization. The rectangle signal fired on 0/5. With no source hint, the one document remained in Photos. There is no screenshot in this set, so the Screenshots exception is tested only with synthetic decision tests. Five images cannot establish precision.
Extraction recommendation: show rule candidates as suggestions until dates, due dates, and amounts have hand-labeled precision/recall. Detector recommendation: keep the conservative four-signal rule and its Photos fallback; run it as bounded background work. Keep ordinary screenshots in Photos → Screenshots, and show only high-certainty receipt, booking, or payment screenshots in Documents. Classification does not move files. Build a larger public negative set and private positive set before auto-routing.
License review
ocrscode and published modelsThe Paperless-style baseline uses scikit-learn but copies no Paperless source code. No hosted decision service or non-commercial weights were used.
Work estimate and decisions
The existing models avoid writing an OCR engine. Turning the prototype into a production Rust linear classifier with private correction storage, calibration, retraining, and per-User thresholds is roughly 2–3 engineer-weeks after labels exist. Integrating Paddle ONNX preprocessing and one bounded worker is roughly 2–4 weeks; a Tesseract subprocess path is roughly 3–5 days. A reliable multi-language OCR model from scratch would take months plus licensed training data. Hardening rule extraction and the image detector is roughly 1–3 weeks each once labeled validation sets exist. These are engineering estimates, not measured project schedules.
The design did not specify the benchmark sample. I chose six first pages with text and six without, selected by a stable path hash. The rule extractor uses the first PDF page and DOCX text. The classifier treats private JSONL order as correction order and requires at least 20 calibration predictions before a 95% bar can be set. The detector requires three signals for an unknown-source image and keeps uncertain images in Photos. None of these prototype thresholds is a product decision.
The new owner comment on #341 confirms Sidecars keep standard names, stay hidden by default, and move with their parent file. I recorded those rules in DESIGN §4. Document XMP fields remain OPEN.
Known gaps
ocrsmodel owner must state redistribution terms before those weights can ship.#417 final branch and gates
Head:
b1199ff5f0f11a323d045c29e349596bac785b93onjob/recog-417. Working tree clean. The branch contains the rerunnablebench/recog/prototype and the confirmed Sidecars decision in DESIGN §4. No product recognition route or UI was built, so no API adversarial round or screenshots apply.Full workspace Rust gates were launched once after the
devmerge. On the shared host, Clippy spent over two hours compiling dependencies;cargo test --workspacestayed on the Cargo build lock. I stopped both at the job timebox, exit 130. There were no diagnostics before interruption. These two workspace gates are incomplete, not passed. Their exact final log excerpts:Changed benchmark crate gates: Clippy exit 0; Rust test exit 0. Exact output:
cargo fmt --manifest-path bench/recog/Cargo.toml --check: exit 0, no output.Root
cargo fmt --check: exit 0, no output.bun run checkandbun run test: exit 0. Exact output:Python tests: exit 0. Exact output:
Cleanup, exact output:
apps/web/.svelte-kitwas deleted. The perf-VM corpus and model copies were deleted after measurement. The remaining gate work for the orchestrator is the full workspace Clippy and Cargo test on a quieter or larger build host; the branch is ready for that review.Resuming round 2 on
job/recog-417fromb1199ff5f0f11a323d045c29e349596bac785b93(currentdev:191b179baac3ef4f5bebfe07ce91c4b7a887ace2). I will evaluate pdf-inspector and current local layout-aware OCR on the 4-vCPU perf-test VM, keep corpus artifacts under the private directory, and report aggregate results only. The classifier remains unchanged while R3 labels are pending.Round-2 finding on the 4-vCPU, 7-GB perf-test VM: pdf-inspector 1.25.2 processed 439/455 PDFs; the same 16 failed the prior Poppler inspection. Against the prior 20-byte-per-page text-layer threshold, 1,557 pages yielded 1,426 correct text routes, 102 correct OCR routes, 24 text pages sent to OCR, and 5 textless pages missed. Full Markdown extraction across these PDFs took 14.443 CPU s and 14.528 wall s (0.0093 s/page), with a 105,880 KiB peak child-process RSS upper bound. Across 383 PDFs with an extractable text layer and nonempty Markdown, mean unique-token recall against Poppler text was 0.9909. This is a text-layer proxy, not a scan transcription score; the five missed pages require a fallback check before product routing. The rerunnable aggregate-only benchmark is commit
662de426underbench/recog/.Round-2 finding on the 4-vCPU VM: Docling.rs 1.74.1 (MIT code, permissively licensed ONNX assets) processed the matched 12 one-page PDFs with one warm worker in 83.08 CPU s and 38.28 wall s (6.92 CPU and 3.19 wall s/page), peak RSS 3,355,964 KiB. It returned nonempty text on 12/12. Against the six text layers, mean unique-token recall was 0.9232 and adjacent-token recall was 0.8410; it emitted Markdown table syntax on 5/12 pages and no HTML tables. These are proxies, not hand-labeled scan or table accuracy. The ongoing 107-textless-page run has already reached about 6 GB RSS, so the final recommendation must use its worst-case memory, not the mixed sample alone. Granite-Docling produced only an embedded image and zero text on one textless page (42.67 wall s, 2.81 GB peak RSS). SmolDocling did not finish a warm one-page conversion within 75 wall seconds, so it fails this VM's 60 s/page screen.
License correction: the MinerU 2.5 model card marks the weights AGPL-3.0, but the current upstream MinerU code LICENSE.md adds commercial-use conditions to Apache-2.0. That current code is not an acceptable dependency for this AGPL-3.0-only repository. I am recording this distinction rather than relying on the older code-license description.
Docling.rs corpus stress finding: a continuous one-worker run retained about 6.7 GB RSS by 70/107 textless pages, so I stopped it before memory exhaustion on the 7.7-GB VM. Restarting the process in 12-page batches reduced the first batch's peak RSS to 1.99 GB. The second batch peaked at 5.08 GB and converted 11/12 pages; it rejected one PDF page whose declared render size exceeded its built-in per-side safety cap. I am keeping that cap and counting the page as a fallback case. This shows why a bounded, recyclable worker and a lower-resolution OCR fallback are needed for extreme page geometry. No corpus path or content is in this comment.
Completed the 107 textless-page Docling.rs CPU run on the 4-vCPU VM with a fresh process every 12 pages. Nine batches took 1,270.02 CPU s and 617.75 wall s in total (11.87 CPU and 5.77 wall s per input page). Peak batch RSS was 5,078,068 KiB. The highest successful per-page wall time was 25.1 s; none exceeded 60 s. It produced Markdown for 106/107 pages and nonempty text for 95/106 outputs. Twenty-four outputs had Markdown table syntax and none had HTML table syntax. One page hit the parser's render-size safety cap; it needs a lower-resolution OCR fallback. Because the scans lack hand transcripts, nonempty text and table syntax are coverage checks, not accuracy scores. The earlier continuous run retained about 6.7 GB RSS, so process recycling is required on this VM.
Final pdf-inspector 1.25.2 routing measurement on the 4-vCPU VM: classification-only for 439 readable PDFs took 3.296 CPU s and 3.424 wall s. Full Markdown extraction took 14.535 CPU s and 14.143 wall s for 1,557 pages, peak child-process RSS ≤106,448 KiB. At PDF level it identified 383/396 text-bearing PDFs and called 0/43 textless PDFs text-bearing. At page level it kept 1,426/1,450 text-bearing pages on the text path, routed 102/107 textless pages to OCR, sent 24 text-bearing pages to OCR, and missed 5 textless pages. On 383 text-bearing PDFs with nonempty Markdown, mean unique-token recall against Poppler text was 0.9909 and adjacent-token recall was 0.9267. Markdown had table syntax in 285 PDFs, but table-cell correctness is unverified. The default Rust crate has MIT code and 86 normal dependencies; the checked dependency licenses contained no GPL-only, non-commercial, or unspecified license. The API accepts PDF bytes in memory, which can come from calternal-fs without constructing a user path.
PP-StructureV3 3.7.0 / PaddlePaddle CPU 3.3.1 finding on the 4-vCPU VM: its full layout/table pipeline loaded about 1.7 GB of official model assets but failed inside the Paddle static CPU predictor before returning one page. With oneDNN enabled it raised NotImplementedError. With oneDNN disabled it raised MemoryError at 5,958,472 KiB peak RSS. I then kept text/layout/table recognition, disabled formula/chart/seal models, and rendered the same page to a 1600-px maximum side; it still raised MemoryError at 4,952,736 KiB peak RSS after 35.27 wall seconds. No text or table output was produced, so there is no valid corpus recall or seconds-per-successful-page figure. This configuration is not a usable worker on this VM. The sampled official model cards checked so far declare Apache-2.0; I will record the complete model/license list in the report.
Round 2 result
Branch
job/recog-417has been merged with localdev; head SHA and final gate output follow below. The aggregate-only report isdocs/research/document-recognition-417.md. Benchmark code is inbench/recog/.MemoryErroreven after disabling formula, chart, and seal models and limiting rendered side to 1,600 px (4.95 GB peak RSS). No page completed.pdf-inspector is a local Rust crate with an in-memory API and 86 normal dependencies in the checked default closure. Its default build excludes OCR and PDFium. AnyDoc is MIT, runs office/PDF conversion locally, and depends on pdf-inspector. Its optional hosted OCR sends the PDF to Firecrawl Parse, so it is outside the local-only path.
Decision: use pdf-inspector for routing and the PDF text layer for text-bearing pages. Retry empty extracted text even when routing says text. Use one Docling.rs background worker for textless pages, unload it when idle, recycle after at most 12 pages, cap it at 6 GiB RSS and 60 s per page. Use the previously measured PP-OCRv5 mobile ONNX path at lower resolution for the one render-cap failure and 11 blank pages. The measured PDF work plus estimated fallback is about 12–15 minutes for the 439 readable PDFs on the VM. The 16 unreadable PDFs and office extraction are outside that estimate. No transcripts or table/reading-order gold labels exist, so quality remains a proxy. The type/sender/tag classifier is unchanged pending the private R3 Paperless-ngx label export.
Known gaps: no product recognition path was implemented; no valid CPU throughput/quality figures exist for models that failed or exceeded the one-page time limit. All staged corpus and model copies were removed from the VM. No per-document data is in this report.
Gates
Head SHA:
31a523176c2b91bbd1bac006732b36e1fb16a359.cargo fmt --check: exit 0, no output.cargo clippy --manifest-path bench/recog/Cargo.toml --all-targets -- -D warnings(exit 0):cargo test --manifest-path bench/recog/Cargo.toml(exit 0):bun run checkinapps/web(exit 0):bun run testinapps/web(exit 0; final test summary):python3 -m py_compile bench/recog/*.pyand the local privacy/extraction tests passed. Full local Python discovery could not importnumpyfor the earlier detector test because model dependencies are installed in the private VM environment. This does not affect the new routing/layout scripts. The latest host rule requires Rust gates per touched crate; the orchestrator runs full-workspace gates at merge time. No API was changed, so the adversarial API round did not apply.cargo cleanremoved 214 files, 105.1 MiB; web build output was removed. No push, deployment, or merge intodevwas made.Orchestrator decision on round 2: accept pdf-inspector routing plus text layer first. Do not default to Docling.rs: calternal.cloud has 7.7 GB of RAM shared with the server, and the owner's rule is minimal memory. Default OCR for textless pages: PaddleOCR mobile ONNX (1.5 GB peak, 4 s per page, best text recovery 0.97), in one worker loaded only while the queue has work and unloaded when idle, with a named RSS ceiling of 2 GB. Docling.rs becomes an Admin option (quality over memory, needs 6 GB or more), off by default. AnyDoc may extract office-file text if it runs fully local. Round 3 (queued) measures the tag classifier on the owner's paperless labels (604/614 matched, 36 tags with 10 or more documents).
Starting recog-417 round 3 on branch job/recog-417. Starting HEAD:
31a523176c; merge base with dev:09e238d527; dev HEAD:28ac39f917. I read the orchestrator decision: PaddleOCR mobile ONNX is default OCR; Docling.rs is an Admin option. I am measuring tag classification from the private Paperless export with aggregate-only reporting.Round 3 data preparation findings (aggregate only): exact matching joined 604/614 label records. The matched subset has 573 tagged records and 35 tags with 10 or more documents; the full export has 582 tagged records and 36 such tags, so nine tagged records did not match. Local PDF text and Office XML extraction produced 544 usable rows; 60 matched rows had no usable text. The round-2 OCR text artifacts were not present in the local private work directory, so this run uses embedded PDF text and Office XML only. The local Photos CLIP text encoder produced 512-dimensional features for all 544 rows using two context windows per row; inference took 115.506 seconds.
The prepared, date-ordered classifier input contains 544 text rows and 96 distinct ranked tags. Six tags from the matched label set have no usable text row after the local text-layer and Office XML pass. They cannot be scored by a text model in this run; the report will show the 96 measured ranks and call out this coverage limit. No raw tag values or document identifiers are included.
The Rust candidate's five-fold result is available. On the 96 tags with usable text, micro precision was 87.6%, micro recall was 33.7%, and micro ECE was 0.52%. One of 96 tags had a training-only 95% bar. Its held-out auto precision was 95.7% at 4.4% coverage of true tag instances. The date-ordered cold-start run found no tag with a qualifying bar through 200 corrections. These are aggregate corpus results; tag ranks and source labels are not included in this progress comment.
Method review found that public rank ids must be based on tag frequency in the full 614-record export, so unmatched documents cannot reorder T01–T36. I corrected the join, verified aggregate input counts (604 matched, 544 usable text rows, 102 tags total, 36 tags with at least 10 records in the full export), and am rerunning the models with the corrected ranks. No original tag values or document identifiers are included here.
Document tag classification, round 3 (#417)
This run measures tag classification on the private Paperless export. It uses
aggregate results only. It does not include document text, titles, file names,
or original tag values.
Input and text
The export has 614 records. Exact basename matching found 604 records in the
corpus. Of those, 573 have tags. The matched set has 102 distinct tags. Thirty
five tags occur on at least 10 matched records. Nine tagged records did not
match. Six tags have no usable text row in this run.
The private input contains 544 text rows and 96 distinct ranked tags. It sorts
rows by the export's creation date. The tool reads embedded PDF text and text
from DOCX and XLSX files. It omits 60 matched records that have no usable text.
The round-2 OCR text output was not present in the private work directory. This
run does not score OCR. It does not include the textless rows in the classifier.
The rank order follows tag frequency in the full 614-record export. This keeps
public ids stable when a record has no corpus match. Public per-tag tables show
only T01–T36. The other measured tags contribute to the all-tag micro and macro
results.
Models and scoring
shared hidden layer and one binary output per tag.
The implementation uses sparse stochastic gradient descent with L2 decay.
CLIP text model, followed by one logistic model per tag. The encoder uses at
most two evenly spaced context windows. Feature extraction took 115.506 s
across 544 rows.
Each fold fits its vocabulary, inverse document frequency, model, Platt
calibration, and confidence bar from training rows only. The Python models use
the same iterative multilabel folds. The Rust model stratifies each tag's
positive and negative rows separately. Each run uses five outer folds.
The confidence bar selects the widest calibration set with at least 20
predictions and at least 95% observed precision. Precision and recall use a
calibrated probability of 0.5. Coverage at 95% precision is the share of true
tag instances that the bar correctly auto-applies. ECE uses ten equal-width
probability bins. Micro ECE uses the combined tag-document probability bins.
Macro ECE is the mean of per-tag ECE values.
The cold-start run sorts usable text rows by creation date. It trains on the
first 10, 25, 50, 100, or 200 rows. The first 80% train the model. The last 20%
set its calibration bar. Later rows are held out. The run does not update the
model as it scores those later rows.
Results
The Paperless-style model has the highest micro recall (59.7%) and micro
precision of 74.7%. TF-IDF plus CLIP has 76.6% micro precision and 53.7% recall.
The Rust TF-IDF candidate has the highest micro precision (87.6%) and the
lowest recall (33.7%). Its macro precision and recall are low because many
infrequent tags receive no positive prediction. Of the 60 other measured tags,
18 receive no positive prediction from the Paperless-style model, 24 from Rust,
and 22 from TF-IDF plus CLIP. The CLIP feature does not improve micro recall
over the Paperless-style model in this corpus.
The Rust model is the only one with a held-out auto-apply precision at or above
95% (95.7%). Its coverage is 4.4% of true tag instances. The Paperless-style
model reaches 6.8% coverage, but its held-out auto-apply precision is 90.9%.
The CLIP model selects no held-out tags at its training-only bars. These results
show that a bar meeting 95% precision on calibration rows may not meet that
target on later documents. C95 below measures the share of true tags passed by
bars selected on training-only calibration rows; it does not promise 95%
precision on future documents.
No frequent tag has an eligible bar in any cold-start run through 200
corrections. The run cannot give a correction count at which most frequent
tags become eligible. The threshold is unknown and beyond the tested 200-row
limit if it is reached. Each N uses the newest fifth of those N rows for
calibration, so 10, 25, and 50 corrections do not provide the minimum 20
calibration candidates used by this benchmark.
Five-fold results
Precision and recall use calibrated probability ≥0.5. C95 is the share of true tags auto-applied by a bar selected on training-only calibration data; ECE uses ten probability bins.
Per-tag results
This table shows the 36 frequent tags only. Ranks follow frequency in the full export. Other measured tags contribute to the all-tag micro and macro results. The positive count is the number of usable text rows carrying the tag.
Cold-start curve
The model trains on the first N dated rows, calibrates on the newest fifth of those rows, and scores all later rows. ‘Barred frequent tags’ counts T01–T36 with a usable calibration bar.
First correction count with a bar for each frequent tag
A dash means no training-only calibration bar was available through 200 corrections.
At least half of frequent tags
Paperless MLP: fewer than 18 of the 36 frequent tags had a bar by 200 corrections.
Rust TF-IDF logistic: fewer than 18 of the 36 frequent tags had a bar by 200 corrections.
TF-IDF + CLIP: fewer than 18 of the 36 frequent tags had a bar by 200 corrections.
Decision for #341 Q6
Use a separate calibrated bar for each tag in the User's own correction
history. Auto-apply only when the bar has at least 95% precision on that
history. Require at least 20 calibration predictions before setting a bar.
Keep a rejection floor above the confidence of any rejected suggestion when
the model is updated.
The experiment does not establish that a small history will keep 95% precision
on future documents. The cold-start table shows how many frequent tags have a
training-only bar at each correction count.
Correspondents and document types
The export has no correspondent or document-type labels. It cannot support a
supervised precision or recall result for those fields. Document text may
contain sender strings or letterhead cues. These are possible zero-shot
signals, not measured evidence. Keep them as quiet suggestions until the User
confirms them. Do not auto-apply them from this corpus.
Scope and next step
The test uses PDF text and Office XML, not OCR output. The round-2 decision
sets PaddleOCR mobile ONNX as default OCR and Docling.rs as an Admin option.
Apply that OCR route in a later run to recover text from scans, then measure
whether tag coverage changes. No product recognition path is built here.
The required local API adversarial pass was attempted with
ADVERSARIAL_API_ONLY=1and the web build skipped. The runner completed dependency setup but did not start its server or send probes: itscargo buildprocess remained infutex_wait_queuefor 14:48 with no server binary and no compiler output, while the shared host had many concurrent Cargo jobs. I sent SIGINT to that build and the runner exited with status 1. This is a setup block, not an API finding; no adversarial API assertions ran. I did not retry the round.Completed
Built the rank-only multi-label tag benchmark for #417. It compares a Paperless-style shared MLP, Rust TF-IDF one-vs-rest logistic regression, and TF-IDF plus the local CLIP text feature. The five-fold tables and cold-start results are in
docs/research/document-recognition-417-round3.md; this comment thread also has the complete tables. Public output uses only T-rank ids.The input contained 614 export records, 604 corpus matches, and 544 usable text rows. There are 102 tags in the full export; 36 occur on at least 10 records. Six tags had no usable text row. The Rust candidate had 87.6% micro precision, 33.7% micro recall, and 4.4% true-tag coverage at training-calibrated 95% bars; held-out auto precision was 95.7%. No frequent tag had an eligible bar in the cold-start runs through 200 corrections, so the count for most tags remains unknown beyond the tested range.
Decisions
Known gaps
ADVERSARIAL_API_ONLY=1and the web build skipped. Setup installed dependencies but its servercargo buildstayed infutex_wait_queuefor 14:48 with no binary and no compiler output while many shared Cargo jobs were active. I stopped it with SIGINT; the runner exited 1 before starting the server or sending probes. No API assertions ran. This is a setup block, not an API finding. I did not retry.Files
bench/recog/: private label preparation, CLIP features, shared Python models, Rust candidate, rank-only report generator, tests, and run documentation.docs/research/document-recognition-417.md: link to the round-three report.docs/research/document-recognition-417-round3.md: method, aggregate and per-tag tables, cold-start curve, and recommendation.Merge and gates
Merged local
devonce before final gates. Head SHA:a0821e5d392c3c0038832d63c1364da1b14039fc.cargo fmt --manifest-path bench/recog/Cargo.toml --check(exit 0; output was empty).cargo clippy --manifest-path bench/recog/Cargo.toml --all-targets -- -D warnings(exit 0), output:cargo test --manifest-path bench/recog/Cargo.toml(exit 0), output:Python benchmark tests (exit 0), output:
py_compilesucceeded with no output.cargo clean --manifest-path bench/recog/Cargo.tomloutput:Round 4 resumed on
job/recog-417from base43c1377c8fd3b1b495841446989a64156f8ae1e4; current heada0821e5d392c3c0038832d63c1364da1b14039fc. I am extending the private, rank-only benchmark for the owner’s 95%/20-confirmation 90% rule, semantic nearest neighbours, extra signals, and OCR recovery. No corpus content or names will enter the public repository or issue.Round-4 OCR finding: Of the 60 matched rows without usable round-3 text, the source set is 56 PDFs, three legacy spreadsheets and one image. PDFium reports password protection for 16 PDFs, so their pages cannot be rendered or OCRed without keys. The first OCR attempt also found that this host has no
pdftoppm; the benchmark now renders with the already used PDFium library. OCR is running on the readable PDFs and image. The locked PDFs and legacy spreadsheets will remain an explicit measured gap; no private names or content are reported.Locked four-vCPU VM measurements (load averages recorded inside
flock /root/perf.lock): Search's pinned MiniLM ONNX encoder on 50 private documents took 1.621 s total, 32.4 ms/document, with 121,552 KB peak process RSS; starting load 0.04, 0.20, 0.50. A separate run of TF-IDF feature extraction, tag heads and cosine voting on 109 held-out private documents measured p50 36.035 ms, p95 42.106 ms, 187,688 KB peak RSS; starting load 0.02, 0.08, 0.36. These are separate processes, so their RSS values cannot simply be added as a measured combined peak. OCR cost is separate.OCR recovery completed. The matched corpus has 604 records. Round 3 had 544 usable text rows and omitted 60. PaddleOCR mobile ONNX recovered nonempty text for 41 of those 60, giving 585 usable rows, 555 with at least one Tag and 100 measured Tag ranks. Sixteen PDFs remain inaccessible because PDFium reports password protection; three legacy spreadsheets are unsupported. First-page source lines were recovered for 537 usable rows. The other 48 use a bounded text prefix as a header proxy. All text, names and vectors remain in the private work directory.
Final round-4 aggregate finding (585 usable rows, 100 Tag ranks): MiniLM k=3/5/10 gave 0% held-out C95; filename lexical voting gave 1.9% C95 at 95.5% held-out auto precision, and its 20-confirmation 90% slice gave 3.8% coverage at 91.3% precision. The calibrated stack gave 8.9% C95 and 10.7% mixed-policy coverage, but only 81.8% held-out auto precision, below both owner bars. In the chronological simulation, no model made all top 5, 10 or 20 Tags eligible through 500 corrections, and automatic precision was unstable where any bar existed. This supports quiet combined suggestions and cautious filename-only automation after sufficient User-specific evidence; the stack must not auto-apply from these calibration bars.
Updated locked VM inference for the final lexical-vote pipeline: on 117 held-out rows and 100 Tag heads, feature extraction plus voting measured p50 11.239 ms, p95 14.307 ms, peak RSS 177,008 KB; starting load 0.64, 0.53, 0.32. Search encoding was measured separately at 32.4 ms/document and 121,552 KB peak RSS on 50 rows. The prior 36.035 ms classifier result used the slower all-logistic candidate before lexical votes replaced its four signal heads; it is superseded.
The final combined resident-process VM measurement supersedes the sum of separate timings. Under
flock /root/perf.lockwith starting load 0.20, 0.42, 0.29, the pinned Search encoder plus TF-IDF heads, lexical and semantic votes, and stack inputs processed 117 held-out rows (100 Tag heads) at p50 48.686 ms and p95 59.395 ms per new record. Peak RSS was 252,156 KB; the process reported 258,576,384 bytes after inference. OCR remains a separate scan-only stage; round 2 measured PaddleOCR mobile ONNX at 4.02 wall s/page and 1.52 GB peak RSS on the same VM.Production-parity correction:
calternal-embeddisables the pinned tokenizer's fixed 128-token padding, uses two ONNX intra-op threads and disables thread spinning. I changed the prototype to those settings, regenerated all 585 Search vectors, and reran all five folds and cold-start points. The recommendation and coverage are unchanged to one decimal place (the small MiniLM precision/recall values move by at most 0.2 percentage points). Underflock /root/perf.lock, the corrected combined resident-process run on 117 records measured p50 75.214 ms, p95 85.331 ms, peak RSS 252,276 KB, with starting load 0.72, 0.50, 0.35. These numbers supersede the earlier 48.686/59.395 ms run, which used the tokenizer's fixed padding and four ONNX threads. The public round-4 report now contains the corrected values.Final workspace Clippy exposed one pre-existing Rust 1.98 lint in
calternal-fs: its one-byte Search readiness marker compared a byte array with[b'1']. Clippy with-D warningsrequests the equivalent*b"1". The marker writer already usesb"1"; I made that one-line, behavior-preserving change outside the benchmark-owned files so the required gate can continue. The initial full gate exited 101 at that lint. I am verifyingcalternal-fsbefore the follow-up workspace run.Gate update: the first full-workspace Clippy run found one
byte_char_sliceswarning incalternal-fson the Search readiness marker. Commitc8736bbf8changes the read-side comparison from[b'1']to*b"1", matching the writer.cargo clippy -p calternal-fs --all-targets -- -D warningsandcargo test -p calternal-fspass (42 tests). Full-workspace Clippy rerun is still compiling dependencies on the shared build host; no second diagnostic has appeared yet. The benchmark's 23 Python tests pass.