Search: sources — PDF text extraction, events, bookmarks, tasks, tags in the index #58

Open
opened 2026-09-24 15:52:59 +00:00 by kayg · 8 comments
Owner

Owner decision S1 (DESIGN §32): search everything that exists. For v0.0.1 add PDF text extraction to the Tantivy indexer (Rust: pdf-extract/lopdf or poppler pdftotext subprocess behind a trait; size and time caps; runs as a background job; derived data). Index as they land: tasks (#55), CalDAV events from the cache (#40), bookmarks and clipped text (#51), tags (#52), log entries (already). Office documents and photo OCR come with the photos plugin (separate issues). Permission filtering stays (home + shares).

Context for the owning job

  • Repo: kayg/calternal (~/Developer/calternal). Read CLAUDE.md, CONTEXT.md and docs/DESIGN.md (§15, §18 budgets, §31, §32) first. Prior art: calternal.js docs/search.md (read-only at /home/kayg/Developer/calternal.js) — its palette UX, > command mode, ranking weights, a11y combobox pattern and recents carry over.
  • Existing code: crates/calternal-search (Tantivy index, watcher + reconcile, providers via calternal-plugin fan-out), the ⌘K registry in apps/web.
  • Owner rules: search must be ultra fast (⌘K results < 50 ms p95 at 100k items; first keystroke to first results < 16 ms for client providers); file over app (indexes are derived and rebuildable); performance first but never at the cost of finesse; calternal.js design system (Claude reviews screenshots; floating window over a dimmed + blurred background, spring motion, reduced-motion respected); never ship sample data; atomic commits (commit every 30–45 min); adversarial testing after API work.
  • Comment on this issue when you start, on findings, when blocked, and when finished. Never close it.
Owner decision S1 (DESIGN §32): search everything that exists. For v0.0.1 add **PDF text extraction** to the Tantivy indexer (Rust: `pdf-extract`/`lopdf` or poppler `pdftotext` subprocess behind a trait; size and time caps; runs as a background job; derived data). Index as they land: tasks (#55), CalDAV events from the cache (#40), bookmarks and clipped text (#51), tags (#52), log entries (already). Office documents and photo OCR come with the photos plugin (separate issues). Permission filtering stays (home + shares). ## Context for the owning job - Repo: kayg/calternal (~/Developer/calternal). Read CLAUDE.md, CONTEXT.md and docs/DESIGN.md (§15, §18 budgets, §31, §32) first. Prior art: calternal.js `docs/search.md` (read-only at /home/kayg/Developer/calternal.js) — its palette UX, `>` command mode, ranking weights, a11y combobox pattern and recents carry over. - Existing code: `crates/calternal-search` (Tantivy index, watcher + reconcile, providers via calternal-plugin fan-out), the ⌘K registry in apps/web. - Owner rules: search must be **ultra fast** (⌘K results < 50 ms p95 at 100k items; first keystroke to first results < 16 ms for client providers); file over app (indexes are derived and rebuildable); performance first but never at the cost of finesse; calternal.js design system (Claude reviews screenshots; floating window over a dimmed + blurred background, spring motion, reduced-motion respected); never ship sample data; atomic commits (commit every 30–45 min); adversarial testing after API work. - Comment on this issue when you start, on findings, when blocked, and when finished. Never close it.
Author
Owner

Starting the search backend job for issues #58, #61, and #62. Branch: job/search-backend; base SHA: 55a2752a4c4c7a087fce55868871a8576fea31a7. I am reading the current index, plugin interfaces, and API contract before implementation.

Starting the search backend job for issues #58, #61, and #62. Branch: `job/search-backend`; base SHA: `55a2752a4c4c7a087fce55868871a8576fea31a7`. I am reading the current index, plugin interfaces, and API contract before implementation.
Author
Owner

Tantivy’s open_or_create rejects an existing index when the schema changes. The current index schema is stored in .system/index/tantivy, so adding fields without handling this would prevent existing instances from starting. I added a handle-relative reset for that fixed derived-data directory and a test that opens an old schema and confirms the new index is marked for a source rebuild. This worktree has files and Notes providers only; it has no Calendar/event, bookmark, or task plugin. Markdown task files and Daily note Log entries are indexed now. Calendar and bookmark results can be indexed when those providers land, as the issue specifies.

Tantivy’s `open_or_create` rejects an existing index when the schema changes. The current index schema is stored in `.system/index/tantivy`, so adding fields without handling this would prevent existing instances from starting. I added a handle-relative reset for that fixed derived-data directory and a test that opens an old schema and confirms the new index is marked for a source rebuild. This worktree has files and Notes providers only; it has no Calendar/event, bookmark, or task plugin. Markdown task files and Daily note Log entries are indexed now. Calendar and bookmark results can be indexed when those providers land, as the issue specifies.
Author
Owner

Finished on branch job/search-backend at head 9e3cb6ed6b677a0af990955821f4aa96b24b3866.

The backend now extracts PDF text in the background with input, page, output, and time caps; indexes task Markdown, individual daily-log entries, and searchable Markdown metadata; and applies the caller's allowed roots. Shared operator parsing and search-open tracking are also available through the API. The PDF source remains the original file; extracted text is derived index data.

At 100k indexed documents, 200 search queries measured 100k search query latency: p50=6369 us p95=13817 us (200 queries), under the 50 ms p95 target. The adversarial run ended with ==== FINDINGS 0.

Gates:

  • cargo fmt --check: exit 0, no output.
  • cargo clippy --all-targets -- -D warnings: Finished dev profile [unoptimized + debuginfo] target(s) in 6.29s.
  • cargo test --workspace: passed; final output included test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s for the remaining doc-test targets.
  • bash packages/api-client/check-generated.sh: passed; OpenAPI generated with openapi-typescript 7.13.0.
  • bash tests/adversarial/run.sh: passed; ==== FINDINGS 0.

Events and bookmarks have no providers in this worktree, so they remain to be indexed when those providers land. Relative date filters use UTC because the API has no user time-zone setting.

Finished on branch `job/search-backend` at head `9e3cb6ed6b677a0af990955821f4aa96b24b3866`. The backend now extracts PDF text in the background with input, page, output, and time caps; indexes task Markdown, individual daily-log entries, and searchable Markdown metadata; and applies the caller's allowed roots. Shared operator parsing and search-open tracking are also available through the API. The PDF source remains the original file; extracted text is derived index data. At 100k indexed documents, 200 search queries measured `100k search query latency: p50=6369 us p95=13817 us (200 queries)`, under the 50 ms p95 target. The adversarial run ended with `==== FINDINGS 0`. Gates: - `cargo fmt --check`: exit 0, no output. - `cargo clippy --all-targets -- -D warnings`: `Finished `dev` profile [unoptimized + debuginfo] target(s) in 6.29s`. - `cargo test --workspace`: passed; final output included `test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s` for the remaining doc-test targets. - `bash packages/api-client/check-generated.sh`: passed; OpenAPI generated with `openapi-typescript 7.13.0`. - `bash tests/adversarial/run.sh`: passed; `==== FINDINGS 0`. Events and bookmarks have no providers in this worktree, so they remain to be indexed when those providers land. Relative date filters use UTC because the API has no user time-zone setting.
Author
Owner

PDF extraction evidence from #303's adversarial round (2026-09-28): a one-page PDF containing search extraction sentinel uploaded successfully, but the Search API did not return it after 80 polls (20 seconds). The media upload probe also reported that its thumbnail worker did not finish within 10 seconds. This was a high-load run, so the missing result may be worker backlog; the finding remains unverified on a quiet host.

PDF extraction evidence from #303's adversarial round (2026-09-28): a one-page PDF containing `search extraction sentinel` uploaded successfully, but the Search API did not return it after 80 polls (20 seconds). The media upload probe also reported that its thumbnail worker did not finish within 10 seconds. This was a high-load run, so the missing result may be worker backlog; the finding remains unverified on a quiet host.
Author
Owner

During the 2026-09-28 real-server adversarial round, the PDF text-search probe uploaded its PDF fixture but Search did not return the fixture's embedded text. This ran during the concurrent Search rebuild failure reported on #345, which stopped the background indexer, so the probe does not isolate PDF extraction from that failure. Recording the observation here; no Search code changed in this job.

During the 2026-09-28 real-server adversarial round, the PDF text-search probe uploaded its PDF fixture but Search did not return the fixture's embedded text. This ran during the concurrent Search rebuild failure reported on #345, which stopped the background indexer, so the probe does not isolate PDF extraction from that failure. Recording the observation here; no Search code changed in this job.
Author
Owner

During the single real-server adversarial round for #291, the PDF text search probe reported: uploaded PDF text was not searchable. The round did not isolate extraction, indexing, or query behavior. This is evidence from one run and was not repeated.

During the single real-server adversarial round for #291, the PDF text search probe reported: `uploaded PDF text was not searchable`. The round did not isolate extraction, indexing, or query behavior. This is evidence from one run and was not repeated.
Author
Owner

The one-time adversarial run on job/touch-369 uploaded the PDF fixture, then the PDF text-search check could not find its expected text. This ran alongside severe shared-host load; the finding is outside #369's UI scope and was not changed here.

The one-time adversarial run on `job/touch-369` uploaded the PDF fixture, then the PDF text-search check could not find its expected text. This ran alongside severe shared-host load; the finding is outside #369's UI scope and was not changed here.
Author
Owner

Hygiene review: several real-server runs uploaded the PDF fixture but Search did not return its embedded text. Extraction versus indexing remains unisolated, so #58 stays open.

Hygiene review: several real-server runs uploaded the PDF fixture but Search did not return its embedded text. Extraction versus indexing remains unisolated, so #58 stays open.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#58
No description provided.