Notes rebuild: staging emits empty-Task and Tantivy segment warnings #1016

Open
opened 2026-10-03 15:25:24 +00:00 by kayg · 7 comments
Owner

Observed during #724 staging upgrade from staging-c82b9aca1 (Notes migration 26) to staging-564114585 (migration 27).

The automatic projection rebuild completed 123/123 Daily notes in 9.205 s. The startup window has 276 calternal_plugin_notes::tasks_store warnings with the exact message Skipped a Task projection with an empty title (#623), and two tantivy::indexer::segment_manager warnings at 2026-10-03T15:16:14.241919Z and 15:16:14.241961Z (segment status diagnostics and a missing segment). There are no ERROR records or panics. Raw logs stay private because the Instance has test Users.

Checks found no data loss: all 120 generated Daily notes have unchanged WebDAV SHA-256 hashes, all 960 expected Calendar Log entries have correct dates/times/Tags, all 120 generated Tasks remain indexed, and two HHMM Search targets return without a timeout. The empty-title warning comes from the existing #623 reference-only Task filter at crates/plugins/notes/src/tasks_store.rs:213. The Tantivy source is indexer/segment_manager.rs in 0.26.2.

Investigate why a successful startup/rebuild emits these warnings. Keep expected reference-only filtering quiet if it is normal. Confirm that concurrent segment merges recover without lost Search rows before changing the Tantivy diagnostics. This is a follow-up from staging verification; no behavior or test expectation was changed to hide it.

Observed during #724 staging upgrade from `staging-c82b9aca1` (Notes migration 26) to `staging-564114585` (migration 27). The automatic projection rebuild completed 123/123 Daily notes in 9.205 s. The startup window has 276 `calternal_plugin_notes::tasks_store` warnings with the exact message `Skipped a Task projection with an empty title (#623)`, and two `tantivy::indexer::segment_manager` warnings at 2026-10-03T15:16:14.241919Z and 15:16:14.241961Z (segment status diagnostics and a missing segment). There are no ERROR records or panics. Raw logs stay private because the Instance has test Users. Checks found no data loss: all 120 generated Daily notes have unchanged WebDAV SHA-256 hashes, all 960 expected Calendar Log entries have correct dates/times/Tags, all 120 generated Tasks remain indexed, and two HHMM Search targets return without a timeout. The empty-title warning comes from the existing #623 reference-only Task filter at `crates/plugins/notes/src/tasks_store.rs:213`. The Tantivy source is `indexer/segment_manager.rs` in 0.26.2. Investigate why a successful startup/rebuild emits these warnings. Keep expected reference-only filtering quiet if it is normal. Confirm that concurrent segment merges recover without lost Search rows before changing the Tantivy diagnostics. This is a follow-up from staging verification; no behavior or test expectation was changed to hide it.
Author
Owner

Starting investigation on branch job/rebuildwarn-1016, based at f2f8491ff5 (origin/dev).

Starting investigation on branch job/rebuildwarn-1016, based at f2f8491ff5c76ab28f140c964542c97362e6b119 (origin/dev).
Author
Owner

Finding: crates/plugins/notes/src/tasks_store.rs::index_source warns on every extracted row with a blank title before distinguishing reference-only checkboxes. The existing #623 contract says those retained Log links are attachments after their target is trashed, so this expected filter runs during rebuild and accounts for the Notes warning burst.

Finding: crates/plugins/notes/src/tasks_store.rs::index_source warns on every extracted row with a blank title before distinguishing reference-only checkboxes. The existing #623 contract says those retained Log links are attachments after their target is trashed, so this expected filter runs during rebuild and accounts for the Notes warning burst.
Author
Owner

Finding: Tantivy 0.26.2 SegmentManager::end_merge returns a recoverable error if merge inputs no longer exist after an update; the merge output is discarded. Regression evidence: concurrent_commits_keep_search_rows_through_segment_merges passed 64 commits / 512 rows, verified fewer than 64 final segments and all 512 rows after wait_merging_threads().

Finding: Tantivy 0.26.2 SegmentManager::end_merge returns a recoverable error if merge inputs no longer exist after an update; the merge output is discarded. Regression evidence: concurrent_commits_keep_search_rows_through_segment_merges passed 64 commits / 512 rows, verified fewer than 64 final segments and all 512 rows after wait_merging_threads().
Author
Owner

Finding: the new warning-capture test reproduced the original path with and captured . The extracted row has an empty title and a reference target, and indexing creates no Task row.

Finding: the new warning-capture test reproduced the original path with and captured . The extracted row has an empty title and a reference target, and indexing creates no Task row.
Author
Owner

Finding: the new warning-capture test reproduced the original path with - [ ] arrow Missing task and captured WARN calternal_plugin_notes::tasks_store: Skipped a Task projection with an empty title (#623). The extracted row has an empty title and a reference target, and indexing creates no Task row.

Finding: the new warning-capture test reproduced the original path with - [ ] arrow [Missing task](missing-task.md) and captured WARN calternal_plugin_notes::tasks_store: Skipped a Task projection with an empty title (#623). The extracted row has an empty title and a reference target, and indexing creates no Task row.
Author
Owner

Finding: deploy/cloud/calternal-cloud.env.example sets RUST_LOG=info, so a fallback-only filter would not silence the segment-manager warnings in that deployment shape. The server will apply the targeted threshold alongside broad RUST_LOG levels and preserve an explicit tantivy::indexer::segment_manager directive for diagnosis.

Finding: deploy/cloud/calternal-cloud.env.example sets RUST_LOG=info, so a fallback-only filter would not silence the segment-manager warnings in that deployment shape. The server will apply the targeted threshold alongside broad RUST_LOG levels and preserve an explicit tantivy::indexer::segment_manager directive for diagnosis.
Author
Owner

READY FOR MERGE: yes

Built

  • Titleless inline checkboxes with a reference target remain filtered from the Task projection. Their expected per-item diagnostic is now debug level; other empty-title projection rows still warn.
  • Added a regression that commits 512 Search rows in 64 batches, waits for Tantivy merge workers, confirms compaction occurred, and verifies all rows remain searchable.
  • Regression tests: reference_only_task_projection_filter_is_quiet, concurrent_commits_keep_search_rows_through_segment_merges, and tantivy_merge_diagnostics_need_an_exact_target_directive.
  • The server applies an error threshold to Tantivy's segment-manager diagnostics even when a broad RUST_LOG value such as info is set. An exact RUST_LOG directive for tantivy::indexer::segment_manager opts back into those warnings. Errors remain visible.
  • Kept the locked Tantivy version at 0.26.2; cargo search reports 0.26.2 as the available version.

Files

  • crates/plugins/notes/src/tasks_store.rs
  • crates/plugins/notes/src/lib.rs
  • crates/plugins/notes/Cargo.toml
  • crates/calternal-search/src/index.rs
  • crates/calternal-server/src/main.rs
  • Cargo.lock

Commits and head

Gates

cargo fmt --check
(no output; exit code 0)

cargo clippy -p calternal-search --all-targets -- -D warnings
Finished `dev` profile [unoptimized + debuginfo] target(s) in 0.94s

cargo test -p calternal-search -- --test-threads=4
test result: 41 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 13.47s
test result: 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.92s
test result: 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.06s
test result: 24 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 205.53s
test result: 5 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.02s
test result: 2 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.02s
test result: 1 passed; 0 failed; 2 ignored; 0 measured; 0 filtered out; finished in 4.24s
test result: 4 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s
test result: 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s

cargo clippy -p calternal-plugin-notes --all-targets -- -D warnings
Finished `dev` profile [unoptimized + debuginfo] target(s) in 1m 13s

cargo test -p calternal-plugin-notes -- --test-threads=4
test result: 187 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 106.66s
test result: 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.34s
test result: 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s

cargo clippy -p calternal-server --all-targets -- -D warnings
Finished `dev` profile [unoptimized + debuginfo] target(s) in 18.99s

cargo test -p calternal-server -- --test-threads=4
test result: 163 passed; 0 failed; 5 ignored; 0 measured; 0 filtered out; finished in 14.67s

bun run build
✓ built in 21.67s
  Wrote site to "build"

Cleanup: cargo clean reported Removed 18051 files, 9.5GiB total; generated web output and installed node_modules were removed.

Known gaps

  • No staging restart was run. The local regression exercised concurrent commits and merges; no API or UI behavior changed, so no adversarial API round or screenshots were needed.
  • The required web build succeeded and emitted MODULE_LEVEL_DIRECTIVE warnings in vendored analytics chart files. No web source was changed.

Decisions not in DESIGN.md

  • Treat a titleless Inline row with reference_target as expected reference-only filtering and log it below the default level. Keep warnings for other titleless rows.
  • Treat Tantivy's segment-manager missing-input diagnostics as recoverable by default after the concurrent-merge row-retention regression passes. Keep exact-target opt-in for diagnosis.
READY FOR MERGE: yes ## Built - Titleless inline checkboxes with a reference target remain filtered from the Task projection. Their expected per-item diagnostic is now debug level; other empty-title projection rows still warn. - Added a regression that commits 512 Search rows in 64 batches, waits for Tantivy merge workers, confirms compaction occurred, and verifies all rows remain searchable. - Regression tests: `reference_only_task_projection_filter_is_quiet`, `concurrent_commits_keep_search_rows_through_segment_merges`, and `tantivy_merge_diagnostics_need_an_exact_target_directive`. - The server applies an error threshold to Tantivy's segment-manager diagnostics even when a broad RUST_LOG value such as info is set. An exact RUST_LOG directive for tantivy::indexer::segment_manager opts back into those warnings. Errors remain visible. - Kept the locked Tantivy version at 0.26.2; cargo search reports 0.26.2 as the available version. ## Files - crates/plugins/notes/src/tasks_store.rs - crates/plugins/notes/src/lib.rs - crates/plugins/notes/Cargo.toml - crates/calternal-search/src/index.rs - crates/calternal-server/src/main.rs - Cargo.lock ## Commits and head - edb1a29bb test search row recovery during concurrent merges (#1016) - 32e5ad432 quiet expected reference Task projection warnings (#1016) - 2b1486ecc quiet recoverable Tantivy merge warnings by default (#1016) - Head: 2b1486ecc36f8ec6374707fed30a16419daefc24 ## Gates ```text cargo fmt --check (no output; exit code 0) cargo clippy -p calternal-search --all-targets -- -D warnings Finished `dev` profile [unoptimized + debuginfo] target(s) in 0.94s cargo test -p calternal-search -- --test-threads=4 test result: 41 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 13.47s test result: 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.92s test result: 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.06s test result: 24 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 205.53s test result: 5 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.02s test result: 2 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.02s test result: 1 passed; 0 failed; 2 ignored; 0 measured; 0 filtered out; finished in 4.24s test result: 4 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s test result: 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s cargo clippy -p calternal-plugin-notes --all-targets -- -D warnings Finished `dev` profile [unoptimized + debuginfo] target(s) in 1m 13s cargo test -p calternal-plugin-notes -- --test-threads=4 test result: 187 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 106.66s test result: 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.34s test result: 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s cargo clippy -p calternal-server --all-targets -- -D warnings Finished `dev` profile [unoptimized + debuginfo] target(s) in 18.99s cargo test -p calternal-server -- --test-threads=4 test result: 163 passed; 0 failed; 5 ignored; 0 measured; 0 filtered out; finished in 14.67s bun run build ✓ built in 21.67s Wrote site to "build" ``` Cleanup: cargo clean reported `Removed 18051 files, 9.5GiB total`; generated web output and installed node_modules were removed. ## Known gaps - No staging restart was run. The local regression exercised concurrent commits and merges; no API or UI behavior changed, so no adversarial API round or screenshots were needed. - The required web build succeeded and emitted MODULE_LEVEL_DIRECTIVE warnings in vendored analytics chart files. No web source was changed. ## Decisions not in DESIGN.md - Treat a titleless Inline row with reference_target as expected reference-only filtering and log it below the default level. Keep warnings for other titleless rows. - Treat Tantivy's segment-manager missing-input diagnostics as recoverable by default after the concurrent-merge row-retention regression passes. Keep exact-target opt-in for diagnosis.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#1016
No description provided.