Search: canonically equivalent NFD query misses NFC file names #1044

Open
opened 2026-10-04 08:47:45 +00:00 by kayg · 9 comments
Owner

Found during #867 / job/7bfix-photos. This is a non-blocking Search oddity, so this job leaves it for a separate fix.

A real local server receives two Tus uploads: café-nfc.txt with body unicodenfcsentinel, and cafe followed by U+0301-nfd.txt with body unicodenfdsentinel. Files uses NFC names. After both ASCII body markers are searchable, GET /api/v1/search?q=caf%C3%A9 returns both Files. GET /api/v1/search?q=cafe%CC%81 returns HTTP 200 with no results. The query spellings are canonically equivalent.

Source evidence: crates/calternal-search/src/query.rs:222 passes query text directly to SimpleTokenizer. crates/calternal-search/src/index.rs:603 does the same for filename tokens. Neither normalizes canonical Unicode sequences first. The same query_words implementation is present at production dev 6074f71d1. This job will add exact dev-binary evidence on #867 when its comparison completes.

Expected: canonically equivalent NFC and NFD terms find the same items. Keep the stable Files path and source bytes unchanged; normalize derived Search terms with the existing Unicode rule (DESIGN §§32, 40). Cover names, titles, headings and body text as well as query terms. Reuse the shared normalization primitive.

Test: index one NFC and one NFD text, search with both query spellings, compare returned stable item identities and verify the original source bytes stay unchanged. Run the focused live check on dev and the fix. The missing ASCII sentinel in round-3 is a separate timing finding: a local run took about eight seconds to find it, while search_chaos.py waits five seconds.

Found during #867 / job/7bfix-photos. This is a non-blocking Search oddity, so this job leaves it for a separate fix. A real local server receives two Tus uploads: café-nfc.txt with body unicodenfcsentinel, and cafe followed by U+0301-nfd.txt with body unicodenfdsentinel. Files uses NFC names. After both ASCII body markers are searchable, GET /api/v1/search?q=caf%C3%A9 returns both Files. GET /api/v1/search?q=cafe%CC%81 returns HTTP 200 with no results. The query spellings are canonically equivalent. Source evidence: crates/calternal-search/src/query.rs:222 passes query text directly to SimpleTokenizer. crates/calternal-search/src/index.rs:603 does the same for filename tokens. Neither normalizes canonical Unicode sequences first. The same query_words implementation is present at production dev 6074f71d1. This job will add exact dev-binary evidence on #867 when its comparison completes. Expected: canonically equivalent NFC and NFD terms find the same items. Keep the stable Files path and source bytes unchanged; normalize derived Search terms with the existing Unicode rule (DESIGN §§32, 40). Cover names, titles, headings and body text as well as query terms. Reuse the shared normalization primitive. Test: index one NFC and one NFD text, search with both query spellings, compare returned stable item identities and verify the original source bytes stay unchanged. Run the focused live check on dev and the fix. The missing ASCII sentinel in round-3 is a separate timing finding: a local run took about eight seconds to find it, while search_chaos.py waits five seconds.
Author
Owner

Started #1044 on branch job/nfd-1044 from dev base c39ffe5d90. The issue evidence and #1045 Search comments confirm canonically equivalent NFC/NFD Search terms can miss indexed filenames or text. I am reusing the existing Search generation rebuild path for the one-time derived Index refresh, and I am checking snippet offsets before choosing how to normalize indexed text.

Started #1044 on branch job/nfd-1044 from dev base c39ffe5d90126527d7aacf2d8b79507929c80616. The issue evidence and #1045 Search comments confirm canonically equivalent NFC/NFD Search terms can miss indexed filenames or text. I am reusing the existing Search generation rebuild path for the one-time derived Index refresh, and I am checking snippet offsets before choosing how to normalize indexed text.
Author
Owner

Migration finding: changing the Tantivy analyzer schema rebuilds the shared Index, but an existing private Index is rejected by QueryIndex::open_user and the existing readiness sweep skips any User whose marker is still current. Search currently stores a one-byte per-User readiness marker and already has a paged scheduler that enqueues durable per-User staged rebuild jobs. I will version that marker with the analyzer change so each old private generation is rebuilt once through the existing bounded job path; source files remain untouched.

Migration finding: changing the Tantivy analyzer schema rebuilds the shared Index, but an existing private Index is rejected by `QueryIndex::open_user` and the existing readiness sweep skips any User whose marker is still current. Search currently stores a one-byte per-User readiness marker and already has a paged scheduler that enqueues durable per-User staged rebuild jobs. I will version that marker with the analyzer change so each old private generation is rebuilt once through the existing bounded job path; source files remain untouched.
Author
Owner

Finding in the existing Search Chaos probe: its NFC/NFD upload case checked the two ASCII body sentinels, but it did not query the shared filename token using both composed and decomposed spellings. During the live round, direct requests for café and cafe\u0301 both returned the two uploaded files (HTTP 200). I added a probe assertion that both spellings return both normalized filenames and a focused test for highlight-safe title matching.

Finding in the existing Search Chaos probe: its NFC/NFD upload case checked the two ASCII body sentinels, but it did not query the shared filename token using both composed and decomposed spellings. During the live round, direct requests for `café` and `cafe\u0301` both returned the two uploaded files (HTTP 200). I added a probe assertion that both spellings return both normalized filenames and a focused test for highlight-safe title matching.
Author
Owner

#1044 final report

Branch: job/nfd-1044 (base c39ffe5d9)
Head: a59b89a6e629856c5e164f9984d7adab1176b425

Built

  • Added a Search tokenizer that keeps combining marks with their word, normalizes each word to NFC, then uses the same Tantivy LowerCaser as the indexed and queried text. It retains byte offsets into the original source, so snippets highlight the original composed or decomposed spelling. Filename prefixes use NFC terms too.
  • Changed the analyzer identifiers so shared Tantivy schemas rebuild. Changed the private Search ready marker from b"1" to b"2". The existing paged User sweep now treats old private generations as stale and enqueues the existing durable staged rebuild job. A published generation gets the new marker.
  • Added NFC/NFD, mixed-script, original-spelling highlight, legacy-marker and Search Chaos query coverage. Extended the existing 100k mixed and 1M stress Search profiles with both canonical query forms; the mixed profile also cycles them during a full rebuild burst and records query latency and process CPU/RSS.

Files

Cargo.lock; crates/calternal-search/Cargo.toml; crates/calternal-search/src/index.rs; crates/calternal-search/src/query.rs; crates/calternal-fs/src/root.rs; crates/calternal-fs/src/lib.rs; tests/adversarial/search_chaos.py; tests/adversarial/test_search_result_matching.py; tests/perf/search_scale.py.

Verification output

cargo fmt --check exited 0 with no output.

cargo clippy -p calternal-search --all-targets -- -D warnings:

Finished `dev` profile [unoptimized + debuginfo] target(s) in 3m 30s

cargo test -p calternal-search:

test result: ok. 43 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 27.33s
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 1.12s
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.08s
test result: ok. 24 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 517.93s
test result: ok. 5 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.06s
test result: ok. 2 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.02s
test result: ok. 1 passed; 0 failed; 2 ignored; 0 measured; 0 filtered out; finished in 9.78s
test result: ok. 4 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.01s
test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s

cargo clippy -p calternal-fs --all-targets -- -D warnings:

Finished `dev` profile [unoptimized + debuginfo] target(s) in 16.91s

cargo test -p calternal-fs:

test result: ok. 59 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 15.31s
test result: ok. 44 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 4.56s
test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s

With OPENSSL_NO_VENDOR=1, cargo clippy -p calternal-server --all-targets -- -D warnings:

Finished `dev` profile [unoptimized + debuginfo] target(s) in 6m 36s

With OPENSSL_NO_VENDOR=1, cargo test -p calternal-server:

test result: ok. 163 passed; 0 failed; 6 ignored; 0 measured; 0 filtered out; finished in 59.57s

Web build and focused probe matcher tests:

✓ built in 49.55s
  Wrote site to "build"
  ✔ done
....
----------------------------------------------------------------------
Ran 4 tests in 0.006s

OK

The benchmark Python files compiled, and a fixture check verified both spellings and filenames. git fetch origin && git merge origin/dev reported Already up to date. Cleanup reported:

Removed 16591 files, 9.8GiB total

Search Chaos and remaining gap

Direct live requests for café and cafe\u0301 each returned HTTP 200 and both expected normalized filenames. The active Search Chaos process had imported the probe before I added its explicit equivalent-query assertion, so that new assertion has only passed its focused matcher unit test, not a live run.

The one time-boxed Search Chaos round reported HTTP timeouts for rename-storm items 12–31. The runner then force-stopped its server after the 60-second SIGTERM window. The host load average was 34.19, 39.36, 38.28 during the round and later 28.86, 28.90, 32.57. These were slow-only observations; I found no authorization, hostile-input acceptance, data-loss or data-corruption finding. The round did not finish cleanly. I did not repeat it.

The profile was extended but not measured: this is a correctness issue, and the current verification policy restricts performance measurements to performance issues on the perf VM. Therefore there are no new numbers for docs/perf/baseline.json.

UX gaps

  • Closed: equivalent query spellings now select the same indexed items, and snippet offsets retain source text.
  • Left: none; no UI screen changed.

Decisions not specified by DESIGN

  • Reused calternal_notes_core::normalize_tag_value for NFC and Tantivy LowerCaser for case handling.
  • Used versioned analyzer names and a one-byte private readiness version (b"2") to invoke the existing bounded, resumable rebuild flow. Source documents stay unchanged.
  • Compared Search Chaos titles after NFC normalization because Files stores the uploaded names in NFC.

READY FOR MERGE

No. The Rust and web gates pass, but Search Chaos did not complete cleanly, and the new live NFC/NFD assertion was not exercised in that run. The merge round should run the focused Search Chaos probe once with the updated script and confirm both query spellings return the same two titles.

## #1044 final report Branch: `job/nfd-1044` (base `c39ffe5d9`) Head: `a59b89a6e629856c5e164f9984d7adab1176b425` ### Built - Added a Search tokenizer that keeps combining marks with their word, normalizes each word to NFC, then uses the same Tantivy `LowerCaser` as the indexed and queried text. It retains byte offsets into the original source, so snippets highlight the original composed or decomposed spelling. Filename prefixes use NFC terms too. - Changed the analyzer identifiers so shared Tantivy schemas rebuild. Changed the private Search ready marker from `b"1"` to `b"2"`. The existing paged User sweep now treats old private generations as stale and enqueues the existing durable staged rebuild job. A published generation gets the new marker. - Added NFC/NFD, mixed-script, original-spelling highlight, legacy-marker and Search Chaos query coverage. Extended the existing 100k mixed and 1M stress Search profiles with both canonical query forms; the mixed profile also cycles them during a full rebuild burst and records query latency and process CPU/RSS. ### Files `Cargo.lock`; `crates/calternal-search/Cargo.toml`; `crates/calternal-search/src/index.rs`; `crates/calternal-search/src/query.rs`; `crates/calternal-fs/src/root.rs`; `crates/calternal-fs/src/lib.rs`; `tests/adversarial/search_chaos.py`; `tests/adversarial/test_search_result_matching.py`; `tests/perf/search_scale.py`. ### Verification output `cargo fmt --check` exited 0 with no output. `cargo clippy -p calternal-search --all-targets -- -D warnings`: ```text Finished `dev` profile [unoptimized + debuginfo] target(s) in 3m 30s ``` `cargo test -p calternal-search`: ```text test result: ok. 43 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 27.33s test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 1.12s test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.08s test result: ok. 24 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 517.93s test result: ok. 5 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.06s test result: ok. 2 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.02s test result: ok. 1 passed; 0 failed; 2 ignored; 0 measured; 0 filtered out; finished in 9.78s test result: ok. 4 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.01s test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s ``` `cargo clippy -p calternal-fs --all-targets -- -D warnings`: ```text Finished `dev` profile [unoptimized + debuginfo] target(s) in 16.91s ``` `cargo test -p calternal-fs`: ```text test result: ok. 59 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 15.31s test result: ok. 44 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 4.56s test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s ``` With `OPENSSL_NO_VENDOR=1`, `cargo clippy -p calternal-server --all-targets -- -D warnings`: ```text Finished `dev` profile [unoptimized + debuginfo] target(s) in 6m 36s ``` With `OPENSSL_NO_VENDOR=1`, `cargo test -p calternal-server`: ```text test result: ok. 163 passed; 0 failed; 6 ignored; 0 measured; 0 filtered out; finished in 59.57s ``` Web build and focused probe matcher tests: ```text ✓ built in 49.55s Wrote site to "build" ✔ done .... ---------------------------------------------------------------------- Ran 4 tests in 0.006s OK ``` The benchmark Python files compiled, and a fixture check verified both spellings and filenames. `git fetch origin && git merge origin/dev` reported `Already up to date.` Cleanup reported: ```text Removed 16591 files, 9.8GiB total ``` ### Search Chaos and remaining gap Direct live requests for `café` and `cafe\u0301` each returned HTTP 200 and both expected normalized filenames. The active Search Chaos process had imported the probe before I added its explicit equivalent-query assertion, so that new assertion has only passed its focused matcher unit test, not a live run. The one time-boxed Search Chaos round reported HTTP timeouts for rename-storm items 12–31. The runner then force-stopped its server after the 60-second SIGTERM window. The host load average was 34.19, 39.36, 38.28 during the round and later 28.86, 28.90, 32.57. These were slow-only observations; I found no authorization, hostile-input acceptance, data-loss or data-corruption finding. The round did not finish cleanly. I did not repeat it. The profile was extended but not measured: this is a correctness issue, and the current verification policy restricts performance measurements to performance issues on the perf VM. Therefore there are no new numbers for `docs/perf/baseline.json`. ### UX gaps - Closed: equivalent query spellings now select the same indexed items, and snippet offsets retain source text. - Left: none; no UI screen changed. ### Decisions not specified by DESIGN - Reused `calternal_notes_core::normalize_tag_value` for NFC and Tantivy `LowerCaser` for case handling. - Used versioned analyzer names and a one-byte private readiness version (`b"2"`) to invoke the existing bounded, resumable rebuild flow. Source documents stay unchanged. - Compared Search Chaos titles after NFC normalization because Files stores the uploaded names in NFC. ### READY FOR MERGE **No.** The Rust and web gates pass, but Search Chaos did not complete cleanly, and the new live NFC/NFD assertion was not exercised in that run. The merge round should run the focused Search Chaos probe once with the updated script and confirm both query spellings return the same two titles.
Author
Owner

Merge-round-8 progress: #1044 is integrated at 7b4b37cc3. The marker now advances to v2 for the NFC analyzer; v1 is stale and enters the existing per-User staged rebuild path. cargo clippy -p calternal-search --all-targets -- -D warnings passed, and cargo test -p calternal-search passed (55 unit tests, 1 ignored; indexer 25 passed; remaining integration suites passed). The staged-rebuild test verifies Search can read the prior index until the replacement is published. Final startup/API checks are pending.

Merge-round-8 progress: #1044 is integrated at `7b4b37cc3`. The marker now advances to v2 for the NFC analyzer; v1 is stale and enters the existing per-User staged rebuild path. `cargo clippy -p calternal-search --all-targets -- -D warnings` passed, and `cargo test -p calternal-search` passed (55 unit tests, 1 ignored; indexer 25 passed; remaining integration suites passed). The staged-rebuild test verifies Search can read the prior index until the replacement is published. Final startup/API checks are pending.
Author
Owner

Merge-round-8 startup probe: on real server data with the analyzer marker reset from 2 to 1, startup advanced it to 2 in 18.2 s. I issued 16 Search API requests while the per-user rebuild was leased; all 16 succeeded against the existing index. NFC and NFD queries both returned 200 and found the NFD-source fixture. The focused search-chaos matrix is still running once against the merged branch.

Merge-round-8 startup probe: on real server data with the analyzer marker reset from 2 to 1, startup advanced it to 2 in 18.2 s. I issued 16 Search API requests while the per-user rebuild was leased; all 16 succeeded against the existing index. NFC and NFD queries both returned 200 and found the NFD-source fixture. The focused search-chaos matrix is still running once against the merged branch.
Author
Owner

Focused Search chaos pass found a live freshness failure: after successful uploads of café-nfc.txt and cafe\u0301-nfd.txt, the API returned HTTP 200 but neither content marker appeared within the 5 s wait, and both NFC and NFD café queries returned zero expected filenames. The 20,000-file watcher-overflow/rebuild phase then reached its SIGKILL/restart barrier. I am checking whether the Unicode updates appear after startup indexing settles; I have not changed the adversarial expectations.

Focused Search chaos pass found a live freshness failure: after successful uploads of `café-nfc.txt` and `cafe\u0301-nfd.txt`, the API returned HTTP 200 but neither content marker appeared within the 5 s wait, and both NFC and NFD `café` queries returned zero expected filenames. The 20,000-file watcher-overflow/rebuild phase then reached its SIGKILL/restart barrier. I am checking whether the Unicode updates appear after startup indexing settles; I have not changed the adversarial expectations.
Author
Owner

Merge-round 8 completed at 2b6c77c14be78e7d1e1030e23b14c63a6772fca7; NFC/NFD analyzer v2 and the bounded startup reindex are integrated. READY FOR STAGING: yes.

Real-server startup probe output:

search startup v2: marker=1->2, active queries=16, successful=16, completion=18212ms, NFC/NFD Search=200

The active-query storm continued to return HTTP 200 while the marker changed from 1 to 2; the rebuild completed once. The focused Search chaos probe did not meet its 5-second visibility deadline for two new Unicode files:

FAIL search chaos: search did not find uploaded marker unicodenfcsentinel
FAIL search chaos: search did not find uploaded marker unicodenfdsentinel
FAIL search chaos: NFD canonical Unicode query missed expected filenames: []

A direct live-server follow-up found the NFC filename in 64ms; NFD filename and body searches became visible by 14.6s. All four queries returned the expected results after indexing. This is a SLOW indexing delay under load, not a persistent NFC/NFD mismatch or a 5xx.

cargo clippy -p calternal-search --all-targets -- -D warnings and cargo test -p calternal-search passed (55 unit tests, 1 ignored; indexer integration 25 passed). Startup Search returned HTTP 200 throughout the rebuild. UX gap left: newly indexed Unicode content can take up to 14.6s to appear during this burst; no data was missing after indexing completed.

Merge-round 8 completed at `2b6c77c14be78e7d1e1030e23b14c63a6772fca7`; NFC/NFD analyzer v2 and the bounded startup reindex are integrated. READY FOR STAGING: yes. Real-server startup probe output: ``` search startup v2: marker=1->2, active queries=16, successful=16, completion=18212ms, NFC/NFD Search=200 ``` The active-query storm continued to return HTTP 200 while the marker changed from 1 to 2; the rebuild completed once. The focused Search chaos probe did not meet its 5-second visibility deadline for two new Unicode files: ``` FAIL search chaos: search did not find uploaded marker unicodenfcsentinel FAIL search chaos: search did not find uploaded marker unicodenfdsentinel FAIL search chaos: NFD canonical Unicode query missed expected filenames: [] ``` A direct live-server follow-up found the NFC filename in 64ms; NFD filename and body searches became visible by 14.6s. All four queries returned the expected results after indexing. This is a SLOW indexing delay under load, not a persistent NFC/NFD mismatch or a 5xx. `cargo clippy -p calternal-search --all-targets -- -D warnings` and `cargo test -p calternal-search` passed (55 unit tests, 1 ignored; indexer integration 25 passed). Startup Search returned HTTP 200 throughout the rebuild. UX gap left: newly indexed Unicode content can take up to 14.6s to appear during this burst; no data was missing after indexing completed.
Author
Owner

Deployed to production 2026-10-05 03:12 CEST in round 8 (2b6c77c14). Staging healthy first; production healthy in 33 s; /api/v1/version reports the build ID; change events 0/30 s.

Deployed to production 2026-10-05 03:12 CEST in round 8 (2b6c77c14). Staging healthy first; production healthy in 33 s; `/api/v1/version` reports the build ID; change events 0/30 s.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#1044
No description provided.