PERF: Tag suggestions rebuild whole-Home object sets and Tag pages use OFFSET #784

Open
opened 2026-10-02 13:10:40 +00:00 by kayg · 5 comments
Owner

Found in the read-mostly server architecture audit #663. Applies DESIGN §58 rules 1, 2 and 8 from queued job/instant-663. Source base origin/dev = c4a61e8cf090170f35b1bed3350d9de20c83ecd5; pending origin/job/merge-round-7a = 2f4482ded066d9c5d9c59130377907f7fd2916c9. This is structural evidence, not a measured latency or confirmed security blocker.

Context: GET /api/v1/tags supplies Tag counts/suggestions used across Tabs. Its owning crate is calternal-tags, so fixing only Notes pagination (#703) leaves this path unchanged.

Evidence:

  • crates/calternal-tags/src/lib.rs:607–624: list_tags awaits tag_counts then returns all counts in one array.
  • index.rs:563–605: tag_counts SELECT DISTINCTs every visible or hidden tagged identity for the User into fetch_all, then builds HashMap<String, HashSet<(kind,id,path)>> entries for every namespace prefix. Hidden-path filtering happens in Rust after allocation.
  • lib.rs:641–672,711–749 and index.rs:474–524: item and assignment pages allow 500 paths/items and use unsigned integer OFFSET cursors. A path page can include many Tag values. There is no response-byte cap in these handlers.
  • Round 7a retains these functions (its Tags changes concern rename/parity).

Reasoned impact: a small suggestion/count request depends on the number of tagged objects in the whole Home and namespace depth. Each namespace prefix copies identity strings into sets. Paging deep into a Tag discards previous rows and can shift when data changes. This is finite normal-data work, not a demonstrated DoS.

Concrete fix: maintain revisioned Tag and namespace-prefix counts at ingest with correct distinct-object semantics, including hidden/internal filtering. Serve bounded suggestions/count pages and item/assignment pages with signed keysets and byte caps. Reuse the Files cursor signing and #665 revision helpers; keep one calternal-tags owner. Invalidate counts atomically on rename, trash, restore, assignment and source changes. Do not change written Tag spelling or case behavior.

Regression tests: verify exact parent counts when one item has several child Tags, hidden/internal exclusions, two-User isolation and rename/trash/restore. Instrument a 100-result suggestion request and require no whole-object fetch/set build. Insert/delete between consecutive keyset pages and prove no skips or duplicate items; include many Tags on one path to verify the byte cap.

Validation for the fix: preserve existing assertions and protocol status codes. Run per-crate fmt/clippy/test, plus calternal-server if the route or provider contract changes. Extend an existing bench profile with warm/cold latency, CPU, RSS and a realistic large-data burst. Measure ≥5 samples on the perf VM under /root/perf.lock with load recorded inside the lock and HDD emulation; compare only a matching baseline. The audit itself did not run a server or benchmark. No product edit is requested from the audit branch.

Found in the read-mostly server architecture audit #663. Applies DESIGN §58 rules 1, 2 and 8 from queued `job/instant-663`. Source base `origin/dev` = `c4a61e8cf090170f35b1bed3350d9de20c83ecd5`; pending `origin/job/merge-round-7a` = `2f4482ded066d9c5d9c59130377907f7fd2916c9`. This is structural evidence, not a measured latency or confirmed security blocker. Context: `GET /api/v1/tags` supplies Tag counts/suggestions used across Tabs. Its owning crate is calternal-tags, so fixing only Notes pagination (#703) leaves this path unchanged. Evidence: - `crates/calternal-tags/src/lib.rs:607–624`: list_tags awaits tag_counts then returns all counts in one array. - `index.rs:563–605`: tag_counts SELECT DISTINCTs every visible or hidden tagged identity for the User into fetch_all, then builds HashMap<String, HashSet<(kind,id,path)>> entries for every namespace prefix. Hidden-path filtering happens in Rust after allocation. - `lib.rs:641–672,711–749` and `index.rs:474–524`: item and assignment pages allow 500 paths/items and use unsigned integer OFFSET cursors. A path page can include many Tag values. There is no response-byte cap in these handlers. - Round 7a retains these functions (its Tags changes concern rename/parity). Reasoned impact: a small suggestion/count request depends on the number of tagged objects in the whole Home and namespace depth. Each namespace prefix copies identity strings into sets. Paging deep into a Tag discards previous rows and can shift when data changes. This is finite normal-data work, not a demonstrated DoS. Concrete fix: maintain revisioned Tag and namespace-prefix counts at ingest with correct distinct-object semantics, including hidden/internal filtering. Serve bounded suggestions/count pages and item/assignment pages with signed keysets and byte caps. Reuse the Files cursor signing and #665 revision helpers; keep one calternal-tags owner. Invalidate counts atomically on rename, trash, restore, assignment and source changes. Do not change written Tag spelling or case behavior. Regression tests: verify exact parent counts when one item has several child Tags, hidden/internal exclusions, two-User isolation and rename/trash/restore. Instrument a 100-result suggestion request and require no whole-object fetch/set build. Insert/delete between consecutive keyset pages and prove no skips or duplicate items; include many Tags on one path to verify the byte cap. Validation for the fix: preserve existing assertions and protocol status codes. Run per-crate fmt/clippy/test, plus calternal-server if the route or provider contract changes. Extend an existing bench profile with warm/cold latency, CPU, RSS and a realistic large-data burst. Measure ≥5 samples on the perf VM under /root/perf.lock with load recorded inside the lock and HDD emulation; compare only a matching baseline. The audit itself did not run a server or benchmark. No product edit is requested from the audit branch.
Author
Owner

Starting work on job/webperf, based on 2f4482ded066d9c5d9c59130377907f7fd2916c9 (job/merge-round-7a). I am reading the matching audit evidence and will report the concrete finding, regression coverage, measurements, and gate output here when finished.

Starting work on `job/webperf`, based on `2f4482ded066d9c5d9c59130377907f7fd2916c9` (`job/merge-round-7a`). I am reading the matching audit evidence and will report the concrete finding, regression coverage, measurements, and gate output here when finished.
Author
Owner

SQLite audit #663. Base c4a61e8cf0; queued round-7a 2f4482ded0.

Additional collation evidence: index.rs:463 uses a BINARY total predicate; :483 uses a NOCASE item predicate. tags_lookup is BINARY. A 100k fixture split between Area/Work and area/work has 50k exact-case matches but 100k NOCASE matches. The first 101-row page can stop early via the owner/path index; the total performs about 1,100,000 VM instructions. This is an odd-but-harmless total inconsistency, not a security blocker. Keep predicates consistent with the decided case semantics; do not alter written spelling or existing expectations without an owner decision. Test mixed-case namespace children and isolation, and index the chosen collation for selective/deep cases. Reuse #784.

Figures are local Python SQLite 3.53.3 query work in 1,000-instruction callback units, not production latency. Rust bundles 3.51.3. Recheck production plans before implementation; no product edit or assertion change.

SQLite audit #663. Base c4a61e8cf090170f35b1bed3350d9de20c83ecd5; queued round-7a 2f4482ded066d9c5d9c59130377907f7fd2916c9. Additional collation evidence: index.rs:463 uses a BINARY total predicate; :483 uses a NOCASE item predicate. tags_lookup is BINARY. A 100k fixture split between Area/Work and area/work has 50k exact-case matches but 100k NOCASE matches. The first 101-row page can stop early via the owner/path index; the total performs about 1,100,000 VM instructions. This is an odd-but-harmless total inconsistency, not a security blocker. Keep predicates consistent with the decided case semantics; do not alter written spelling or existing expectations without an owner decision. Test mixed-case namespace children and isolation, and index the chosen collation for selective/deep cases. Reuse #784. Figures are local Python SQLite 3.53.3 query work in 1,000-instruction callback units, not production latency. Rust bundles 3.51.3. Recheck production plans before implementation; no product edit or assertion change.
Author
Owner

Finding: tag suggestions fetched every tag assignment and built per-object HashSets in application code. Nested distinct counts now aggregate in SQLite, retaining hidden-path filtering and User isolation; added a regression case for parent counts and nested tags. The audit's signed keyset cursor and response byte-cap changes for item and assignment pages remain open. No Tag benchmark has been run.

Finding: tag suggestions fetched every tag assignment and built per-object HashSets in application code. Nested distinct counts now aggregate in SQLite, retaining hidden-path filtering and User isolation; added a regression case for parent counts and nested tags. The audit's signed keyset cursor and response byte-cap changes for item and assignment pages remain open. No Tag benchmark has been run.
Author
Owner

F2 — P1: Tag-count SQL loses required word separators

Owner: #784. Introduced by 1063f7d83.
Evidence: crates/calternal-tags/src/index.rs:592, :593, :599, :601.
Rust's escaped newline removes the newline and following indentation. The
query therefore contains UNION ALLSELECT, THEN remainingELSE,
FROM prefixesGROUP BY and FROM distinct_itemsGROUP BY. These are SQL syntax
errors. crates/calternal-tags/src/lib.rs:612 maps this query failure to an
internal error on GET /api/v1/tags, including an empty Home. Tag suggestions
and counts cannot load. No server response or test output is claimed.

Fix: use a multiline raw SQL literal or explicit spaces before each escaped
newline. Keep distinct namespace counts, hidden-path filtering and User isolation.
Rule: CLAUDE.md requires API failures to be fixed; DESIGN §32 requires Tag search.
Test idea: run the existing nested-count regression and add a route test for
an empty Home and one nested Tag. Both must return 200 with correct counts.
Search before reporting: Tag, SQL whitespace and #784. Use #784 for the fix.

## F2 — P1: Tag-count SQL loses required word separators Owner: #784. Introduced by `1063f7d83`. Evidence: `crates/calternal-tags/src/index.rs:592`, `:593`, `:599`, `:601`. Rust's escaped newline removes the newline and following indentation. The query therefore contains `UNION ALLSELECT`, `THEN remainingELSE`, `FROM prefixesGROUP BY` and `FROM distinct_itemsGROUP BY`. These are SQL syntax errors. `crates/calternal-tags/src/lib.rs:612` maps this query failure to an internal error on `GET /api/v1/tags`, including an empty Home. Tag suggestions and counts cannot load. No server response or test output is claimed. Fix: use a multiline raw SQL literal or explicit spaces before each escaped newline. Keep distinct namespace counts, hidden-path filtering and User isolation. Rule: CLAUDE.md requires API failures to be fixed; DESIGN §32 requires Tag search. Test idea: run the existing nested-count regression and add a route test for an empty Home and one nested Tag. Both must return 200 with correct counts. Search before reporting: `Tag`, `SQL whitespace` and #784. Use #784 for the fix.
Author
Owner

Finding: tag_counts used Rust continued-line string escapes. Rust removes the line break and indentation, which joined SQL tokens such as UNION ALLSELECT and prefixesGROUP BY; the Tags list could return an internal error, including for an empty Home.

Fix: replaced the query with a raw multiline SQL literal and added a direct list-handler regression for empty and nested counts. The existing hidden-source count test now also asserts that an empty Home has no counts.

Verification: cargo clippy -p calternal-tags --all-targets -- -D warnings passed. cargo test -p calternal-tags passed: 13 passed, 0 failed; doc tests 0 passed, 0 failed. Fix commit: b918d26be.

Finding: `tag_counts` used Rust continued-line string escapes. Rust removes the line break and indentation, which joined SQL tokens such as `UNION ALLSELECT` and `prefixesGROUP BY`; the Tags list could return an internal error, including for an empty Home. Fix: replaced the query with a raw multiline SQL literal and added a direct list-handler regression for empty and nested counts. The existing hidden-source count test now also asserts that an empty Home has no counts. Verification: `cargo clippy -p calternal-tags --all-targets -- -D warnings` passed. `cargo test -p calternal-tags` passed: 13 passed, 0 failed; doc tests 0 passed, 0 failed. Fix commit: `b918d26be`.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#784
No description provided.