PERF: Photos rule 2 — bound lists by rows and bytes with signed keyset pages (#663) #682

Open
opened 2026-10-02 05:36:31 +00:00 by kayg · 2 comments
Owner

Context: #663 matrix and DESIGN §58 rule 2. Scope: Photos must bound lists by rows and bytes with signed keyset pages. Shared client machinery belongs only to #665–#668.

Evidence at c4a61e8cf0:

  • Endpoint: GET /api/v1/photos/timeline.
  • Source: crates/plugins/photos/src/index.rs:30. 90 days × 200 tiles permits 18,000 tiles. timeline/days uses unsigned offset (:653); buckets returns all days (:613). No total 100-row/byte keyset page.
  • Representative endpoint numbers, server/client revisions, fixture size, lock/HDD qualification and limits are in the #663 production/HDD table. Those numbers do not prove this rule passes; the structural gap above is separate evidence.

Expected and regression tests:
At 10k items (and largest realistic plugin data), every response has ≤100 item rows and an explicit total byte cap; nested arrays count toward both. Signed cursors bind User, view/filter/sort/zone and keyset position. Edits between pages do not skip or repeat unchanged items; malformed/cross-User/replayed scope cursors cannot change authorization. No OFFSET and no startup whole-corpus drain. Client retains first page plus one lookahead and a bounded resident row/byte window. Preserve deep links, selection, Copy link and current list order. Test large values and variable row size at the byte boundary, empty pages, filter switches, deleted cursor anchor and cache eviction.

Performance test:
Extend the existing bench profile for this hot path and the #549/#641 harness. Use production builds on root@10.69.69.63, bench/hdd-emu.sh and flock -w 14400 /root/perf.lock. Record load inside the lock; ≥5 samples, median/p95/max, average CPU/RSS and one realistic large-data/burst case. Separate warm, cold, accepted and durable boundaries. First usable 10k view ≤1.5 s; cached open/warm return ≤100 ms and accepted action ≤150 ms where applicable. Compare only a matching baseline in docs/perf/baseline.json; missing profiles require a new recorded baseline, not a made-up comparison.

Reuse/ownership:

Reuse #555 userStorage, #549 route caches, Files signed keysets/change feed and the existing Db reader_pool. #641 owns blaze measurements; #642 owns Settings opening; #640 owns Mail layouts; #639 owns linked Note opening. Integrate their active/completed branches before changing related code.

Acceptance:
Existing tests and status expectations stay intact. Run the per-crate gates (and calternal-server for route/contract changes), web gates if changed, and one time-boxed real-server regression/adversarial round for any new API contract. Preserve calternal-fs as the only filesystem interface and the server as the only writer. UI changes need pointer/touch/keyboard/screen-reader coverage and real-production captures at 390/820/1440 in both themes. Do not change shared motion for keyboard input.

Representative measurement from the audit (not a full-view budget result):
/api/v1/photos/timeline?days=30&tiles_per_day=48, five serial requests after fixture readiness, all HTTP 200. Median/p95/max 46.0/48.2/48.2 ms; five-request burst median/p95/max 62.2/62.7/62.7 ms. Serial window server CPU 10 ms, RSS 443625472 bytes; no ETag on these sampled responses. Load inside lock 3.45/2.03/0.9.
Shared release server source cc25c441b7a974185622a1dee853cf38686d2b67, binary SHA-256 2f3567d91c34839851247bc0acbc25a56aaacd14dca269b8f0342ddf83447ed9; embedded production SPA; Chromium browser on the build host through SSH/HTTPS. Server/Home/Index on perf VM HDD emulator: direct-I/O loop, 8 ms read/write dm-delay, 200 IOPS and 150 MiB/s caps. Every measured phase held flock -w 14400 /root/perf.lock. Qualification QD1 115.3 IOPS/8.028 ms median, QD16 200.7 IOPS/96.993 ms. Fixture: 366 Daily notes, 10,980 Logs, 100 Files/Photos, 20 Notes/Tasks, three Budgets and 100 transactions; Mail empty, Admin one User.
Structural source evidence above is the newer audit base, not the measured binary revision. No claim that these revisions are equivalent. The baseline in docs/perf/baseline.json uses another fixture/build/transport; no regression ratio is valid here. See #663 for matching baseline endpoint values and coverage gaps.

Context: #663 matrix and DESIGN §58 rule 2. Scope: Photos must bound lists by rows and bytes with signed keyset pages. Shared client machinery belongs only to #665–#668. Evidence at c4a61e8cf090170f35b1bed3350d9de20c83ecd5: - Endpoint: `GET /api/v1/photos/timeline`. - Source: `crates/plugins/photos/src/index.rs:30`. 90 days × 200 tiles permits 18,000 tiles. timeline/days uses unsigned offset (:653); buckets returns all days (:613). No total 100-row/byte keyset page. - Representative endpoint numbers, server/client revisions, fixture size, lock/HDD qualification and limits are in the #663 production/HDD table. Those numbers do not prove this rule passes; the structural gap above is separate evidence. Expected and regression tests: At 10k items (and largest realistic plugin data), every response has ≤100 item rows and an explicit total byte cap; nested arrays count toward both. Signed cursors bind User, view/filter/sort/zone and keyset position. Edits between pages do not skip or repeat unchanged items; malformed/cross-User/replayed scope cursors cannot change authorization. No OFFSET and no startup whole-corpus drain. Client retains first page plus one lookahead and a bounded resident row/byte window. Preserve deep links, selection, Copy link and current list order. Test large values and variable row size at the byte boundary, empty pages, filter switches, deleted cursor anchor and cache eviction. Performance test: Extend the existing bench profile for this hot path and the #549/#641 harness. Use production builds on root@10.69.69.63, bench/hdd-emu.sh and flock -w 14400 /root/perf.lock. Record load inside the lock; ≥5 samples, median/p95/max, average CPU/RSS and one realistic large-data/burst case. Separate warm, cold, accepted and durable boundaries. First usable 10k view ≤1.5 s; cached open/warm return ≤100 ms and accepted action ≤150 ms where applicable. Compare only a matching baseline in docs/perf/baseline.json; missing profiles require a new recorded baseline, not a made-up comparison. Reuse/ownership: Reuse #555 userStorage, #549 route caches, Files signed keysets/change feed and the existing Db reader_pool. #641 owns blaze measurements; #642 owns Settings opening; #640 owns Mail layouts; #639 owns linked Note opening. Integrate their active/completed branches before changing related code. Acceptance: Existing tests and status expectations stay intact. Run the per-crate gates (and calternal-server for route/contract changes), web gates if changed, and one time-boxed real-server regression/adversarial round for any new API contract. Preserve calternal-fs as the only filesystem interface and the server as the only writer. UI changes need pointer/touch/keyboard/screen-reader coverage and real-production captures at 390/820/1440 in both themes. Do not change shared motion for keyboard input. Representative measurement from the audit (not a full-view budget result): `/api/v1/photos/timeline?days=30&tiles_per_day=48`, five serial requests after fixture readiness, all HTTP 200. Median/p95/max 46.0/48.2/48.2 ms; five-request burst median/p95/max 62.2/62.7/62.7 ms. Serial window server CPU 10 ms, RSS 443625472 bytes; no ETag on these sampled responses. Load inside lock 3.45/2.03/0.9. Shared release server source `cc25c441b7a974185622a1dee853cf38686d2b67`, binary SHA-256 `2f3567d91c34839851247bc0acbc25a56aaacd14dca269b8f0342ddf83447ed9`; embedded production SPA; Chromium browser on the build host through SSH/HTTPS. Server/Home/Index on perf VM HDD emulator: direct-I/O loop, 8 ms read/write dm-delay, 200 IOPS and 150 MiB/s caps. Every measured phase held `flock -w 14400 /root/perf.lock`. Qualification QD1 115.3 IOPS/8.028 ms median, QD16 200.7 IOPS/96.993 ms. Fixture: 366 Daily notes, 10,980 Logs, 100 Files/Photos, 20 Notes/Tasks, three Budgets and 100 transactions; Mail empty, Admin one User. Structural source evidence above is the newer audit base, not the measured binary revision. No claim that these revisions are equivalent. The baseline in docs/perf/baseline.json uses another fixture/build/transport; no regression ratio is valid here. See #663 for matching baseline endpoint values and coverage gaps.
Author
Owner

Client runtime audit for #663; keep this evidence with #671/#682.

origin/dev c4a61e8cf0; timeline.svelte.ts is unchanged on merge-round-7a 2f4482ded0. PhotoTimelineData.loaded (:48), #aspects (:61) and #dayIndex (:62) retain visited days/tiles. #store and #append (:275 onward) add/merge pages; there is no eviction for unchanged offscreen days. Bucket refresh drops changed or removed days only. ordered (:136–145) rebuilds all loaded tile order; PhotosView.svelte:151–153 rebuilds order and tileByKey as data changes.

The DOM timeline is virtualised; the resident model is not bounded. A long scroll gradually retains the library and increases later page/order work. Use the shared snapshot/window limits from #666, evict offscreen tile/aspect/index data, retain small bucket geometry, and ensure selection/viewer navigation can reload evicted items. Test long traversal of many days and a very large single day, bounded resident rows/bytes, backward travel and selection after eviction. This is code evidence, not a measured RSS regression.

Client runtime audit for #663; keep this evidence with #671/#682. origin/dev c4a61e8cf090170f35b1bed3350d9de20c83ecd5; timeline.svelte.ts is unchanged on merge-round-7a 2f4482ded066d9c5d9c59130377907f7fd2916c9. PhotoTimelineData.loaded (:48), #aspects (:61) and #dayIndex (:62) retain visited days/tiles. #store and #append (:275 onward) add/merge pages; there is no eviction for unchanged offscreen days. Bucket refresh drops changed or removed days only. ordered (:136–145) rebuilds all loaded tile order; PhotosView.svelte:151–153 rebuilds order and tileByKey as data changes. The DOM timeline is virtualised; the resident model is not bounded. A long scroll gradually retains the library and increases later page/order work. Use the shared snapshot/window limits from #666, evict offscreen tile/aspect/index data, retain small bucket geometry, and ensure selection/viewer navigation can reload evicted items. Test long traversal of many days and a very large single day, bounded resident rows/bytes, backward travel and selection after eviction. This is code evidence, not a measured RSS regression.
Author
Owner

SQLite audit #663 on job/perf-arch-db. Base c4a61e8cf0; queued merge-round-7a 2f4482ded0 checked separately.

Photos SQL work evidence: photos/src/routes.rs:832 ranks all selected days with ROW_NUMBER before filtering day_rank. With 10k synthetic groups in one day, a 100-tile read does about 1,003,000 VM instructions; the ordered early-LIMIT comparison does about 4,000. Both use photos_groups_timeline and the files_index_identity join. The original also materializes/sorts ranked results. A response row cap does not bound work. Use per-day bounded seek/LIMIT probes (or an equivalent proven early-stop plan), preserve Files freshness/access joins and group ordering, and test a dense day plus several days/roots. The single-day comparison is a query hypothesis, not a drop-in multi-day fix. Round 7a retains the ranking query.

Local query-work figures use Python SQLite 3.53.3 and synthetic data; VM callbacks count 1,000-instruction units. They are not production latency, CPU/RSS or HDD budget samples. The locked Rust bundled SQLite is 3.51.3. Repeat with audit-sqlite.py --plans and confirm the production-version plan before implementing. No existing assertion or product behavior was changed. Reuse this issue instead of filing a duplicate.

SQLite audit #663 on job/perf-arch-db. Base c4a61e8cf090170f35b1bed3350d9de20c83ecd5; queued merge-round-7a 2f4482ded066d9c5d9c59130377907f7fd2916c9 checked separately. Photos SQL work evidence: photos/src/routes.rs:832 ranks all selected days with ROW_NUMBER before filtering day_rank. With 10k synthetic groups in one day, a 100-tile read does about 1,003,000 VM instructions; the ordered early-LIMIT comparison does about 4,000. Both use photos_groups_timeline and the files_index_identity join. The original also materializes/sorts ranked results. A response row cap does not bound work. Use per-day bounded seek/LIMIT probes (or an equivalent proven early-stop plan), preserve Files freshness/access joins and group ordering, and test a dense day plus several days/roots. The single-day comparison is a query hypothesis, not a drop-in multi-day fix. Round 7a retains the ranking query. Local query-work figures use Python SQLite 3.53.3 and synthetic data; VM callbacks count 1,000-instruction units. They are not production latency, CPU/RSS or HDD budget samples. The locked Rust bundled SQLite is 3.51.3. Repeat with audit-sqlite.py --plans and confirm the production-version plan before implementing. No existing assertion or product behavior was changed. Reuse this issue instead of filing a duplicate.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#682
No description provided.