Mail: reduce full-snapshot SELECT cost for large cached mailboxes #1070

Open
opened 2026-10-04 22:01:07 +00:00 by kayg · 0 comments
Owner

Mail round 2 (#1038/#1067) removes cached SELECT dependence on Index and Security writers. The required one-account/12-client 1,800-second phase passes. The existing full-snapshot cost remains visible at large mailbox sizes (DESIGN §§53, 59).

Measured on the perf VM with /root/perf.lock, using a Mail test executable built locally in release mode from the repaired product. No build ran on the VM. The fixture has three Connected Accounts under one User. Every cached SELECT sample and the 12-client burst keep the ordinary writer occupied.

Profile Warm SELECT p50 / p95 Whole-worker mean CPU Whole-worker mean / peak RSS Largest 12-client SELECT duration
10,001 messages, median of 3 serial runs 134.51 / 176.76 ms 147.54% 111.03 / 153.00 MiB 1,696.68 ms
100,000 messages 1,953.88 / 5,540.95 ms 163.04% 848.05 / 1,205.99 MiB 18,724.86 ms
3 workers, each with 100,000 messages and its own Index; median latency/per-worker resources 1,790.76 / 4,051.31 ms 113.24% 891.19 / 1,205.86 MiB 24,210.30 ms

CPU and RSS include fixture seeding, metadata and durable queue phases; they are not SELECT-only resource numbers. CPU can exceed 100% because SQLite work uses multiple threads. Serial p50/p95 and the 12-client duration measure SELECT itself. The load averages inside the lock were 6.125 / 4.5923 / 2.1152 before and 7.5220 / 5.5894 / 3.3701 after. The starting average includes the preceding debug diagnostic profile; do not describe this as a cold quiet-host baseline.

docs/perf/baseline.json has no comparable Mail proxy SELECT profile. Its mail.accounts values (HTTP account listing, p50 1.3 / p95 3.8 ms) cannot establish a SELECT regression. This finding records an existing large-snapshot performance gap, not a measured baseline regression. The profile completed with no test failure. It does not replace real-provider 100,000-message sync or 50-client acceptance.

Reproduce with the current prebuilt release Mail test executable (never compile on the VM):

flock /root/perf.lock python3 bench/mail-sync.py <prebuilt-release-mail-test-executable> --label perf-vm-release --test profile_mail_proxy --profile-prefix 'MAIL_PROXY_PROFILE '

Reduce large cached-view allocation and repeated projection work. Reuse the shared IMAP snapshot/row primitives. Keep stable UIDs, epochs, flags, current User/service visibility, queued intent and the single-writer clock checks. Keep metadata/body limits and the ordinary/Security durability policies unchanged. Extend this existing profile and its held-writer regression so the optimization has before/after evidence.

Evidence: artifacts/round2-select-perf-vm-release.log in the mailround-1038 worktree. Product repair 9c8d1e7630aafae04c43036cba7b80fd64e2b932; topology regression a6e0ff6a5e8ab501a3c5739b0ada54b24dc5e0fa.

Mail round 2 (#1038/#1067) removes cached SELECT dependence on Index and Security writers. The required one-account/12-client 1,800-second phase passes. The existing full-snapshot cost remains visible at large mailbox sizes (DESIGN §§53, 59). Measured on the perf VM with `/root/perf.lock`, using a Mail test executable built locally in release mode from the repaired product. No build ran on the VM. The fixture has three Connected Accounts under one User. Every cached SELECT sample and the 12-client burst keep the ordinary writer occupied. | Profile | Warm SELECT p50 / p95 | Whole-worker mean CPU | Whole-worker mean / peak RSS | Largest 12-client SELECT duration | |---|---|---|---|---| | 10,001 messages, median of 3 serial runs | 134.51 / 176.76 ms | 147.54% | 111.03 / 153.00 MiB | 1,696.68 ms | | 100,000 messages | 1,953.88 / 5,540.95 ms | 163.04% | 848.05 / 1,205.99 MiB | 18,724.86 ms | | 3 workers, each with 100,000 messages and its own Index; median latency/per-worker resources | 1,790.76 / 4,051.31 ms | 113.24% | 891.19 / 1,205.86 MiB | 24,210.30 ms | CPU and RSS include fixture seeding, metadata and durable queue phases; they are not SELECT-only resource numbers. CPU can exceed 100% because SQLite work uses multiple threads. Serial p50/p95 and the 12-client duration measure SELECT itself. The load averages inside the lock were 6.125 / 4.5923 / 2.1152 before and 7.5220 / 5.5894 / 3.3701 after. The starting average includes the preceding debug diagnostic profile; do not describe this as a cold quiet-host baseline. `docs/perf/baseline.json` has no comparable Mail proxy SELECT profile. Its `mail.accounts` values (HTTP account listing, p50 1.3 / p95 3.8 ms) cannot establish a SELECT regression. This finding records an existing large-snapshot performance gap, not a measured baseline regression. The profile completed with no test failure. It does not replace real-provider 100,000-message sync or 50-client acceptance. Reproduce with the current prebuilt release Mail test executable (never compile on the VM): ``` flock /root/perf.lock python3 bench/mail-sync.py <prebuilt-release-mail-test-executable> --label perf-vm-release --test profile_mail_proxy --profile-prefix 'MAIL_PROXY_PROFILE ' ``` Reduce large cached-view allocation and repeated projection work. Reuse the shared IMAP snapshot/row primitives. Keep stable UIDs, epochs, flags, current User/service visibility, queued intent and the single-writer clock checks. Keep metadata/body limits and the ordinary/Security durability policies unchanged. Extend this existing profile and its held-writer regression so the optimization has before/after evidence. Evidence: `artifacts/round2-select-perf-vm-release.log` in the mailround-1038 worktree. Product repair `9c8d1e7630aafae04c43036cba7b80fd64e2b932`; topology regression `a6e0ff6a5e8ab501a3c5739b0ada54b24dc5e0fa`.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#1070
No description provided.