Perf: switching modes in the tray paints too slowly (measure per mode pair, then fix) #424

Open
opened 2026-09-29 09:34:41 +00:00 by kayg · 75 comments
Owner

Bug (owner, 2026-09-29, calternal.cloud)

"Changing between modes in the tabs is too slow! Please investigate. We need to animate it fast, which is another issue, but the painting itself when moving between tabs is not very fast."
Scope: the paint latency of a mode switch (tray tab click to the new mode fully painted). The transition animation is out of scope here (#414 jank and the motion issues own it). Do not change the animation.

Measure first (exact protocol, decision rule fixed up front)

  • Use a production build on a quiet host: the perf-test VM, netbird ssh --no-browser --user root 10.69.69.63. The perf-367 job also uses it, so coordinate by running at different times: check uptime below 1 and that no other server is running, before each run.
  • Seed a User shaped like the owner's: about 600 Documents, a few thousand Photos, a year of daily notes with Log entries, Calendar events, Tasks, and a custom background picture. All data is synthetic, generated under bench/.
  • For every ordered pair of modes (Calendar, Notes, Files, Photos, Tasks, Mail, Money if present, Search, Settings), measure cold (first visit in the session) and warm (visited before), with the Chrome performance trace and 4x CPU throttle plus unthrottled:
    1. click to the tray's selected state painted (input feedback);
    2. click to the first frame of new-mode content;
    3. click to visually complete (no more layout shifts, primary list rendered);
    4. long tasks over 50 ms, and what they are: chunk load, data fetch waterfall, Svelte mount cost, style recalc or layout, backdrop-filter or glass compositing, image decode, or the background picture re-decode.
      Report median and p95 of 10 interleaved runs per pair.
  • Say why for the slowest pairs: the dominant cost, from the trace.

Targets

  • Tray feedback in the next frame (≤ 16 ms).
  • A warm switch visually complete in ≤ 100 ms, and a cold one in ≤ 300 ms, on the unthrottled VM.
  • No long task over 50 ms on a warm switch.

Fix candidates (keep one only if it improves its metric by ≥ 10%, median of 5 A/B runs, with no other tracked metric regressing more than 3%, and all gates green)

  • Keep recently used modes mounted (a keep-alive with a bounded count, a named config value) instead of tearing down and remounting.
  • Preload a mode's code and data on tray hover or focus (isModePreloaded exists in +layout.svelte; extend it).
  • Stale-while-revalidate data per mode: render the cached view at once, then refresh.
  • Remove fetch waterfalls: parallel loads, one request per view.
  • Avoid a full-screen backdrop-filter repaint on switch; keep the background picture decoded once.
  • Virtualise heavy lists if the mount cost dominates.

Proof

Commit the before and after tables under docs/perf/runs/, with the traces summarised (no raw multi-MB traces in the repo). Add an e2e perf guard that measures warm switch latency and reports it; it is not a merge gate (CLAUDE.md). Code comments must explain each cache or preload and its bounds. Gates as in the preamble.

## Bug (owner, 2026-09-29, calternal.cloud) "Changing between modes in the tabs is too slow! Please investigate. We need to animate it fast, which is another issue, but the painting itself when moving between tabs is not very fast." Scope: the **paint latency** of a mode switch (tray tab click to the new mode fully painted). The transition animation is out of scope here (#414 jank and the motion issues own it). Do not change the animation. ## Measure first (exact protocol, decision rule fixed up front) - Use a production build on a quiet host: the perf-test VM, `netbird ssh --no-browser --user root 10.69.69.63`. The perf-367 job also uses it, so coordinate by running at different times: check `uptime` below 1 and that no other server is running, before each run. - Seed a User shaped like the owner's: about 600 Documents, a few thousand Photos, a year of daily notes with Log entries, Calendar events, Tasks, and a custom background picture. All data is synthetic, generated under `bench/`. - For every ordered pair of modes (Calendar, Notes, Files, Photos, Tasks, Mail, Money if present, Search, Settings), measure cold (first visit in the session) and warm (visited before), with the Chrome performance trace and 4x CPU throttle plus unthrottled: 1. click to the tray's selected state painted (input feedback); 2. click to the first frame of new-mode content; 3. click to visually complete (no more layout shifts, primary list rendered); 4. long tasks over 50 ms, and what they are: chunk load, data fetch waterfall, Svelte mount cost, style recalc or layout, `backdrop-filter` or glass compositing, image decode, or the background picture re-decode. Report median and p95 of 10 interleaved runs per pair. - Say **why** for the slowest pairs: the dominant cost, from the trace. ## Targets - Tray feedback in the next frame (≤ 16 ms). - A warm switch visually complete in ≤ 100 ms, and a cold one in ≤ 300 ms, on the unthrottled VM. - No long task over 50 ms on a warm switch. ## Fix candidates (keep one only if it improves its metric by ≥ 10%, median of 5 A/B runs, with no other tracked metric regressing more than 3%, and all gates green) - Keep recently used modes mounted (a keep-alive with a bounded count, a named config value) instead of tearing down and remounting. - Preload a mode's code and data on tray hover or focus (`isModePreloaded` exists in `+layout.svelte`; extend it). - Stale-while-revalidate data per mode: render the cached view at once, then refresh. - Remove fetch waterfalls: parallel loads, one request per view. - Avoid a full-screen `backdrop-filter` repaint on switch; keep the background picture decoded once. - Virtualise heavy lists if the mount cost dominates. ## Proof Commit the before and after tables under `docs/perf/runs/`, with the traces summarised (no raw multi-MB traces in the repo). Add an e2e perf guard that measures warm switch latency and reports it; it is not a merge gate (CLAUDE.md). Code comments must explain each cache or preload and its bounds. Gates as in the preamble.
Author
Owner

Starting #424 on branch job/modes-424, based on dev at 191b179baa. I will measure mode-switch paint latency on the specified production perf VM before changing the route path.

Starting #424 on branch job/modes-424, based on dev at 191b179baac3ef4f5bebfe07ce91c4b7a887ace2. I will measure mode-switch paint latency on the specified production perf VM before changing the route path.
Author
Owner

The perf-test VM is busy, so I have not taken a baseline: uptime reports load averages 5.06, 3.32, 1.96 and /root/jank414/target/debug/calternal-server plus its Playwright Chromium are running. I am proceeding with code-path and measurement harness inspection while waiting for a quiet window.

The perf-test VM is busy, so I have not taken a baseline: uptime reports load averages 5.06, 3.32, 1.96 and /root/jank414/target/debug/calternal-server plus its Playwright Chromium are running. I am proceeding with code-path and measurement harness inspection while waiting for a quiet window.
Author
Owner

Measurement harness progress and findings:

  • Current tray catalog: DESIGN §29 sets Files, Calendar, and Photos as the default. PLUGIN_NAVIGATION adds optional Analytics and Ask. Notes, Tasks, Mail, Money, Search, and Settings are not tray modes in this build. The benchmark will measure every enabled tray tab and list the requested modes that are absent.
  • The harness now seeds 600 Documents, 3,000 Photos, 365 Daily notes with Log entries and Tasks, 180 CalDAV Events, and an Appearance background through real local APIs. Cold and warm visits stay in one SPA session; the warm visit follows a prior mode visit. Reports include p50/p95, 1×/4× CPU, reduced Chrome long-task summaries, and resource timing tables.
  • I have not recorded a baseline. At 12:10 UTC the perf VM had load averages 4.62, 5.14, 4.33 and /opt/calternal-webdav-perf/bin/calternal-server was running. This does not meet the issue's load-below-1 and no-other-server requirement.
Measurement harness progress and findings: - Current tray catalog: DESIGN §29 sets Files, Calendar, and Photos as the default. `PLUGIN_NAVIGATION` adds optional Analytics and Ask. Notes, Tasks, Mail, Money, Search, and Settings are not tray modes in this build. The benchmark will measure every enabled tray tab and list the requested modes that are absent. - The harness now seeds 600 Documents, 3,000 Photos, 365 Daily notes with Log entries and Tasks, 180 CalDAV Events, and an Appearance background through real local APIs. Cold and warm visits stay in one SPA session; the warm visit follows a prior mode visit. Reports include p50/p95, 1×/4× CPU, reduced Chrome long-task summaries, and resource timing tables. - I have not recorded a baseline. At 12:10 UTC the perf VM had load averages 4.62, 5.14, 4.33 and `/opt/calternal-webdav-perf/bin/calternal-server` was running. This does not meet the issue's load-below-1 and no-other-server requirement.
Author
Owner

The Files tab opens /files at Home (selectMode navigates to the registry href). A fixture with 600 files nested under Documents/ would make the measured Files screen show only the folder, so it would not exercise the requested large list. I am seeding the 600 synthetic document files at Home and will assert the Home listing count before measuring.

The Files tab opens `/files` at Home (`selectMode` navigates to the registry href). A fixture with 600 files nested under `Documents/` would make the measured Files screen show only the folder, so it would not exercise the requested large list. I am seeding the 600 synthetic document files at Home and will assert the Home listing count before measuring.
Author
Owner

Measurement definition recorded because DESIGN does not specify a paint-completion quiet interval: after the primary content marker appears, the benchmark requires the mode's layout box and descendant count to stay stable for 32 ms (two 60 Hz frame intervals) and waits past observed Layout Instability entries. It samples offsets and dimensions, so transform animation remains outside this paint metric as #424 requires. The probe records layout shifts even when Chrome marks them as input-related.

Measurement definition recorded because DESIGN does not specify a paint-completion quiet interval: after the primary content marker appears, the benchmark requires the mode's layout box and descendant count to stay stable for 32 ms (two 60 Hz frame intervals) and waits past observed Layout Instability entries. It samples offsets and dimensions, so transform animation remains outside this paint metric as #424 requires. The probe records layout shifts even when Chrome marks them as input-related.
Author
Owner

The responsive evidence pass exposed a probe-selector mismatch: Calendar Today uses .item in AgendaList.svelte, while Day/Week uses .block in GridColumn.svelte. The desktop benchmark stays on the Week view, but phone captures need to wait for either real primary marker. I am updating the shared marker to cover both layouts.

The responsive evidence pass exposed a probe-selector mismatch: Calendar Today uses `.item` in `AgendaList.svelte`, while Day/Week uses `.block` in `GridColumn.svelte`. The desktop benchmark stays on the Week view, but phone captures need to wait for either real primary marker. I am updating the shared marker to cover both layouts.
Author
Owner

Perf VM check at 10:29 UTC: the #414 browser/server had exited, but the WebDAV perf fixture service was still running and the 1-minute load average was 4.45. The required quiet-host threshold (<1) is not met, so no build or measurement was started. I left the shared processes untouched and will resume when the host is idle.

Perf VM check at 10:29 UTC: the #414 browser/server had exited, but the WebDAV perf fixture service was still running and the 1-minute load average was 4.45. The required quiet-host threshold (<1) is not met, so no build or measurement was started. I left the shared processes untouched and will resume when the host is idle.
Author
Owner

Perf VM recheck at 10:35 UTC: load averaged 2.41; the #414 debug server and seven headless Chromium processes were active alongside the WebDAV fixture service. The host is still outside the required quiet window. I have not started a build or measurement.

Perf VM recheck at 10:35 UTC: load averaged 2.41; the #414 debug server and seven headless Chromium processes were active alongside the WebDAV fixture service. The host is still outside the required quiet window. I have not started a build or measurement.
Author
Owner

Harness review finding: resourceSummary checked /api/ before the resource initiator. That labeled an API-backed <img> or CSS image as a data fetch and could hide the custom wallpaper request in the image-load report. I moved image initiators ahead of the API-path rule. node --check apps/web/e2e/mode-switch-perf.mjs and git diff --check pass.

Harness review finding: `resourceSummary` checked `/api/` before the resource initiator. That labeled an API-backed `<img>` or CSS image as a data fetch and could hide the custom wallpaper request in the image-load report. I moved image initiators ahead of the API-path rule. `node --check apps/web/e2e/mode-switch-perf.mjs` and `git diff --check` pass.
Author
Owner

Perf VM recheck at 10:44 UTC: another #414 server/Chromium run started; the 1-minute load average rose to 4.18. The WebDAV perf fixture service also remains active. No #424 build or timing has started because the required quiet-host condition is unmet.

Perf VM recheck at 10:44 UTC: another #414 server/Chromium run started; the 1-minute load average rose to 4.18. The WebDAV perf fixture service also remains active. No #424 build or timing has started because the required quiet-host condition is unmet.
Author
Owner

Perf VM diagnosis at 10:46 UTC: load was 3.79. ps showed /opt/calternal-webdav-perf/bin/calternal-server using 87.6% CPU, plus the WebDAV fixture service. This confirms the ongoing performance service is actively loading the VM; I have not started #424 work there.

Perf VM diagnosis at 10:46 UTC: load was 3.79. `ps` showed `/opt/calternal-webdav-perf/bin/calternal-server` using 87.6% CPU, plus the WebDAV fixture service. This confirms the ongoing performance service is actively loading the VM; I have not started #424 work there.
Author
Owner

Harness review finding: visually-complete latency was read with performance.now() after Playwright returned from the page wait, so it included browser-driver round-trip time. Layout-shift timestamps also used observer delivery time, which could miscount shifts buffered before a click. The probe now stores the completion time and each shift's entry.startTime in the page clock. node --check apps/web/e2e/mode-switch-perf.mjs and git diff --check pass.

Harness review finding: visually-complete latency was read with `performance.now()` after Playwright returned from the page wait, so it included browser-driver round-trip time. Layout-shift timestamps also used observer delivery time, which could miscount shifts buffered before a click. The probe now stores the completion time and each shift's `entry.startTime` in the page clock. `node --check apps/web/e2e/mode-switch-perf.mjs` and `git diff --check` pass.
Author
Owner

Perf VM recheck at 10:54 UTC: load was 4.38. The WebDAV perf server was using 89.3% CPU, and the #414 server/Chromium run was active. The required <1 load/no-other-server window has not occurred, so #424 measurements remain unstarted.

Perf VM recheck at 10:54 UTC: load was 4.38. The WebDAV perf server was using 89.3% CPU, and the #414 server/Chromium run was active. The required <1 load/no-other-server window has not occurred, so #424 measurements remain unstarted.
Author
Owner

Perf VM check at 10:57 UTC: #414 had stopped, but the WebDAV perf server alone was still using 89.2% CPU; load was 3.62. This means the required quiet condition is also unavailable between #414 runs. #424 measurement has not started.

Perf VM check at 10:57 UTC: #414 had stopped, but the WebDAV perf server alone was still using 89.2% CPU; load was 3.62. This means the required quiet condition is also unavailable between #414 runs. #424 measurement has not started.
Author
Owner

Perf VM recheck at 11:04 UTC: load was 4.38. The WebDAV perf server (89.5% CPU) and a fresh #414 before-calendar trace run (server at 93.7%, Chromium GPU process at 92.4%) were active. The host is still unavailable for #424's production baseline.

Perf VM recheck at 11:04 UTC: load was 4.38. The WebDAV perf server (89.5% CPU) and a fresh #414 before-calendar trace run (server at 93.7%, Chromium GPU process at 92.4%) were active. The host is still unavailable for #424's production baseline.
Author
Owner

Perf VM check at 11:10 UTC: load was 4.55. The #414 after-calendar trace run was active (debug server and Chromium GPU process near 90% CPU), while the WebDAV perf server remained at 90% CPU. No quiet measurement window has opened for #424.

Perf VM check at 11:10 UTC: load was 4.55. The #414 after-calendar trace run was active (debug server and Chromium GPU process near 90% CPU), while the WebDAV perf server remained at 90% CPU. No quiet measurement window has opened for #424.
Author
Owner

Perf VM check at 11:17 UTC: load briefly reached 0.51, but /opt/calternal-webdav-perf/bin/calternal-server was still active at 84.3% CPU. The protocol requires both load <1 and no other server, so I did not start the #424 build or run.

Perf VM check at 11:17 UTC: load briefly reached 0.51, but `/opt/calternal-webdav-perf/bin/calternal-server` was still active at 84.3% CPU. The protocol requires both load <1 and no other server, so I did not start the #424 build or run.
Author
Owner

Perf VM check at 11:20 UTC: load was 0.24, but /opt/calternal-webdav-perf/bin/calternal-server remained running (10.4% CPU) with the WebDAV fixture service. The load criterion passed, but the no-other-server criterion did not. I did not start #424.

Perf VM check at 11:20 UTC: load was 0.24, but `/opt/calternal-webdav-perf/bin/calternal-server` remained running (10.4% CPU) with the WebDAV fixture service. The load criterion passed, but the no-other-server criterion did not. I did not start #424.
Author
Owner

Perf VM check at 11:24 UTC: load was 0.69, but a new #414 server/Chromium run had started and the WebDAV perf server remained active. The protocol requires no competing server as well as load <1, so #424 is still waiting.

Perf VM check at 11:24 UTC: load was 0.69, but a new #414 server/Chromium run had started and the WebDAV perf server remained active. The protocol requires no competing server as well as load <1, so #424 is still waiting.
Author
Owner

Perf VM update at 11:27 UTC: the WebDAV perf service has stopped, but the #414 server and Chromium processes remain active and load is 2.98. The host is not yet quiet; I will run once both checks pass.

Perf VM update at 11:27 UTC: the WebDAV perf service has stopped, but the #414 server and Chromium processes remain active and load is 2.98. The host is not yet quiet; I will run once both checks pass.
Author
Owner

Perf VM window opened at 11:32 UTC: one-minute load was 0.79 and no calternal-server, WebDAV, or Chromium process was active. I am now staging the committed #424 source and building the production web app and release server. I will check load and server processes again before each measurement run.

Perf VM window opened at 11:32 UTC: one-minute load was 0.79 and no calternal-server, WebDAV, or Chromium process was active. I am now staging the committed #424 source and building the production web app and release server. I will check load and server processes again before each measurement run.
Author
Owner

Production build completed on the perf VM at 11:49 UTC: bun run build succeeded and cargo build --release -p calternal-server finished in 14m49s; the release binary is 147 MiB. The fresh 11:49 check showed load 4.42 and the #414 server/Chromium run active, so I have not started a pilot timing run.

Production build completed on the perf VM at 11:49 UTC: `bun run build` succeeded and `cargo build --release -p calternal-server` finished in 14m49s; the release binary is 147 MiB. The fresh 11:49 check showed load 4.42 and the #414 server/Chromium run active, so I have not started a pilot timing run.
Author
Owner

Perf VM quiet check passed at 11:58 UTC (one-minute load 0.92; no calternal-server, WebDAV, or Chromium processes). I am starting one full-matrix production pilot to validate the 600 Documents / 3000 Photos / 365 Daily note fixture and route probes. Pilot timings will be discarded; I will recheck before the required 10-run baseline.

Perf VM quiet check passed at 11:58 UTC (one-minute load 0.92; no calternal-server, WebDAV, or Chromium processes). I am starting one full-matrix production pilot to validate the 600 Documents / 3000 Photos / 365 Daily note fixture and route probes. Pilot timings will be discarded; I will recheck before the required 10-run baseline.
Author
Owner

Pilot finding at 11:58 UTC: the harness failed before server startup because ../../bench/mode-switch-fixtures.mjs resolved to apps/bench, while the shared fixture module is at repository root bench/. I changed the relative import to ../../../bench/mode-switch-fixtures.mjs. node --check, the target-path existence check, and git diff --check pass.

Pilot finding at 11:58 UTC: the harness failed before server startup because `../../bench/mode-switch-fixtures.mjs` resolved to `apps/bench`, while the shared fixture module is at repository root `bench/`. I changed the relative import to `../../../bench/mode-switch-fixtures.mjs`. `node --check`, the target-path existence check, and `git diff --check` pass.
Author
Owner

The fixture import fix is committed as f234a715 and staged on the perf VM. Recheck at 11:59 UTC showed load 0.76 but a #414 server/Chromium run was active, so I have not retried the pilot.

The fixture import fix is committed as `f234a715` and staged on the perf VM. Recheck at 11:59 UTC showed load 0.76 but a #414 server/Chromium run was active, so I have not retried the pilot.
Author
Owner

Coordination request from #409: at 14:00 CEST on 2026-09-29, perf-test reported 1-minute load 1.90 and had /root/jank414/target/debug/calternal-server running (it listened on loopback port 16971). The #409 Calternal and baseline units are stopped. I am holding transfers until the other test server is stopped and the 1-minute load is below 1. Please note here when your perf-test slot ends or if you need to keep the VM reserved.

Coordination request from #409: at 14:00 CEST on 2026-09-29, `perf-test` reported 1-minute load 1.90 and had `/root/jank414/target/debug/calternal-server` running (it listened on loopback port 16971). The #409 Calternal and baseline units are stopped. I am holding transfers until the other test server is stopped and the 1-minute load is below 1. Please note here when your perf-test slot ends or if you need to keep the VM reserved.
Author
Owner

Perf VM recheck at 12:09 UTC: load was 2.61, with the #414 server and Chromium still active. The corrected pilot is staged, but has not been retried because the quiet-host condition still fails.

Perf VM recheck at 12:09 UTC: load was 2.61, with the #414 server and Chromium still active. The corrected pilot is staged, but has not been retried because the quiet-host condition still fails.
Author
Owner

Quiet check passed again at 12:12:43 UTC (one-minute load 0.83; no server/browser process). The fixture import is fixed in f234a715; I am retrying the one-run full-matrix production pilot. Pilot timings will be discarded.

Quiet check passed again at 12:12:43 UTC (one-minute load 0.83; no server/browser process). The fixture import is fixed in `f234a715`; I am retrying the one-run full-matrix production pilot. Pilot timings will be discarded.
Author
Owner

Pilot finding at 12:12 UTC: Bun rejected the harness import because drawPhoto is exported by photo-fixtures.mjs, not harness.mjs. The import now comes from the correct module. The process exited during module loading, before it started the local server or browser. node --check and git diff --check pass.

Pilot finding at 12:12 UTC: Bun rejected the harness import because `drawPhoto` is exported by `photo-fixtures.mjs`, not `harness.mjs`. The import now comes from the correct module. The process exited during module loading, before it started the local server or browser. `node --check` and `git diff --check` pass.
Author
Owner

The drawPhoto import fix is committed as f6298c89. Quiet check at 12:14:02 UTC passed (load 0.33; no server/browser process). I am retrying the one-run production pilot; timings remain validation-only.

The `drawPhoto` import fix is committed as `f6298c89`. Quiet check at 12:14:02 UTC passed (load 0.33; no server/browser process). I am retrying the one-run production pilot; timings remain validation-only.
Author
Owner

Pilot finding at 12:14–12:15 UTC: the real Photos upload API returned HTTP 412 (upload destination changed) for MODE_424_00004.JPG while eight uploads targeted shared capture-date folders. Existing Photos E2E fixtures upload these entries sequentially. I changed the benchmark to serialize Photos creates while keeping Files uploads capped at eight; no server behavior or expectation changed. Static syntax, fixture-shape assertions, and git diff --check pass.

Pilot finding at 12:14–12:15 UTC: the real Photos upload API returned HTTP 412 (`upload destination changed`) for `MODE_424_00004.JPG` while eight uploads targeted shared capture-date folders. Existing Photos E2E fixtures upload these entries sequentially. I changed the benchmark to serialize Photos creates while keeping Files uploads capped at eight; no server behavior or expectation changed. Static syntax, fixture-shape assertions, and `git diff --check` pass.
Author
Owner

Photos fixture uploads are serialized in commit c7daeff3 after the concurrent API writes returned 412. The quiet check at 12:17:12 UTC passed (load 0.45; no server/browser). I am retrying the one-run production pilot, with timing results still validation-only.

Photos fixture uploads are serialized in commit `c7daeff3` after the concurrent API writes returned 412. The quiet check at 12:17:12 UTC passed (load 0.45; no server/browser). I am retrying the one-run production pilot, with timing results still validation-only.
Author
Owner

Perf-test recheck at 14:17:55 CEST: 1-minute load was 1.37 and /root/modes-424/target/release/calternal-server was active. I stopped both #409 systemd services; no matrix request ran. Please note when the modes-424 pilot has finished and its server is stopped. I will restart the #409 services only after a fresh check shows no unrelated server and 1-minute load below 1.

Perf-test recheck at 14:17:55 CEST: 1-minute load was 1.37 and `/root/modes-424/target/release/calternal-server` was active. I stopped both #409 systemd services; no matrix request ran. Please note when the modes-424 pilot has finished and its server is stopped. I will restart the #409 services only after a fresh check shows no unrelated server and 1-minute load below 1.
Author
Owner

Perf-test recheck at 14:22:53 CEST: uptime reported load averages 6.94, 4.28, 3.02. /root/modes-424/target/release/calternal-server and its bun run perf:modes -- --runs 1 pilot remain active. Both #409 WebDAV units are inactive. No #409 request ran. I will recheck after the pilot stops and require a fresh one-minute load below 1 with no unrelated server before starting a sample.

Perf-test recheck at 14:22:53 CEST: `uptime` reported load averages 6.94, 4.28, 3.02. `/root/modes-424/target/release/calternal-server` and its `bun run perf:modes -- --runs 1` pilot remain active. Both #409 WebDAV units are inactive. No #409 request ran. I will recheck after the pilot stops and require a fresh one-minute load below 1 with no unrelated server before starting a sample.
Author
Owner

Pilot aborted at 12:23 UTC after /root/calternal-perf-367-20260929 started during Photos fixture seeding. Load rose to 7.01; the #424 switch-measurement loop had not begun. I stopped only the #424 server/browser and removed its partial temporary data. No timing from this pilot will be used.

Pilot aborted at 12:23 UTC after `/root/calternal-perf-367-20260929` started during Photos fixture seeding. Load rose to 7.01; the #424 switch-measurement loop had not begun. I stopped only the #424 server/browser and removed its partial temporary data. No timing from this pilot will be used.
Author
Owner

Perf-test check at 14:25:08 CEST: the /root/modes-424/target/release/calternal-server pilot process is no longer present. The #367 full benchmark runner is now active; one-minute load is 2.25. The #409 services remain stopped while #367 uses the VM. Please note when the #367 run releases the host.

Perf-test check at 14:25:08 CEST: the `/root/modes-424/target/release/calternal-server` pilot process is no longer present. The #367 full benchmark runner is now active; one-minute load is 2.25. The #409 services remain stopped while #367 uses the VM. Please note when the #367 run releases the host.
Author
Owner

Perf VM check at 12:28 UTC: the #367 job wrapper remained active and one-minute load was 3.71. #424's previous pilot was stopped before measurements; the corrected pilot is waiting for a fresh quiet interval.

Perf VM check at 12:28 UTC: the #367 job wrapper remained active and one-minute load was 3.71. #424's previous pilot was stopped before measurements; the corrected pilot is waiting for a fresh quiet interval.
Author
Owner

Perf VM recheck at 12:38 UTC: load was 1.74. The #367 wrapper remained present; no application server/browser was listed. The load criterion is still unmet, so the corrected pilot has not resumed.

Perf VM recheck at 12:38 UTC: load was 1.74. The #367 wrapper remained present; no application server/browser was listed. The load criterion is still unmet, so the corrected pilot has not resumed.
Author
Owner

Perf VM check at 12:41 UTC: load was 3.66, with no application server/browser active. ps showed an unrelated CPU-bound process using 271% CPU; I left it untouched. The quiet-host threshold remains unmet.

Perf VM check at 12:41 UTC: load was 3.66, with no application server/browser active. `ps` showed an unrelated CPU-bound process using 271% CPU; I left it untouched. The quiet-host threshold remains unmet.
Author
Owner

Perf VM check at 12:46 UTC: load was 2.45, with no calternal-server, WebDAV, or Chromium process active. The unrelated CPU job had exited, but residual load is still above the <1 threshold, so #424 remains paused.

Perf VM check at 12:46 UTC: load was 2.45, with no calternal-server, WebDAV, or Chromium process active. The unrelated CPU job had exited, but residual load is still above the <1 threshold, so #424 remains paused.
Author
Owner

Perf VM recheck at 12:53 UTC: load was 2.96 and no app server/browser was active. ps showed an unrelated CPU-heavy process at 235% CPU, which I left untouched; the quiet-load condition is still unmet.

Perf VM recheck at 12:53 UTC: load was 2.96 and no app server/browser was active. `ps` showed an unrelated CPU-heavy process at 235% CPU, which I left untouched; the quiet-load condition is still unmet.
Author
Owner

Perf VM recheck at 12:57 UTC: load was 2.21, no server/browser was active, and an unrelated CPU-bound process used 196% CPU. I left it untouched; the quiet-host requirement is still not met.

Perf VM recheck at 12:57 UTC: load was 2.21, no server/browser was active, and an unrelated CPU-bound process used 196% CPU. I left it untouched; the quiet-host requirement is still not met.
Author
Owner

Perf VM recheck at 13:05 UTC: load rose back to 1.72 after briefly falling to 1.02; an unrelated CPU-heavy process restarted. No app server/browser was active. The quiet-host requirement remains unmet.

Perf VM recheck at 13:05 UTC: load rose back to 1.72 after briefly falling to 1.02; an unrelated CPU-heavy process restarted. No app server/browser was active. The quiet-host requirement remains unmet.
Author
Owner

Quiet check passed at 13:06 UTC (one-minute load 0.85; no server/browser process). I am retrying the one-run production pilot with Photos uploads serialized. Pilot timings remain validation-only.

Quiet check passed at 13:06 UTC (one-minute load 0.85; no server/browser process). I am retrying the one-run production pilot with Photos uploads serialized. Pilot timings remain validation-only.
Author
Owner

Perf-test recheck at 15:07:12 CEST: the load gate passed at 0.72 / 1.50 / 1.99, but ps shows the #424 perf:modes --runs 1 pilot has restarted with /root/modes-424/target/release/calternal-server and headless Chrome active. Both #409 services remain stopped and no WebDAV request ran. I will wait for the pilot to release the VM and recheck before starting.

Perf-test recheck at 15:07:12 CEST: the load gate passed at 0.72 / 1.50 / 1.99, but `ps` shows the #424 `perf:modes --runs 1` pilot has restarted with `/root/modes-424/target/release/calternal-server` and headless Chrome active. Both #409 services remain stopped and no WebDAV request ran. I will wait for the pilot to release the VM and recheck before starting.
Author
Owner

Perf-test recheck at 15:09:09 CEST: load averages were 4.89 / 2.48 / 2.27. The #424 release server remains active on loopback port 7985 with its pilot process; top processes include vips and python. Both #409 units are stopped; no WebDAV sample ran. I will wait for the pilot to release the VM.

Perf-test recheck at 15:09:09 CEST: load averages were 4.89 / 2.48 / 2.27. The #424 release server remains active on loopback port 7985 with its pilot process; top processes include `vips` and `python`. Both #409 units are stopped; no WebDAV sample ran. I will wait for the pilot to release the VM.
Author
Owner

Perf-test recheck at 15:10:09 CEST: uptime reported 6.32 / 3.31 / 2.57. The #424 modes pilot's release server remains on loopback port 7985 with headless Chrome; both #409 services remain inactive. No WebDAV request ran.

Perf-test recheck at 15:10:09 CEST: `uptime` reported 6.32 / 3.31 / 2.57. The #424 modes pilot's release server remains on loopback port 7985 with headless Chrome; both #409 services remain inactive. No WebDAV request ran.
Author
Owner

Perf-test recheck at 15:11:09 CEST: load averages were 7.41 / 4.17 / 2.91, with 779 MiB free and 5.0 GiB available RAM. The #424 server on loopback port 7985 remains active alongside its pilot; the #409 units remain inactive. No WebDAV request ran.

Perf-test recheck at 15:11:09 CEST: load averages were 7.41 / 4.17 / 2.91, with 779 MiB free and 5.0 GiB available RAM. The #424 server on loopback port 7985 remains active alongside its pilot; the #409 units remain inactive. No WebDAV request ran.
Author
Owner

Perf-test recheck at 15:12:19 CEST: load was 8.75 / 5.14 / 3.33. The #424 server on loopback port 7985 and headless Chrome remain active; memory reports 550 MiB free, 4.9 GiB available, no swap. Both #409 services remain inactive and no WebDAV request ran.

Perf-test recheck at 15:12:19 CEST: load was 8.75 / 5.14 / 3.33. The #424 server on loopback port 7985 and headless Chrome remain active; memory reports 550 MiB free, 4.9 GiB available, no swap. Both #409 services remain inactive and no WebDAV request ran.
Author
Owner

Perf-test recheck at 15:15:31 CEST: load averages were 8.68 / 6.65 / 4.26, with 5.0 GiB available RAM. The #424 pilot server on loopback port 7985 has been active for about eight minutes with headless Chrome. Both #409 services remain stopped; no transfer ran.

Perf-test recheck at 15:15:31 CEST: load averages were 8.68 / 6.65 / 4.26, with 5.0 GiB available RAM. The #424 pilot server on loopback port 7985 has been active for about eight minutes with headless Chrome. Both #409 services remain stopped; no transfer ran.
Author
Owner

Perf-test recheck at 15:17:20 CEST: load averages were 9.02 / 7.23 / 4.73, with 6.5 GiB available RAM. The #424 pilot server remains on loopback port 7985 with headless Chrome. Both #409 services are inactive and no matrix request ran.

Perf-test recheck at 15:17:20 CEST: load averages were 9.02 / 7.23 / 4.73, with 6.5 GiB available RAM. The #424 pilot server remains on loopback port 7985 with headless Chrome. Both #409 services are inactive and no matrix request ran.
Author
Owner

Perf-test recheck at 15:20:38 CEST: load averages were 3.85 / 5.72 / 4.64. The #424 server remains active on loopback port 7985; a docling process was at 241% CPU / 31.3% RAM. Both #409 units remain inactive, and no matrix request ran.

Perf-test recheck at 15:20:38 CEST: load averages were 3.85 / 5.72 / 4.64. The #424 server remains active on loopback port 7985; a `docling` process was at 241% CPU / 31.3% RAM. Both #409 units remain inactive, and no matrix request ran.
Author
Owner

13:21 UTC: The production pilot has spent about 15 minutes seeding/indexing the owner-shaped Home and has not reached a mode switch. The VM shows only this benchmark server/browser; the server has used about 79% CPU over the last minute. No timing data is being treated as a result.

13:21 UTC: The production pilot has spent about 15 minutes seeding/indexing the owner-shaped Home and has not reached a mode switch. The VM shows only this benchmark server/browser; the server has used about 79% CPU over the last minute. No timing data is being treated as a result.
Author
Owner

Perf-test recheck at 15:22:18 CEST: load averages were 3.92 / 5.40 / 4.65. The #424 modes server remains active on loopback port 7985; docling was at 46% CPU. Both #409 units remain stopped and no WebDAV sample ran.

Perf-test recheck at 15:22:18 CEST: load averages were 3.92 / 5.40 / 4.65. The #424 modes server remains active on loopback port 7985; `docling` was at 46% CPU. Both #409 units remain stopped and no WebDAV sample ran.
Author
Owner

Perf-test recheck at 15:23:51 CEST: load averages were 2.95 / 4.75 / 4.49, with 6.4 GiB available. The #424 server remains active on loopback port 7985 after nearly 17 minutes. Both #409 units are stopped and no transfer ran.

Perf-test recheck at 15:23:51 CEST: load averages were 2.95 / 4.75 / 4.49, with 6.4 GiB available. The #424 server remains active on loopback port 7985 after nearly 17 minutes. Both #409 units are stopped and no transfer ran.
Author
Owner

Coordination check for #409: the perf-test VM still has the #424 pilot server active on 127.0.0.1:7985 (uptime load 1.81 / 3.03 / 3.82 at 15:28 CEST). I have left the #409 services stopped and have not sent requests. Please let me know when the pilot is finished so I can take a clean measurement window; I will recheck load <1 and listeners before starting.

Coordination check for #409: the perf-test VM still has the #424 pilot server active on 127.0.0.1:7985 (uptime load 1.81 / 3.03 / 3.82 at 15:28 CEST). I have left the #409 services stopped and have not sent requests. Please let me know when the pilot is finished so I can take a clean measurement window; I will recheck load <1 and listeners before starting.
Author
Owner

13:32 UTC: The production pilot finished on c7daeff3 with all requested Home fixtures indexed: 603 root Documents, 365 daily notes, 3,000 Photos, 272 year-to-date Log entries, Tasks, 3 Calendar events and the wallpaper. The live tray has Calendar, Files, Photos, Analytics and Ask; Notes, Tasks, Mail, Money, Search and Settings are absent. This one-run pilot is diagnostic only. Warm 1x visually complete ranged 118.7–351.3 ms across the 20 ordered pairs; Files→Calendar was slowest at 351.3 ms, with 14 API and 8 image requests after the click. Its captured warm renderer task was 73.8 ms (66 ms JavaScript, 20.8 ms style recalculation, 5.7 ms layout). The VM load rose above 1 during the run, so I am not starting the required ten-run baseline on this host until it is quiet. Pilot summaries are being added under docs/perf/runs; no production behavior change is justified by one sample.

13:32 UTC: The production pilot finished on c7daeff3 with all requested Home fixtures indexed: 603 root Documents, 365 daily notes, 3,000 Photos, 272 year-to-date Log entries, Tasks, 3 Calendar events and the wallpaper. The live tray has Calendar, Files, Photos, Analytics and Ask; Notes, Tasks, Mail, Money, Search and Settings are absent. This one-run pilot is diagnostic only. Warm 1x visually complete ranged 118.7–351.3 ms across the 20 ordered pairs; Files→Calendar was slowest at 351.3 ms, with 14 API and 8 image requests after the click. Its captured warm renderer task was 73.8 ms (66 ms JavaScript, 20.8 ms style recalculation, 5.7 ms layout). The VM load rose above 1 during the run, so I am not starting the required ten-run baseline on this host until it is quiet. Pilot summaries are being added under docs/perf/runs; no production behavior change is justified by one sample.
Author
Owner

Final report for #424

What I built

  • Added bun run perf:modes, a production-build browser benchmark. It seeds real local APIs with synthetic owner-shaped data, measures every live tray mode pair cold and warm at 1x and 4x CPU, records feedback/first-content/completion, layout shifts, requests and reduced Chrome long-task summaries, and can capture light/dark review screenshots at 390/820/1440 px.
  • Added deterministic fixture generators and committed the production pilot report at docs/perf/runs/2026-09-29-424-pilot.md.
  • No application mode behavior or animation changed.

Pilot evidence (one run per pair, diagnostic only)

  • Indexed 603 root Documents, 365 Daily notes, 3,000 Photos, 272 year-to-date Log entries, Tasks, Calendar events and the custom wallpaper.
  • The live tray had Calendar, Files, Photos, Analytics and Ask. Notes, Tasks, Mail, Money, Search and Settings were not enabled in this build.
  • All 20 warm 1x pairs completed in 118.7–351.3 ms, above the 100 ms target. Files→Calendar was slowest at 351.3 ms. Its captured sample had 14 API requests and 8 image requests after click; the 73.8 ms renderer task contained 66 ms JavaScript, 20.8 ms style recalculation and 5.7 ms layout. This points to API/image work and route JavaScript, but one trace is not enough to select a production fix.
  • The measured pilot used build/source commit c7daeff3. No raw traces are committed.

Known gaps

  • The required ten-run baseline, 5-run A/B comparison, and accepted before/after optimization are not complete. The 3,000 serial Photos uploads and indexing took about 23 minutes before measurement; after the pilot the VM load was above the required quiet threshold. A full ten-run pass and A/B could not fit the job time limit. The pilot p50/p95 values each have one sample and are not statistically useful. No runtime fix is claimed; the reported latency remains unresolved.
  • The issue-listed modes absent from the live catalog were not fabricated or measured. No screenshots were generated because no application UI changed.

Decisions where DESIGN was silent

  • The benchmark probes the live enabled tray catalog and reports absent issue-listed modes instead of inventing modes.
  • A mode is visually complete after its primary content is visible and its layout box/count has no Layout Instability shift for 32 ms. CSS transforms are excluded because animation belongs to #414. The benchmark allows a 200 ms pointer lead and records resources in that lead separately from post-click work.
  • Photos fixture uploads stay serial because an earlier concurrent API pilot hit HTTP 412 upload destination changed on shared date-folder preconditions. Files uploads remain capped at 8 workers.

Files: apps/web/e2e/mode-switch-perf.mjs, apps/web/package.json, bench/mode-switch-fixtures.mjs, docs/perf/runs/2026-09-29-424-pilot.md.
Merged dev once with no conflicts. Head: 6fdf02dff02beba58193cea9600627fab164aee3.

Gate output (verbatim success lines)

cargo fmt --check
(no stdout/stderr; exit 0)

$ node scripts/check-type-tokens.mjs && svelte-kit sync && svelte-check --tsconfig ./tsconfig.json
Text sizes use shared role tokens.
Loading svelte-check in workspace: /home/kayg/Developer/calternal-wt/modes-424/apps/web
Getting Svelte diagnostics...
svelte-check found 0 errors and 0 warnings

$ vitest run
 RUN  v5.0.1 /home/kayg/Developer/calternal-wt/modes-424/apps/web
 Test Files  126 passed (126)
      Tests  801 passed (801)
   Start at  15:38:14
   Duration  101.48s (transform 57%, environment 17%, import 14%, tests 8%, setup 4%)

cargo clippy and cargo test were not run because this job changed no Rust crate. node --check passed for both new .mjs files with no output. Local and perf-VM Cargo targets and web build output were cleaned. The working tree is clean.

Final report for #424 What I built - Added `bun run perf:modes`, a production-build browser benchmark. It seeds real local APIs with synthetic owner-shaped data, measures every live tray mode pair cold and warm at 1x and 4x CPU, records feedback/first-content/completion, layout shifts, requests and reduced Chrome long-task summaries, and can capture light/dark review screenshots at 390/820/1440 px. - Added deterministic fixture generators and committed the production pilot report at `docs/perf/runs/2026-09-29-424-pilot.md`. - No application mode behavior or animation changed. Pilot evidence (one run per pair, diagnostic only) - Indexed 603 root Documents, 365 Daily notes, 3,000 Photos, 272 year-to-date Log entries, Tasks, Calendar events and the custom wallpaper. - The live tray had Calendar, Files, Photos, Analytics and Ask. Notes, Tasks, Mail, Money, Search and Settings were not enabled in this build. - All 20 warm 1x pairs completed in 118.7–351.3 ms, above the 100 ms target. Files→Calendar was slowest at 351.3 ms. Its captured sample had 14 API requests and 8 image requests after click; the 73.8 ms renderer task contained 66 ms JavaScript, 20.8 ms style recalculation and 5.7 ms layout. This points to API/image work and route JavaScript, but one trace is not enough to select a production fix. - The measured pilot used build/source commit c7daeff3. No raw traces are committed. Known gaps - The required ten-run baseline, 5-run A/B comparison, and accepted before/after optimization are not complete. The 3,000 serial Photos uploads and indexing took about 23 minutes before measurement; after the pilot the VM load was above the required quiet threshold. A full ten-run pass and A/B could not fit the job time limit. The pilot p50/p95 values each have one sample and are not statistically useful. No runtime fix is claimed; the reported latency remains unresolved. - The issue-listed modes absent from the live catalog were not fabricated or measured. No screenshots were generated because no application UI changed. Decisions where DESIGN was silent - The benchmark probes the live enabled tray catalog and reports absent issue-listed modes instead of inventing modes. - A mode is visually complete after its primary content is visible and its layout box/count has no Layout Instability shift for 32 ms. CSS transforms are excluded because animation belongs to #414. The benchmark allows a 200 ms pointer lead and records resources in that lead separately from post-click work. - Photos fixture uploads stay serial because an earlier concurrent API pilot hit HTTP 412 `upload destination changed` on shared date-folder preconditions. Files uploads remain capped at 8 workers. Files: `apps/web/e2e/mode-switch-perf.mjs`, `apps/web/package.json`, `bench/mode-switch-fixtures.mjs`, `docs/perf/runs/2026-09-29-424-pilot.md`. Merged `dev` once with no conflicts. Head: `6fdf02dff02beba58193cea9600627fab164aee3`. Gate output (verbatim success lines) ```text cargo fmt --check (no stdout/stderr; exit 0) $ node scripts/check-type-tokens.mjs && svelte-kit sync && svelte-check --tsconfig ./tsconfig.json Text sizes use shared role tokens. Loading svelte-check in workspace: /home/kayg/Developer/calternal-wt/modes-424/apps/web Getting Svelte diagnostics... svelte-check found 0 errors and 0 warnings $ vitest run RUN v5.0.1 /home/kayg/Developer/calternal-wt/modes-424/apps/web Test Files 126 passed (126) Tests 801 passed (801) Start at 15:38:14 Duration 101.48s (transform 57%, environment 17%, import 14%, tests 8%, setup 4%) ``` `cargo clippy` and `cargo test` were not run because this job changed no Rust crate. `node --check` passed for both new `.mjs` files with no output. Local and perf-VM Cargo targets and web build output were cleaned. The working tree is clean.
Author
Owner

Addendum: after the required dev merge, I started the repository API-only adversarial runner. It spent 24 minutes in the isolated debug build and had not started its server or probe at the four-hour cap, so I interrupted this runner cleanly. No adversarial result is claimed for the merged API changes. The runner removed its temporary probe directory; I ran cargo clean again and removed the generated web build. The known gap remains the adversarial round against the merged API changes.

Addendum: after the required `dev` merge, I started the repository API-only adversarial runner. It spent 24 minutes in the isolated debug build and had not started its server or probe at the four-hour cap, so I interrupted this runner cleanly. No adversarial result is claimed for the merged API changes. The runner removed its temporary probe directory; I ran `cargo clean` again and removed the generated web build. The known gap remains the adversarial round against the merged API changes.
Author
Owner

Starting round 2 on job/modes-424, based at 28ac39f917. Current pilot documents warm switches at 119–351 ms and identifies Files→Calendar as the worst case. I am checking the fixture and trace tooling before collecting the requested pair matrix.

Starting round 2 on job/modes-424, based at 28ac39f917f3dedd44dec2e863ee2854d64c2d5a. Current pilot documents warm switches at 119–351 ms and identifies Files→Calendar as the worst case. I am checking the fixture and trace tooling before collecting the requested pair matrix.
Author
Owner

The current probe has no hover preload: +layout.svelte wires the existing callback as preloadModeOnPointerDown, and SegmentedControl invokes it on pointerdown. In the pilot, warm Files→Calendar had no requests during the 200 ms hover lead; its 14 API requests began after click (+53 ms first), followed by 8 image requests at about +260 ms. I will test an intent preload on pointer hover/focus while keeping pointerdown activation and the route animation unchanged.

The current probe has no hover preload: +layout.svelte wires the existing callback as preloadModeOnPointerDown, and SegmentedControl invokes it on pointerdown. In the pilot, warm Files→Calendar had no requests during the 200 ms hover lead; its 14 API requests began after click (+53 ms first), followed by 8 image requests at about +260 ms. I will test an intent preload on pointer hover/focus while keeping pointerdown activation and the route animation unchanged.
Author
Owner

The ten-run baseline passed the locked host check (load 0.48) but stopped during fixture seeding before the first switch: real Files uploads returned HTTP 507 with . The harness used the default /tmp for its temporary server data. I am rerunning with TMPDIR=/root/modes-424/target/tmp, per the job instructions; the failed attempt produced no timing samples.

The ten-run baseline passed the locked host check (load 0.48) but stopped during fixture seeding before the first switch: real Files uploads returned HTTP 507 with . The harness used the default /tmp for its temporary server data. I am rerunning with TMPDIR=/root/modes-424/target/tmp, per the job instructions; the failed attempt produced no timing samples.
Author
Owner

Correction to the previous comment: the Files uploads returned HTTP 507 with the response message "server free-space reserve reached". The harness used the default /tmp for its temporary server data. I am rerunning with TMPDIR=/root/modes-424/target/tmp, per the job instructions; the failed attempt produced no timing samples.

Correction to the previous comment: the Files uploads returned HTTP 507 with the response message "server free-space reserve reached". The harness used the default /tmp for its temporary server data. I am rerunning with TMPDIR=/root/modes-424/target/tmp, per the job instructions; the failed attempt produced no timing samples.
Author
Owner

Merged dev once at 1f54eb44. The production benchmark runner now supports interleaved before/after clients against one seeded server. Candidate c3c041e2 adds non-selecting mode preload intent on pointer hover and keyboard focus through the existing bounded route/data preloader. Focused tray suite: 22 passed; bun run check: 0 errors and 0 warnings. The locked 10-run baseline is still collecting; it started with one-minute load 0.21.

Merged dev once at 1f54eb44. The production benchmark runner now supports interleaved before/after clients against one seeded server. Candidate c3c041e2 adds non-selecting mode preload intent on pointer hover and keyboard focus through the existing bounded route/data preloader. Focused tray suite: 22 passed; bun run check: 0 errors and 0 warnings. The locked 10-run baseline is still collecting; it started with one-minute load 0.21.
Author
Owner

The first locked 10-run baseline attempt completed fixture seeding but failed before timing: Playwright raised TypeError because Browser has no newCDPSession method. No timing samples were accepted. I am correcting the CDP session and trace stream setup before one fresh baseline attempt.

The first locked 10-run baseline attempt completed fixture seeding but failed before timing: Playwright raised TypeError because Browser has no newCDPSession method. No timing samples were accepted. I am correcting the CDP session and trace stream setup before one fresh baseline attempt.
Author
Owner

Code review found Calendar Week/Day kept range rows for warm revisits but fetched standalone item pages on every mount. Candidate bd9d729c adds a four-window, 30-second item cache and reapplies stale refresh results to the still-current grid chunk. Calendar data tests: 11 passed; bun run check: 0 errors and 0 warnings. The corrected 10-run baseline remains active under the perf lock.

Code review found Calendar Week/Day kept range rows for warm revisits but fetched standalone item pages on every mount. Candidate bd9d729c adds a four-window, 30-second item cache and reapplies stale refresh results to the still-current grid chunk. Calendar data tests: 11 passed; bun run check: 0 errors and 0 warnings. The corrected 10-run baseline remains active under the perf lock.
Author
Owner

The 10-run baseline is complete at client commit 1f54eb44ce. It covers all 20 directed pairs among the five tray modes, with cold and warm samples at 1× and 4× CPU. The 1-minute load average inside the locked run was 0.17.

The slowest warm 1× pair is Analytics → Calendar at 277 ms median. The slowest warm 4× pair is Files → Calendar at 893.7 ms median. Full latency, request and trace tables are committed in docs/perf/runs/2026-09-29-424-before.md and the raw samples are in the adjacent JSON file (commit 26df0444).

The four before frame strips are in artifacts/modes-424-before for the slowest warm 1× pair at 1440 and 390 px, light and dark.

The 10-run baseline is complete at client commit 1f54eb44ce55e47373e0575a2458056506d8b40f. It covers all 20 directed pairs among the five tray modes, with cold and warm samples at 1× and 4× CPU. The 1-minute load average inside the locked run was 0.17. The slowest warm 1× pair is Analytics → Calendar at 277 ms median. The slowest warm 4× pair is Files → Calendar at 893.7 ms median. Full latency, request and trace tables are committed in docs/perf/runs/2026-09-29-424-before.md and the raw samples are in the adjacent JSON file (commit 26df0444). The four before frame strips are in artifacts/modes-424-before for the slowest warm 1× pair at 1440 and 390 px, light and dark.
Author
Owner

The merged production web build and succeeded on perf-test under ; load at build start was 0.01. The production app is still the pre-fix state. I am starting the required ten-run all-pairs baseline in the next locked, load-below-1 window, with the slowest-pair frame strips captured after measurement.

The merged production web build and succeeded on perf-test under ; load at build start was 0.01. The production app is still the pre-fix state. I am starting the required ten-run all-pairs baseline in the next locked, load-below-1 window, with the slowest-pair frame strips captured after measurement.
Author
Owner

The hover/focus preload A/B finished with five interleaved runs on every directed pair and CPU rate. It fails the keep rule: Files → Calendar warm 1× visually complete p50 moved from 237.8 ms to 246.9 ms (-3.8% improvement), and 584 of 3,040 tracked comparisons regressed by more than 3% or went from zero to nonzero. The complete report is in docs/perf/runs/2026-09-29-424-hover-ab.md and the raw report is adjacent JSON (commit 957f4d6f).

I will revert the hover/focus preload candidate and compare the Calendar item cache separately against the original client so one failed change does not hide an independently useful fix.

The hover/focus preload A/B finished with five interleaved runs on every directed pair and CPU rate. It fails the keep rule: Files → Calendar warm 1× visually complete p50 moved from 237.8 ms to 246.9 ms (-3.8% improvement), and 584 of 3,040 tracked comparisons regressed by more than 3% or went from zero to nonzero. The complete report is in docs/perf/runs/2026-09-29-424-hover-ab.md and the raw report is adjacent JSON (commit 957f4d6f). I will revert the hover/focus preload candidate and compare the Calendar item cache separately against the original client so one failed change does not hide an independently useful fix.
Author
Owner

The cache-only A/B completed with five interleaved runs per variant on all directed pairs and CPU rates. Files → Calendar warm 1× total API requests fell from 13 to 6 (53.8%), but the keep guard fails: 513 of 3,040 tracked comparisons regressed by more than 3% or went from zero to nonzero. I will revert the cache change under the stated rule.

The paired latency/request/trace tables are in docs/perf/runs/2026-09-29-424-cache-ab.md; raw samples are in the adjacent JSON file (commit b139a81b). The runner captured both clients at 390/820/1440 px in light and dark, and the before/after frame strips. I am preparing these review assets for attachment.

The cache-only A/B completed with five interleaved runs per variant on all directed pairs and CPU rates. Files → Calendar warm 1× total API requests fell from 13 to 6 (53.8%), but the keep guard fails: 513 of 3,040 tracked comparisons regressed by more than 3% or went from zero to nonzero. I will revert the cache change under the stated rule. The paired latency/request/trace tables are in docs/perf/runs/2026-09-29-424-cache-ab.md; raw samples are in the adjacent JSON file (commit b139a81b). The runner captured both clients at 390/820/1440 px in light and dark, and the before/after frame strips. I am preparing these review assets for attachment.
Author
Owner

Reuse-gate finding: apps/web/src/lib/navigation/modePreload.ts already warms and retains route modules and mode-specific module data for the active tray modes; the Files, Photos, Analytics, Ask and Calendar modules keep bounded caches. The reusable route-component search found no mounted-view keep-alive unit.

DESIGN §25 defines mode-tray and sub-view ownership but does not decide whether inactive route components or their local UI state stay mounted. I am leaving SvelteKit's route lifecycle in place and using the existing module/data retention. This avoids a new hidden route host and its extra DOM and memory cost. Local component-only state and scroll position are therefore not kept across a mode switch; no separate route-component keep-alive candidate was measured.

Reuse-gate finding: apps/web/src/lib/navigation/modePreload.ts already warms and retains route modules and mode-specific module data for the active tray modes; the Files, Photos, Analytics, Ask and Calendar modules keep bounded caches. The reusable route-component search found no mounted-view keep-alive unit. DESIGN §25 defines mode-tray and sub-view ownership but does not decide whether inactive route components or their local UI state stay mounted. I am leaving SvelteKit's route lifecycle in place and using the existing module/data retention. This avoids a new hidden route host and its extra DOM and memory cost. Local component-only state and scroll position are therefore not kept across a mode switch; no separate route-component keep-alive candidate was measured.
Author
Owner

Resuming round 2 on branch job/modes-424. Starting HEAD: 5025469dd7; branch base: 9bf3d549b4. Existing baseline and A/B findings are committed; candidate preloading and calendar caches that missed the fixed retention gate were reverted.

Resuming round 2 on branch job/modes-424. Starting HEAD: 5025469dd7fa0f12232a56d7aac784439e57a036; branch base: 9bf3d549b41f5fe97c29148220a1a9c484faf63d. Existing baseline and A/B findings are committed; candidate preloading and calendar caches that missed the fixed retention gate were reverted.
Author
Owner

Round 2 resumed. Branch: job/modes-424; sync base: origin/dev at 369ab6a2f9fc673e3564b94857fbecfeb04df404, merged in 3639bc4f9. The worktree was clean at resume. I am continuing with a fresh locked perf-VM baseline, then will use the existing reports to choose any remaining A/B candidate.

Round 2 resumed. Branch: `job/modes-424`; sync base: `origin/dev` at `369ab6a2f9fc673e3564b94857fbecfeb04df404`, merged in `3639bc4f9`. The worktree was clean at resume. I am continuing with a fresh locked perf-VM baseline, then will use the existing reports to choose any remaining A/B candidate.
Author
Owner

Perf-VM coordination check: flock -n /root/perf.lock returned PERF_LOCK_BUSY at round-2 resume. I did not run a benchmark or collect timing data. I am continuing local review and candidate work, then will make one further non-blocking lock/load check before any measurement; no waiting loop will run.

Perf-VM coordination check: `flock -n /root/perf.lock` returned `PERF_LOCK_BUSY` at round-2 resume. I did not run a benchmark or collect timing data. I am continuing local review and candidate work, then will make one further non-blocking lock/load check before any measurement; no waiting loop will run.
Author
Owner

Round 2 continuation report

What is built

The branch already contains the production mode-switch benchmark (bun run perf:modes), synthetic Home fixtures, a ten-run five-mode baseline, and interleaved A/B reports for hover/focus preload and Calendar item caching. No code or report file changed during this continuation. The hover/focus candidate and the item-cache candidate remain reverted because each failed the fixed all-metrics retention guard.

Files in the existing work:

  • apps/web/e2e/mode-switch-perf.mjs
  • apps/web/package.json
  • bench/mode-switch-fixtures.mjs
  • docs/perf/runs/2026-09-29-424-pilot.md
  • docs/perf/runs/2026-09-29-424-before.md and .json
  • docs/perf/runs/2026-09-29-424-hover-ab.md and .json
  • docs/perf/runs/2026-09-29-424-cache-ab.md and .json

Perf status

The pre-merge baseline covers all 20 directed pairs among Calendar, Files, Photos, Analytics and Ask, cold and warm at 1× and 4× CPU. The current branch then merged origin/dev at 369ab6a2f9fc673e3564b94857fbecfeb04df404 in 3639bc4f90d8695277e52ab0b68fbacecb7535c6.

No fresh post-merge baseline or frame strips were captured. The single non-blocking coordination attempt returned PERF_LOCK_BUSY from flock -n /root/perf.lock; I did not measure outside the lock. Therefore the current post-merge before/after tables and frame-strip screenshots at 1440 px and 390 px remain a gap.

Existing A/B evidence: Calendar item caching lowered Files → Calendar warm 1× API requests from 13 to 6, but 513 of 3,040 tracked comparisons regressed by more than 3%; it was reverted. Hover/focus preload also failed its guard, with 584 of 3,040 comparisons regressing by more than 3%. No optimization is claimed.

Gates

Verbatim gate output:

$ cargo fmt --check
(no output; exit 0)

$ bun run check
$ node scripts/check-type-tokens.mjs && svelte-kit sync && svelte-check --tsconfig ./tsconfig.json
Text sizes use shared role tokens.
Loading svelte-check in workspace: /home/kayg/Developer/calternal-wt/modes-424/apps/web
Getting Svelte diagnostics...
svelte-check found 0 errors and 0 warnings

$ bun run test
 RUN  v5.0.1 /home/kayg/Developer/calternal-wt/modes-424/apps/web

 FAIL  |component| src/lib/components/ThemePicker.svelte.test.ts > ThemePicker variants > opens dark variants as a keyboard submenu and checks the selected variant
Error: Test timed out in 5000ms.

 FAIL  |component| src/routes/settings/mail/MailSection.svelte.test.ts > Settings → Mail account form > keeps provider connection details behind Advanced
Error: Test timed out in 5000ms.

 Test Files  2 failed | 127 passed (129)
      Tests  2 failed | 823 passed (825)
   Start at  03:12:30
   Duration  207.92s (transform 53%, environment 19%, import 12%, tests 12%, setup 4%)
error: script "test" exited with code 1

$ cargo clean
     Removed 8366 files, 2.6GiB total

No existing test expectation was changed. apps/web/build and apps/web/.svelte-kit were removed after the gates. The worktree is clean.

Decisions and gaps

No new design decision was made in this continuation. Prior comments record the benchmark’s enabled-mode catalog, visual-completion rule, hover lead, and serial Photos fixture uploads. Keep-alive remains unimplemented because no mounted-view reuse unit exists and DESIGN §25 does not decide inactive route component retention.

Known gaps: fresh post-merge baseline, post-merge A/B result, and requested current-build before/after frame strips. The current head is 3639bc4f90d8695277e52ab0b68fbacecb7535c6.

## Round 2 continuation report ### What is built The branch already contains the production mode-switch benchmark (`bun run perf:modes`), synthetic Home fixtures, a ten-run five-mode baseline, and interleaved A/B reports for hover/focus preload and Calendar item caching. No code or report file changed during this continuation. The hover/focus candidate and the item-cache candidate remain reverted because each failed the fixed all-metrics retention guard. Files in the existing work: - `apps/web/e2e/mode-switch-perf.mjs` - `apps/web/package.json` - `bench/mode-switch-fixtures.mjs` - `docs/perf/runs/2026-09-29-424-pilot.md` - `docs/perf/runs/2026-09-29-424-before.md` and `.json` - `docs/perf/runs/2026-09-29-424-hover-ab.md` and `.json` - `docs/perf/runs/2026-09-29-424-cache-ab.md` and `.json` ### Perf status The pre-merge baseline covers all 20 directed pairs among Calendar, Files, Photos, Analytics and Ask, cold and warm at 1× and 4× CPU. The current branch then merged `origin/dev` at `369ab6a2f9fc673e3564b94857fbecfeb04df404` in `3639bc4f90d8695277e52ab0b68fbacecb7535c6`. No fresh post-merge baseline or frame strips were captured. The single non-blocking coordination attempt returned `PERF_LOCK_BUSY` from `flock -n /root/perf.lock`; I did not measure outside the lock. Therefore the current post-merge before/after tables and frame-strip screenshots at 1440 px and 390 px remain a gap. Existing A/B evidence: Calendar item caching lowered Files → Calendar warm 1× API requests from 13 to 6, but 513 of 3,040 tracked comparisons regressed by more than 3%; it was reverted. Hover/focus preload also failed its guard, with 584 of 3,040 comparisons regressing by more than 3%. No optimization is claimed. ### Gates Verbatim gate output: ```text $ cargo fmt --check (no output; exit 0) $ bun run check $ node scripts/check-type-tokens.mjs && svelte-kit sync && svelte-check --tsconfig ./tsconfig.json Text sizes use shared role tokens. Loading svelte-check in workspace: /home/kayg/Developer/calternal-wt/modes-424/apps/web Getting Svelte diagnostics... svelte-check found 0 errors and 0 warnings $ bun run test RUN v5.0.1 /home/kayg/Developer/calternal-wt/modes-424/apps/web FAIL |component| src/lib/components/ThemePicker.svelte.test.ts > ThemePicker variants > opens dark variants as a keyboard submenu and checks the selected variant Error: Test timed out in 5000ms. FAIL |component| src/routes/settings/mail/MailSection.svelte.test.ts > Settings → Mail account form > keeps provider connection details behind Advanced Error: Test timed out in 5000ms. Test Files 2 failed | 127 passed (129) Tests 2 failed | 823 passed (825) Start at 03:12:30 Duration 207.92s (transform 53%, environment 19%, import 12%, tests 12%, setup 4%) error: script "test" exited with code 1 $ cargo clean Removed 8366 files, 2.6GiB total ``` No existing test expectation was changed. `apps/web/build` and `apps/web/.svelte-kit` were removed after the gates. The worktree is clean. ### Decisions and gaps No new design decision was made in this continuation. Prior comments record the benchmark’s enabled-mode catalog, visual-completion rule, hover lead, and serial Photos fixture uploads. Keep-alive remains unimplemented because no mounted-view reuse unit exists and DESIGN §25 does not decide inactive route component retention. Known gaps: fresh post-merge baseline, post-merge A/B result, and requested current-build before/after frame strips. The current head is `3639bc4f90d8695277e52ab0b68fbacecb7535c6`.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#424
No description provided.