Blaze test: holding Down/Up through any list shows each item fully loaded #641

Open
opened 2026-10-01 16:22:29 +00:00 by kayg · 100 comments
Owner

Owner test (2026-10-01), now the main performance test. Put the cursor on the first item of a list and hold Down (or Up) at full keyboard repeat speed. The selection moves through the list at once. At every step, the content pane must show the selected item fully loaded. No skeleton, spinner, blank pane, half-rendered card or "Loading" text may appear while you move fast. When you release the key, the content of the last item is already there.

The owner calls this the "blaze test". It is the measure of whether calternal performance is good.

Where it applies

Surface List Content that must be complete at each step
Settings left sidebar sections the full section page (all cards, values, toggles)
Mail message list (middle column) the message (headers + body + images that are cached)
Notes note list / Navigator the note card with its text
Files rows (Space opens the preview) the inspector / preview of the selected file
Photos grid → viewer Left/Right the full-size photo (not a blur placeholder after the first frame)
Calendar previous/next day/week the grid with all events
Money accounts sidebar the account register
Search results with preview the preview
Tabs Cmd/Ctrl+1…9 spam each Tab fully painted
Admin, Jobs, Tasks list → detail the detail

Approach rules (owner priorities: performance first, never at the cost of finesse)

  • Render from data that is already in memory. Load the list's neighbours before the selection reaches them: prefetch in the direction of travel, keep a bounded in-memory cache, and preload route chunks of the sub-views when the parent opens (on idle).
  • Never show a loading state for a step that the cache can serve. Do not debounce the content to "the item where the user stopped": every step paints the selected item.
  • If an item truly is not available (cold, offline), keep the previous content dimmed for at most one frame budget, never a skeleton flash.
  • Keep main-thread work per step under one frame (16.7 ms at 60 Hz). No long task over 50 ms during the run.
  • Memory stays bounded (LRU). Measure RSS before and after a 500-step run.

Measurement (bench/blaze.mjs, shared by every surface)

  • Playwright + CDP Input.dispatchKeyEvent, key repeat at 33 ms (common default) and 15 ms (fastest macOS repeat). 200+ steps. Chromium and WebKit, 1440 px and 390 px where the surface exists on phones.
  • Every animation frame, record: selected item ID, ID shown in the content pane, and whether the pane is complete (no [aria-busy=true], no skeleton/spinner class, expected content element present).
  • Metrics: mismatch frames (content ≠ selection), incomplete frames, steps that were never painted complete, long tasks, dropped frames, JS heap and RSS.
  • Target: 0 incomplete frames on a warm run; content matches the selection in the same frame as the key, or the next frame; 0 long tasks over 50 ms. Run on the perf VM (root@10.69.69.63) with the HDD emulation from #549 (bench/hdd-emu.sh) and flock /root/perf.lock, cold and warm, with heavy data (the large fixtures).
  • A real-keyboard confirmation on the macOS VM comes later: write the exact steps for the owner, do not set up the VM.

Sub-issues own the work per surface. This issue owns the harness and the targets.

**Owner test (2026-10-01), now the main performance test.** Put the cursor on the first item of a list and hold Down (or Up) at full keyboard repeat speed. The selection moves through the list at once. **At every step, the content pane must show the selected item fully loaded.** No skeleton, spinner, blank pane, half-rendered card or "Loading" text may appear while you move fast. When you release the key, the content of the last item is already there. The owner calls this the "blaze test". It is the measure of whether calternal performance is good. ### Where it applies | Surface | List | Content that must be complete at each step | |---|---|---| | Settings | left sidebar sections | the full section page (all cards, values, toggles) | | Mail | message list (middle column) | the message (headers + body + images that are cached) | | Notes | note list / Navigator | the note card with its text | | Files | rows (Space opens the preview) | the inspector / preview of the selected file | | Photos | grid → viewer Left/Right | the full-size photo (not a blur placeholder after the first frame) | | Calendar | previous/next day/week | the grid with all events | | Money | accounts sidebar | the account register | | Search | results with preview | the preview | | Tabs | Cmd/Ctrl+1…9 spam | each Tab fully painted | | Admin, Jobs, Tasks | list → detail | the detail | ### Approach rules (owner priorities: performance first, never at the cost of finesse) - Render from data that is already in memory. Load the list's neighbours **before** the selection reaches them: prefetch in the direction of travel, keep a bounded in-memory cache, and preload route chunks of the sub-views when the parent opens (on idle). - Never show a loading state for a step that the cache can serve. Do not debounce the content to "the item where the user stopped": every step paints the selected item. - If an item truly is not available (cold, offline), keep the previous content dimmed for at most one frame budget, never a skeleton flash. - Keep main-thread work per step under one frame (16.7 ms at 60 Hz). No long task over 50 ms during the run. - Memory stays bounded (LRU). Measure RSS before and after a 500-step run. ### Measurement (bench/blaze.mjs, shared by every surface) - Playwright + CDP `Input.dispatchKeyEvent`, key repeat at **33 ms** (common default) and **15 ms** (fastest macOS repeat). 200+ steps. Chromium and WebKit, 1440 px and 390 px where the surface exists on phones. - Every animation frame, record: selected item ID, ID shown in the content pane, and whether the pane is complete (no `[aria-busy=true]`, no skeleton/spinner class, expected content element present). - Metrics: **mismatch frames** (content ≠ selection), **incomplete frames**, steps that were never painted complete, long tasks, dropped frames, JS heap and RSS. - **Target:** 0 incomplete frames on a warm run; content matches the selection in the same frame as the key, or the next frame; 0 long tasks over 50 ms. Run on the perf VM (`root@10.69.69.63`) with the HDD emulation from #549 (`bench/hdd-emu.sh`) and `flock /root/perf.lock`, cold and warm, with heavy data (the large fixtures). - A real-keyboard confirmation on the macOS VM comes later: write the exact steps for the owner, do not set up the VM. Sub-issues own the work per surface. This issue owns the harness and the targets.
Author
Owner

Starting #641/#642 on job/blaze-settings at base 687ff70313. I will commit the shared blaze harness first, then implement and measure the Settings path.

Starting #641/#642 on job/blaze-settings at base 687ff703136e71e89f8dfba139e93cd0788b25c1. I will commit the shared blaze harness first, then implement and measure the Settings path.
Author
Owner

The shared harness is committed as 134158825e44c1b4c98140efefaf21a33c62c5fb (bench: blaze harness for surface selection). node --check bench/blaze.mjs passed and bun run build completed against the production web app. The worktree does not contain bench/hdd-emu.sh (test -x bench/hdd-emu.sh returned false); I am checking the perf VM for the #549 helper before measuring.

The shared harness is committed as 134158825e44c1b4c98140efefaf21a33c62c5fb (`bench: blaze harness for surface selection`). `node --check bench/blaze.mjs` passed and `bun run build` completed against the production web app. The worktree does not contain `bench/hdd-emu.sh` (`test -x bench/hdd-emu.sh` returned false); I am checking the perf VM for the #549 helper before measuring.
Author
Owner

The runner and report are committed, but the required perf VM cannot currently be reached from this worktree: ssh -o BatchMode=yes -o ConnectTimeout=5 root@10.69.69.63 ... returned No route to host. The worktree also has no bench/hdd-emu.sh. I will use the allowed local fallback and label its numbers local; this does not claim HDD-emulated perf VM coverage.

The runner and report are committed, but the required perf VM cannot currently be reached from this worktree: `ssh -o BatchMode=yes -o ConnectTimeout=5 root@10.69.69.63 ...` returned `No route to host`. The worktree also has no `bench/hdd-emu.sh`. I will use the allowed local fallback and label its numbers `local`; this does not claim HDD-emulated perf VM coverage.
Author
Owner

Status update for #641: I fetched origin/job/blaze-settings and checked git log --oneline job/blaze-settings; the branch still ends at 687ff703 and does not contain the expected bench: blaze harness commit. I am continuing the Mail work and will recheck this ref before the profile run.

Status update for #641: I fetched `origin/job/blaze-settings` and checked `git log --oneline job/blaze-settings`; the branch still ends at `687ff703` and does not contain the expected `bench: blaze harness` commit. I am continuing the Mail work and will recheck this ref before the profile run.
Author
Owner

Starting job/blaze-surfaces from origin/dev at 687ff70313. I am tracing and improving the Files, Photos, Money, and Tab switching blaze paths, using the shared bench/blaze.mjs harness when job/blaze-settings lands it.

Starting job/blaze-surfaces from origin/dev at 687ff703136e71e89f8dfba139e93cd0788b25c1. I am tracing and improving the Files, Photos, Money, and Tab switching blaze paths, using the shared bench/blaze.mjs harness when job/blaze-settings lands it.
Author
Owner

Baseline finding for #641. The production build is running locally because bench/hdd-emu.sh is absent from this worktree and SSH to root@10.69.69.63 returns No route to host. No perf-VM result is claimed.

Exact command:

PERF_LOCATION=local CALTERNAL_E2E_ASSET_OVERRIDE=1 CALTERNAL_SERVER_BIN="$CARGO_TARGET_DIR/debug/calternal-server" TMPDIR="$PWD/target/tmp" bun bench/blaze.mjs --surface settings --json artifacts/blaze/settings-before-641.json

At desktop width, Settings ignores ArrowDown/ArrowUp. Each 200-step direction run therefore left selection at Account; every step was never painted complete. The list is not present in the phone drill-in layout, so those runs were recorded as skipped.

Browser Width Repeat Phase Mismatch frames Incomplete frames Steps never complete Longest task p95 frame
Chromium 1440 33 ms cold / warm 0 0 200 445 ms 16.7 / 16.8 ms
Chromium 1440 15 ms cold / warm 0 0 200 644 ms 16.8 / 16.8 ms
WebKit 1440 33 ms cold / warm 0 0 200 n/a 23 / 23 ms
WebKit 1440 15 ms cold / warm 0 0 200 n/a 197 / 178 ms

Full frame-level evidence: artifacts/blaze/settings-before-641.json (local, not committed).

Baseline finding for #641. The production build is running locally because `bench/hdd-emu.sh` is absent from this worktree and SSH to `root@10.69.69.63` returns `No route to host`. No perf-VM result is claimed. Exact command: ```sh PERF_LOCATION=local CALTERNAL_E2E_ASSET_OVERRIDE=1 CALTERNAL_SERVER_BIN="$CARGO_TARGET_DIR/debug/calternal-server" TMPDIR="$PWD/target/tmp" bun bench/blaze.mjs --surface settings --json artifacts/blaze/settings-before-641.json ``` At desktop width, Settings ignores ArrowDown/ArrowUp. Each 200-step direction run therefore left selection at Account; every step was never painted complete. The list is not present in the phone drill-in layout, so those runs were recorded as skipped. | Browser | Width | Repeat | Phase | Mismatch frames | Incomplete frames | Steps never complete | Longest task | p95 frame | | --- | ---: | ---: | --- | ---: | ---: | ---: | ---: | ---: | | Chromium | 1440 | 33 ms | cold / warm | 0 | 0 | 200 | 445 ms | 16.7 / 16.8 ms | | Chromium | 1440 | 15 ms | cold / warm | 0 | 0 | 200 | 644 ms | 16.8 / 16.8 ms | | WebKit | 1440 | 33 ms | cold / warm | 0 | 0 | 200 | n/a | 23 / 23 ms | | WebKit | 1440 | 15 ms | cold / warm | 0 | 0 | 200 | n/a | 197 / 178 ms | Full frame-level evidence: `artifacts/blaze/settings-before-641.json` (local, not committed).
Author
Owner

Trace finding: packages/ui/src/components/viewer/QuickLook.svelte mounted only the selected image and did not decode adjacent image sources. Files selection also had no neighbour thumbnail/full-preview warmup. I am adding one bounded decoded-image LRU for those existing paths. The first focused unit-test attempt is currently blocked because vitest is not installed in this worktree (bun run test -- ../packages/ui/src/imageCache.test.ts: command not found).

Trace finding: `packages/ui/src/components/viewer/QuickLook.svelte` mounted only the selected image and did not decode adjacent image sources. Files selection also had no neighbour thumbnail/full-preview warmup. I am adding one bounded decoded-image LRU for those existing paths. The first focused unit-test attempt is currently blocked because `vitest` is not installed in this worktree (`bun run test -- ../packages/ui/src/imageCache.test.ts`: command not found).
Author
Owner

Progress on job/blaze-surfaces (base 687ff703136e71e89f8dfba139e93cd0788b25c1): committed 319e353f158a65c5bc643ce17cb9c9e81e95ce0d (perf: predecode Files and Photos image neighbours). The shared viewer now waits for img.decode() before declaring a full image ready, and Files and Quick Look warm a direction-aware bounded decoded-image cache. bun run check completed with svelte-check found 0 errors and 0 warnings; focused cache tests passed (3 tests). Next I am adding the Files/Photos/Money/tab configs to the shared blaze sampler and surface coverage.

Progress on `job/blaze-surfaces` (base `687ff703136e71e89f8dfba139e93cd0788b25c1`): committed `319e353f158a65c5bc643ce17cb9c9e81e95ce0d` (`perf: predecode Files and Photos image neighbours`). The shared viewer now waits for `img.decode()` before declaring a full image ready, and Files and Quick Look warm a direction-aware bounded decoded-image cache. `bun run check` completed with `svelte-check found 0 errors and 0 warnings`; focused cache tests passed (3 tests). Next I am adding the Files/Photos/Money/tab configs to the shared blaze sampler and surface coverage.
Author
Owner

Finding while testing #641: SvelteKit's shallow pushState changes the address bar but does not update $app/state.page.url. Reproduction in the production build: two ArrowDown presses changed the address to /settings/maintenance while every Settings row lost aria-current="page" and the Account content remained selected. I removed that approach and kept Settings' documented same-route goto behavior so the URL, selected row and section body stay synchronized. The repeat-key e2e is rerunning against the corrected navigation.

Finding while testing #641: SvelteKit's shallow `pushState` changes the address bar but does not update `$app/state.page.url`. Reproduction in the production build: two ArrowDown presses changed the address to `/settings/maintenance` while every Settings row lost `aria-current="page"` and the Account content remained selected. I removed that approach and kept Settings' documented same-route `goto` behavior so the URL, selected row and section body stay synchronized. The repeat-key e2e is rerunning against the corrected navigation.
Author
Owner

Progress on #641 — commits 2061200b5 and 69d192b28 are on job/blaze-surfaces.

The shared Quick Look path now shares bounded decoded-image, text-response and PDF-document caches with directional neighbour prefetch. The Files/Photos/Money/Tabs configs use the same frame sampler; it records key-to-paint latency, long tasks, server CPU/RSS and load average, and asserts that every input has a complete matching view. The focused web suite passed: 5 files, 23 tests. bun run check reported 0 errors and 0 warnings. I am moving to production-build runs now.

Progress on #641 — commits `2061200b5` and `69d192b28` are on `job/blaze-surfaces`. The shared Quick Look path now shares bounded decoded-image, text-response and PDF-document caches with directional neighbour prefetch. The Files/Photos/Money/Tabs configs use the same frame sampler; it records key-to-paint latency, long tasks, server CPU/RSS and load average, and asserts that every input has a complete matching view. The focused web suite passed: 5 files, 23 tests. `bun run check` reported 0 errors and 0 warnings. I am moving to production-build runs now.
Author
Owner

Production-build local follow-up after the Settings changes. The perf VM is still unreachable (No route to host), and bench/hdd-emu.sh is absent, so these are host-contended local numbers.

Exact command:

export TMPDIR="$PWD/target/tmp" CARGO_PROFILE_DEV_DEBUG=line-tables-only CARGO_INCREMENTAL=0 CARGO_BUILD_JOBS=4 CALTERNAL_E2E_ASSET_OVERRIDE=1 CALTERNAL_SERVER_BIN="$CARGO_TARGET_DIR/debug/calternal-server"
node apps/web/e2e/settings-blaze-641.mjs
pass browser viewport repeat mismatched frames incomplete frames steps without a distinct paint longest task p95 frame
cold after Chromium 1440×900 33 ms 0 8 31 198 ms 16.8 ms
warm after Chromium 1440×900 33 ms 0 0 24 132 ms 16.8 ms

Every sampled warm frame is complete and matches the selected section. The local traversal still has 24 input steps with no distinct frame before the next input; this is visible in the attached JSON and remains open for the required perf-VM matrix. The benchmark attributes a paint to the interval before the next key event, so later revisits cannot count for an earlier step.

Production-build local follow-up after the Settings changes. The perf VM is still unreachable (`No route to host`), and `bench/hdd-emu.sh` is absent, so these are host-contended local numbers. Exact command: ```sh export TMPDIR="$PWD/target/tmp" CARGO_PROFILE_DEV_DEBUG=line-tables-only CARGO_INCREMENTAL=0 CARGO_BUILD_JOBS=4 CALTERNAL_E2E_ASSET_OVERRIDE=1 CALTERNAL_SERVER_BIN="$CARGO_TARGET_DIR/debug/calternal-server" node apps/web/e2e/settings-blaze-641.mjs ``` | pass | browser | viewport | repeat | mismatched frames | incomplete frames | steps without a distinct paint | longest task | p95 frame | |---|---|---:|---:|---:|---:|---:|---:|---:| | cold after | Chromium | 1440×900 | 33 ms | 0 | 8 | 31 | 198 ms | 16.8 ms | | warm after | Chromium | 1440×900 | 33 ms | 0 | 0 | 24 | 132 ms | 16.8 ms | Every sampled warm frame is complete and matches the selected section. The local traversal still has 24 input steps with no distinct frame before the next input; this is visible in the attached JSON and remains open for the required perf-VM matrix. The benchmark attributes a paint to the interval before the next key event, so later revisits cannot count for an earlier step.
Author
Owner

Perf host access finding: the locked probe ssh root@10.69.69.63 returned No route to host, so the requested perf VM run is unavailable from this worker. I will run the profiles locally against the production build and label the measurements as local. git fetch origin && git merge origin/dev returned Already up to date.

Perf host access finding: the locked probe `ssh root@10.69.69.63` returned `No route to host`, so the requested perf VM run is unavailable from this worker. I will run the profiles locally against the production build and label the measurements as local. `git fetch origin && git merge origin/dev` returned `Already up to date.`
Author
Owner

The shared blaze runner failed before recording Mail samples in its WebKit pass. apps/web/e2e/harness.mjs::authenticator calls context.newCDPSession, which is Chromium-only; the run exited with CDP session is only available in Chromium. I am updating the shared runner to register and sign in with Chromium, close that browser, then transfer the authenticated cookies to the WebKit measurement context. This keeps one browser active at a time and does not change app authentication.

The shared blaze runner failed before recording Mail samples in its WebKit pass. `apps/web/e2e/harness.mjs::authenticator` calls `context.newCDPSession`, which is Chromium-only; the run exited with `CDP session is only available in Chromium`. I am updating the shared runner to register and sign in with Chromium, close that browser, then transfer the authenticated cookies to the WebKit measurement context. This keeps one browser active at a time and does not change app authentication.
Author
Owner

Harness finding: Playwright strict mode rejected Files readiness because the virtualized selector matched 27 rendered rows. The shared wait now uses the first matching row, and Files waits for an indexed row before opening Quick Look. The sampler now records per-key main-thread time through the queued UI flush, reports p95/max, and fails at 16 ms or above. Files is configured to traverse all 1,200 fixture rows in both directions; Photos and Money use 520 and 220 steps. This harness slice is committed as 2fe7a29bf; the full Files matrix is running locally.

Harness finding: Playwright strict mode rejected Files readiness because the virtualized selector matched 27 rendered rows. The shared wait now uses the first matching row, and Files waits for an indexed row before opening Quick Look. The sampler now records per-key main-thread time through the queued UI flush, reports p95/max, and fails at 16 ms or above. Files is configured to traverse all 1,200 fixture rows in both directions; Photos and Money use 520 and 220 steps. This harness slice is committed as `2fe7a29bf`; the full Files matrix is running locally.
Author
Owner

Finished #641 on branch job/blaze-settings. Head: 94a9433216.

Built the shared surface-agnostic blaze runner and Settings coverage. The harness is in commit 134158825e44c1b4c98140efefaf21a33c62c5fb, the first commit on this branch. Settings now supports ArrowUp/ArrowDown rail traversal, paints the pending selection immediately, keeps visited sections mounted until the overlay closes, caches safe Settings resources per User, and preloads the Settings route. The warm E2E test asserts every sampled frame has matching, complete content.

Exact production-build measurement commands:

# Before
PERF_LOCATION=local CALTERNAL_E2E_ASSET_OVERRIDE=1 CALTERNAL_SERVER_BIN="$CARGO_TARGET_DIR/debug/calternal-server" TMPDIR="$PWD/target/tmp" bun bench/blaze.mjs --surface settings --json artifacts/blaze/settings-before-641.json

# After
export TMPDIR="$PWD/target/tmp" CARGO_PROFILE_DEV_DEBUG=line-tables-only CARGO_INCREMENTAL=0 CARGO_BUILD_JOBS=4 PERF_LOCATION=local CALTERNAL_E2E_ASSET_OVERRIDE=1 CALTERNAL_SERVER_BIN="$CARGO_TARGET_DIR/debug/calternal-server"
node bench/blaze.mjs --surface settings --json artifacts/blaze/settings-after-641.json

The table shows cold/warm pairs. Each cell is baseline → after. “Never complete” counts section steps with no distinct complete paint before the next input.

Browser / repeat Mismatch frames, cold/warm Incomplete frames, cold/warm Never complete, cold/warm Longest task ms, cold/warm p95 frame ms, cold/warm Server RSS before/after KiB
Chromium / 33 ms 0/0 → 0/0 0/0 → 15/0 200/200 → 72/28 445/445 → 220/234 16.7/16.8 → 16.8/16.8 173028/224888 → 175228/220884
Chromium / 15 ms 0/0 → 0/0 0/0 → 14/0 200/200 → 74/36 644/644 → 212/95 16.8/16.8 → 116.6/133.3 157528/220956 → 177156/221408
WebKit / 33 ms 0/0 → 0/0 0/0 → 272/0 200/200 → 145/31 n/a → n/a 23/23 → 32/29 169616/217620 → 162716/226436
WebKit / 15 ms 0/0 → 0/0 0/0 → 5/0 200/200 → 53/45 n/a → n/a 197/178 → 25/25 174980/219740 → 140896/218112

These are local results from production assets, not perf-VM results. The baseline load averages were 14.92/16.16/15.93; the after run recorded 26.18/25.23/24.95. Linux cache dropping was not run. SSH to root@10.69.69.63 returned “No route to host”; bench/hdd-emu.sh is absent. Warm runs had zero incomplete frames in both browsers, but fast-repeat inputs still skipped distinct section paints on this busy host. Chromium also recorded local long tasks above 50 ms. The Settings rail is not present at 390 px because the phone layout is a drill-in list; the harness records those runs as not applicable.

Visual screenshots cover Settings, its sidebar, and the account menu at 390, 820, and 1440 px in light and dark themes. They remain in ignored artifacts/settings-key-541/. The fj issue CLI has no attachment option, so they are not attached to this comment.

Final web gates, verbatim output excerpts:

$ node scripts/check-user-storage.mjs && node scripts/check-type-tokens.mjs && node scripts/check-motion-tokens.mjs && svelte-kit sync && svelte-check --tsconfig ./tsconfig.json
User browser caches use userStorage; only documented device/public-link exceptions remain.
Text sizes and UI shape values use shared role tokens.
UI transitions and animation options use shared motion tokens or documented exceptions.
Loading svelte-check in workspace: /home/kayg/Developer/calternal-wt/blaze-settings/apps/web
Getting Svelte diagnostics...

svelte-check found 0 errors and 0 warnings
$ vitest run

 Test Files  148 passed (148)
      Tests  1011 passed (1011)
   Start at  21:39:17
   Duration  258.67s (transform 52%, environment 20%, import 14%, tests 11%, setup 4%)

Environment  |component| jsdom was created 46 times · 269.31s total, 26% of tracked time

Decisions not specified in DESIGN.md: preserve each visited Settings subtree only for the lifetime of the overlay; keep safe card snapshots in userStorage, but keep raw Admin TOML only in per-User memory until the overlay closes; and enable arrow navigation only for Settings while retaining nav/button semantics. The 2026-10-01 owner override makes keyboard actions use pointer motion timings; reduced-motion preferences still settle motion.

No Rust crate or server API changed, so no Rust gates or API adversarial probe were applicable. cargo clean removed 7237 files (4.6GiB), and apps/web/build was deleted.

Finished #641 on branch job/blaze-settings. Head: 94a9433216da2de7b0c57d1a03cfab6f1500df00. Built the shared surface-agnostic blaze runner and Settings coverage. The harness is in commit 134158825e44c1b4c98140efefaf21a33c62c5fb, the first commit on this branch. Settings now supports ArrowUp/ArrowDown rail traversal, paints the pending selection immediately, keeps visited sections mounted until the overlay closes, caches safe Settings resources per User, and preloads the Settings route. The warm E2E test asserts every sampled frame has matching, complete content. Exact production-build measurement commands: ~~~sh # Before PERF_LOCATION=local CALTERNAL_E2E_ASSET_OVERRIDE=1 CALTERNAL_SERVER_BIN="$CARGO_TARGET_DIR/debug/calternal-server" TMPDIR="$PWD/target/tmp" bun bench/blaze.mjs --surface settings --json artifacts/blaze/settings-before-641.json # After export TMPDIR="$PWD/target/tmp" CARGO_PROFILE_DEV_DEBUG=line-tables-only CARGO_INCREMENTAL=0 CARGO_BUILD_JOBS=4 PERF_LOCATION=local CALTERNAL_E2E_ASSET_OVERRIDE=1 CALTERNAL_SERVER_BIN="$CARGO_TARGET_DIR/debug/calternal-server" node bench/blaze.mjs --surface settings --json artifacts/blaze/settings-after-641.json ~~~ The table shows cold/warm pairs. Each cell is baseline → after. “Never complete” counts section steps with no distinct complete paint before the next input. | Browser / repeat | Mismatch frames, cold/warm | Incomplete frames, cold/warm | Never complete, cold/warm | Longest task ms, cold/warm | p95 frame ms, cold/warm | Server RSS before/after KiB | | --- | --- | --- | --- | --- | --- | --- | | Chromium / 33 ms | 0/0 → 0/0 | 0/0 → 15/0 | 200/200 → 72/28 | 445/445 → 220/234 | 16.7/16.8 → 16.8/16.8 | 173028/224888 → 175228/220884 | | Chromium / 15 ms | 0/0 → 0/0 | 0/0 → 14/0 | 200/200 → 74/36 | 644/644 → 212/95 | 16.8/16.8 → 116.6/133.3 | 157528/220956 → 177156/221408 | | WebKit / 33 ms | 0/0 → 0/0 | 0/0 → 272/0 | 200/200 → 145/31 | n/a → n/a | 23/23 → 32/29 | 169616/217620 → 162716/226436 | | WebKit / 15 ms | 0/0 → 0/0 | 0/0 → 5/0 | 200/200 → 53/45 | n/a → n/a | 197/178 → 25/25 | 174980/219740 → 140896/218112 | These are local results from production assets, not perf-VM results. The baseline load averages were 14.92/16.16/15.93; the after run recorded 26.18/25.23/24.95. Linux cache dropping was not run. SSH to root@10.69.69.63 returned “No route to host”; bench/hdd-emu.sh is absent. Warm runs had zero incomplete frames in both browsers, but fast-repeat inputs still skipped distinct section paints on this busy host. Chromium also recorded local long tasks above 50 ms. The Settings rail is not present at 390 px because the phone layout is a drill-in list; the harness records those runs as not applicable. Visual screenshots cover Settings, its sidebar, and the account menu at 390, 820, and 1440 px in light and dark themes. They remain in ignored artifacts/settings-key-541/. The fj issue CLI has no attachment option, so they are not attached to this comment. Final web gates, verbatim output excerpts: ~~~text $ node scripts/check-user-storage.mjs && node scripts/check-type-tokens.mjs && node scripts/check-motion-tokens.mjs && svelte-kit sync && svelte-check --tsconfig ./tsconfig.json User browser caches use userStorage; only documented device/public-link exceptions remain. Text sizes and UI shape values use shared role tokens. UI transitions and animation options use shared motion tokens or documented exceptions. Loading svelte-check in workspace: /home/kayg/Developer/calternal-wt/blaze-settings/apps/web Getting Svelte diagnostics... svelte-check found 0 errors and 0 warnings ~~~ ~~~text $ vitest run Test Files 148 passed (148) Tests 1011 passed (1011) Start at 21:39:17 Duration 258.67s (transform 52%, environment 20%, import 14%, tests 11%, setup 4%) Environment |component| jsdom was created 46 times · 269.31s total, 26% of tracked time ~~~ Decisions not specified in DESIGN.md: preserve each visited Settings subtree only for the lifetime of the overlay; keep safe card snapshots in userStorage, but keep raw Admin TOML only in per-User memory until the overlay closes; and enable arrow navigation only for Settings while retaining nav/button semantics. The 2026-10-01 owner override makes keyboard actions use pointer motion timings; reduced-motion preferences still settle motion. No Rust crate or server API changed, so no Rust gates or API adversarial probe were applicable. cargo clean removed 7237 files (4.6GiB), and apps/web/build was deleted.
Author
Owner

Files profile finding: one local matrix case waited 120 seconds for the first indexed row from the disk-backed fixture and timed out, so the profile emitted no result table. The runner now logs the browser/viewport/cadence and retains only API paths/statuses, failed-request paths, visible row count, and server warning lines on readiness failure. Files gets one bounded five-minute scan window. The diagnostic change is committed as a65d312a5; I am starting one full retry.

Files profile finding: one local matrix case waited 120 seconds for the first indexed row from the disk-backed fixture and timed out, so the profile emitted no result table. The runner now logs the browser/viewport/cadence and retains only API paths/statuses, failed-request paths, visible row count, and server warning lines on readiness failure. Files gets one bounded five-minute scan window. The diagnostic change is committed as `a65d312a5`; I am starting one full retry.
Author
Owner

The local 10k Mail run confirmed a warmup race. The Mail API returned 50 rows and a valid next cursor on both the first and second page. Before the fix, the 15 ms run reached list index 99 while later pages were still pending. The preloader now shares any in-flight page request and waits for 240 rows before sampling; the new run records 250 loaded rows in every pass.

On WebKit, the warm 15 ms pass recorded 0 selected/content mismatches and 0 incomplete frames. It still recorded 277 key steps with no complete painted frame, with p95 frame time 763 ms and host load average 35.88–39.38. The 33 ms pass recorded 17 mismatches. These are local results under heavy shared-host load, not a quiet-host baseline. The perf VM is unreachable from this worktree (ssh reports No route to host), so the required HDD-emulated comparison remains unavailable here.

The local 10k Mail run confirmed a warmup race. The Mail API returned 50 rows and a valid next cursor on both the first and second page. Before the fix, the 15 ms run reached list index 99 while later pages were still pending. The preloader now shares any in-flight page request and waits for 240 rows before sampling; the new run records 250 loaded rows in every pass. On WebKit, the warm 15 ms pass recorded 0 selected/content mismatches and 0 incomplete frames. It still recorded 277 key steps with no complete painted frame, with p95 frame time 763 ms and host load average 35.88–39.38. The 33 ms pass recorded 17 mismatches. These are local results under heavy shared-host load, not a quiet-host baseline. The perf VM is unreachable from this worktree (`ssh` reports `No route to host`), so the required HDD-emulated comparison remains unavailable here.
Author
Owner

Local #641 Columns run completed with the shared harness and a 10,000-message mailbox (250 rows loaded). At load average 27–35, Chromium recorded 102–179 selected/content mismatch frames and 129–217 incomplete frames per 200-key pass; WebKit recorded up to 196 of each. Some frames showed the selected row one message ahead of the reader, and others had no reader frame. This host is too busy for an acceptance latency result, but it exposes a path that needs another pass: sustained input can outpace the current prefetch/reader updates. I am checking that before reporting final measurements. Perf VM access still returns No route to host.

Local #641 Columns run completed with the shared harness and a 10,000-message mailbox (250 rows loaded). At load average 27–35, Chromium recorded 102–179 selected/content mismatch frames and 129–217 incomplete frames per 200-key pass; WebKit recorded up to 196 of each. Some frames showed the selected row one message ahead of the reader, and others had no reader frame. This host is too busy for an acceptance latency result, but it exposes a path that needs another pass: sustained input can outpace the current prefetch/reader updates. I am checking that before reporting final measurements. Perf VM access still returns `No route to host`.
Author
Owner

The post-prefetch local Columns run cleared all selected/body mismatches in Chromium (0 across all four passes). WebKit exposed a second issue under repeated Down: in the 33 ms cold pass, the selected row stopped advancing at index 98 while the reader showed index 99; later repeat events continued to target the old focused row. The handler moves focus only after tick(), which lets a repeat event arrive before the target row receives focus. I am updating the handler to base movement on the latest selected thread and to focus an already-rendered adjacent row synchronously. Local load was 8.9–22 average; this is functional race evidence, not a clean latency run.

The post-prefetch local Columns run cleared all selected/body mismatches in Chromium (0 across all four passes). WebKit exposed a second issue under repeated Down: in the 33 ms cold pass, the selected row stopped advancing at index 98 while the reader showed index 99; later repeat events continued to target the old focused row. The handler moves focus only after `tick()`, which lets a repeat event arrive before the target row receives focus. I am updating the handler to base movement on the latest selected thread and to focus an already-rendered adjacent row synchronously. Local load was 8.9–22 average; this is functional race evidence, not a clean latency run.
Author
Owner

WebKit finding and fix: the session cookie was present in transferred storage state, but WebKit returned 401 for the Files APIs on the plain HTTP test origin because it does not send the __Host- Secure cookie there. I added a temporary HTTPS localhost proxy to the shared E2E harness and routed WebKit matrix cases through it; the real local server remains the only writer. A focused WebKit Files check now reports identity 200 and one indexed row. The fix is committed as c9ca3774d.

WebKit finding and fix: the session cookie was present in transferred storage state, but WebKit returned 401 for the Files APIs on the plain HTTP test origin because it does not send the `__Host-` Secure cookie there. I added a temporary HTTPS localhost proxy to the shared E2E harness and routed WebKit matrix cases through it; the real local server remains the only writer. A focused WebKit Files check now reports identity 200 and one indexed row. The fix is committed as `c9ca3774d`.
Author
Owner

Round 2 traversal work starting on job/blaze-settings at 94a9433216da2de7b0c57d1a03cfab6f1500df00, based on origin/dev 687ff703136e71e89f8dfba139e93cd0788b25c1. I will measure Settings rail traversal on the production build and post the per-key-step paint table before changing the traversal path.

Round 2 traversal work starting on `job/blaze-settings` at `94a9433216da2de7b0c57d1a03cfab6f1500df00`, based on `origin/dev` `687ff703136e71e89f8dfba139e93cd0788b25c1`. I will measure Settings rail traversal on the production build and post the per-key-step paint table before changing the traversal path.
Author
Owner

The keyboard focus fix removed the WebKit row-98 stall and allowed the 200-step sweep to reach index 199. The local WebKit profile still recorded frames with no selected row during scrolling, while the reader remained on a different body. The selection path updates the native scroll position, but the virtual list waits for the later scroll event to update its own scroll state. I am syncing the virtual scroll state in the same keyboard step so its rendered window follows the selected ID immediately. These frames were recorded at load average 18–25 and remain local stress evidence.

The keyboard focus fix removed the WebKit row-98 stall and allowed the 200-step sweep to reach index 199. The local WebKit profile still recorded frames with no selected row during scrolling, while the reader remained on a different body. The selection path updates the native scroll position, but the virtual list waits for the later scroll event to update its own scroll state. I am syncing the virtual scroll state in the same keyboard step so its rendered window follows the selected ID immediately. These frames were recorded at load average 18–25 and remain local stress evidence.
Author
Owner

Production-build Settings traversal baseline, before application changes. Exact run:

export PERF_LOCATION=local CALTERNAL_E2E_ASSET_OVERRIDE=1 CALTERNAL_SERVER_BIN="$CARGO_TARGET_DIR/debug/calternal-server" TMPDIR="$PWD/target/tmp" BLAZE_BROWSERS=chromium BLAZE_WIDTHS=1440x900 BLAZE_CADENCES=15,33
taskset -c 0,1 node bench/blaze.mjs --surface settings --json artifacts/blaze/settings-round2-before-641.json

The full run samples 400 keydowns per phase (200 Down, then 200 Up). Of those, 40 keydowns changed the selected section; the other 360 repeated at a list boundary and kept the same selected section. The table below reports every selection-changing keydown and whether a complete matching section appeared before the next keydown. Local load was 19.56 before and 27.80 after (1/5/15-minute averages 19.56/14.41/19.15 before).

Phase / repeat Mismatch frames Incomplete frames Steps without distinct complete paint (all keydowns) Changed section steps painted complete Longest task p95 frame
cold / 33 ms 0 17 47 9 / 40 171 ms 16.8 ms
warm / 33 ms 0 0 32 10 / 40 95 ms 16.8 ms
cold / 15 ms 0 20 93 3 / 40 250 ms 33.4 ms
warm / 15 ms 0 0 39 10 / 40 83 ms 133.4 ms

Per-step warm 15 ms table (the 30 none rows did not show a distinct complete paint before the next key):

Step Key Section change Complete paint after key
1 ↓ account → apps none
2 ↓ apps → maintenance none
3 ↓ maintenance → appearance 26.6 ms
4 ↓ appearance → editor none
5 ↓ editor → notifications none
6 ↓ notifications → calendars none
7 ↓ calendars → mail 4.4 ms
8 ↓ mail → ai none
9 ↓ ai → photos none
10 ↓ photos → files none
11 ↓ files → plugins none
12 ↓ plugins → admin/users 10.1 ms
13 ↓ admin/users → admin/invitations none
14 ↓ admin/invitations → admin/sign-in none
15 ↓ admin/sign-in → admin/configuration none
16 ↓ admin/configuration → admin/backups none
17 ↓ admin/backups → admin/maintenance none
18 ↓ admin/maintenance → admin/apps none
19 ↓ admin/apps → admin/plugins none
20 ↓ admin/plugins → admin/system 35.8 ms
201 ↑ admin/system → admin/plugins 60.8 ms
202 ↑ admin/plugins → admin/apps 6.4 ms
203 ↑ admin/apps → admin/maintenance none
204 ↑ admin/maintenance → admin/backups none
205 ↑ admin/backups → admin/configuration none
206 ↑ admin/configuration → admin/sign-in none
207 ↑ admin/sign-in → admin/invitations none
208 ↑ admin/invitations → admin/users 9.7 ms
209 ↑ admin/users → plugins none
210 ↑ plugins → files 27.8 ms
211 ↑ files → photos none
212 ↑ photos → ai none
213 ↑ ai → mail 11.7 ms
214 ↑ mail → calendars none
215 ↑ calendars → notifications none
216 ↑ notifications → editor 29.5 ms
217 ↑ editor → appearance none
218 ↑ appearance → maintenance none
219 ↑ maintenance → apps none
220 ↑ apps → account none

Artifact: artifacts/blaze/settings-round2-before-641.json (local, ignored). The run is host-contended and is used to establish the per-step failure, not as quiet-host latency evidence.

Production-build Settings traversal baseline, before application changes. Exact run: ```sh export PERF_LOCATION=local CALTERNAL_E2E_ASSET_OVERRIDE=1 CALTERNAL_SERVER_BIN="$CARGO_TARGET_DIR/debug/calternal-server" TMPDIR="$PWD/target/tmp" BLAZE_BROWSERS=chromium BLAZE_WIDTHS=1440x900 BLAZE_CADENCES=15,33 taskset -c 0,1 node bench/blaze.mjs --surface settings --json artifacts/blaze/settings-round2-before-641.json ``` The full run samples 400 keydowns per phase (200 Down, then 200 Up). Of those, 40 keydowns changed the selected section; the other 360 repeated at a list boundary and kept the same selected section. The table below reports every selection-changing keydown and whether a complete matching section appeared before the next keydown. Local load was 19.56 before and 27.80 after (1/5/15-minute averages 19.56/14.41/19.15 before). | Phase / repeat | Mismatch frames | Incomplete frames | Steps without distinct complete paint (all keydowns) | Changed section steps painted complete | Longest task | p95 frame | | --- | ---: | ---: | ---: | ---: | ---: | ---: | | cold / 33 ms | 0 | 17 | 47 | 9 / 40 | 171 ms | 16.8 ms | | warm / 33 ms | 0 | 0 | 32 | 10 / 40 | 95 ms | 16.8 ms | | cold / 15 ms | 0 | 20 | 93 | 3 / 40 | 250 ms | 33.4 ms | | warm / 15 ms | 0 | 0 | 39 | 10 / 40 | 83 ms | 133.4 ms | Per-step warm 15 ms table (the 30 `none` rows did not show a distinct complete paint before the next key): | Step | Key | Section change | Complete paint after key | | ---: | :---: | --- | ---: | | 1 | ↓ | account → apps | none | | 2 | ↓ | apps → maintenance | none | | 3 | ↓ | maintenance → appearance | 26.6 ms | | 4 | ↓ | appearance → editor | none | | 5 | ↓ | editor → notifications | none | | 6 | ↓ | notifications → calendars | none | | 7 | ↓ | calendars → mail | 4.4 ms | | 8 | ↓ | mail → ai | none | | 9 | ↓ | ai → photos | none | | 10 | ↓ | photos → files | none | | 11 | ↓ | files → plugins | none | | 12 | ↓ | plugins → admin/users | 10.1 ms | | 13 | ↓ | admin/users → admin/invitations | none | | 14 | ↓ | admin/invitations → admin/sign-in | none | | 15 | ↓ | admin/sign-in → admin/configuration | none | | 16 | ↓ | admin/configuration → admin/backups | none | | 17 | ↓ | admin/backups → admin/maintenance | none | | 18 | ↓ | admin/maintenance → admin/apps | none | | 19 | ↓ | admin/apps → admin/plugins | none | | 20 | ↓ | admin/plugins → admin/system | 35.8 ms | | 201 | ↑ | admin/system → admin/plugins | 60.8 ms | | 202 | ↑ | admin/plugins → admin/apps | 6.4 ms | | 203 | ↑ | admin/apps → admin/maintenance | none | | 204 | ↑ | admin/maintenance → admin/backups | none | | 205 | ↑ | admin/backups → admin/configuration | none | | 206 | ↑ | admin/configuration → admin/sign-in | none | | 207 | ↑ | admin/sign-in → admin/invitations | none | | 208 | ↑ | admin/invitations → admin/users | 9.7 ms | | 209 | ↑ | admin/users → plugins | none | | 210 | ↑ | plugins → files | 27.8 ms | | 211 | ↑ | files → photos | none | | 212 | ↑ | photos → ai | none | | 213 | ↑ | ai → mail | 11.7 ms | | 214 | ↑ | mail → calendars | none | | 215 | ↑ | calendars → notifications | none | | 216 | ↑ | notifications → editor | 29.5 ms | | 217 | ↑ | editor → appearance | none | | 218 | ↑ | appearance → maintenance | none | | 219 | ↑ | maintenance → apps | none | | 220 | ↑ | apps → account | none | Artifact: `artifacts/blaze/settings-round2-before-641.json` (local, ignored). The run is host-contended and is used to establish the per-step failure, not as quiet-host latency evidence.
Author
Owner

Sampler finding: the first Chromium desktop cold pass recorded 1,877 identity mismatch frames. The saved frames show selectedValue: null while the Quick Look content key is present after the selected Files row leaves the virtualized DOM. The sampler used .fc-item.focus for selection, so this invalidated identity matching and paint latency. Files now reads selected identity/position from the mounted Quick Look viewer (.ql[data-viewer-key] / .ql[data-viewer-position]) and compares it with the stage identity/position. I will rerun the full matrix with corrected sampling.

Sampler finding: the first Chromium desktop cold pass recorded 1,877 identity mismatch frames. The saved frames show `selectedValue: null` while the Quick Look content key is present after the selected Files row leaves the virtualized DOM. The sampler used `.fc-item.focus` for selection, so this invalidated identity matching and paint latency. Files now reads selected identity/position from the mounted Quick Look viewer (`.ql[data-viewer-key]` / `.ql[data-viewer-position]`) and compares it with the stage identity/position. I will rerun the full matrix with corrected sampling.
Author
Owner

Blaze test result for #640

Extended the shared bench/blaze.mjs harness with Mail Columns, Split, and Morph surfaces. It seeds 10,000 test messages, waits for 240 list rows and their bodies, then records 200 Down and 200 Up events at 33 ms and 15 ms. The sampler checks the selected ID, reader body ID, sanitized HTML frame readiness, frame loss, browser heap, server CPU/RSS, and host load.

The initial run showed that the old 48-body cache and list-only warmup could not keep the reader warm through the sweep. I expanded the bounded cache, added one four-worker prefetch queue, warmed the first 240 bodies, and fixed repeat focus and virtual-scroll state races. Final Chromium profiles have zero selected/body mismatch and incomplete frames across all three layouts. WebKit still records incomplete frames under the overloaded local host; these results do not verify the every-step requirement.

Before / after

Profile Before: mismatch / incomplete per pass After Chromium After WebKit Final local 1-minute load average
Columns · 1440 114 / 128 average (8 passes) 0 / 0 (4 passes) 74 / 74 average (4 passes) 22–33
Split · 1440 No pre-change layout existed 0 / 0 (4 passes) 186.5 / 186.5 average (4 passes) 26–34
Morph · 1440 and 820 No pre-change layout existed 0 / 0 (8 passes) 23.8 / 23.8 average (8 passes) 27–33

The before Columns profile averaged 337.9 key steps without a complete paint per pass. In the final run, Chromium still had missed intermediate animation-frame samples under load, but every sampled frame had the selected body ready. WebKit’s p95 frame intervals ranged from 289 to 690 ms. The local perf VM connection returned No route to host, so these profiles did not use bench/hdd-emu.sh; the 50 ms warm target remains unverified. The current profile JSON is in artifacts/blaze/ in the worktree.

Warm observed paint latency, selected/body mismatch totals and p95 frame interval are summarized in #640. The final profiles all reached 250 rows from the 10,000-message test fixture.

Gate output:

svelte-check found 0 errors and 0 warnings
Test Files  150 passed (150)
      Tests  1015 passed (1015)
Mail layouts e2e passed; screenshots are in /home/kayg/Developer/calternal-wt/maillayouts/artifacts/mail-layouts
✓ built in 58.72s
Wrote site to "build"
✔ done

Head: 68bd0b4999e4c21dc613f6d9b995b5d319653012.

## Blaze test result for #640 Extended the shared `bench/blaze.mjs` harness with Mail Columns, Split, and Morph surfaces. It seeds 10,000 test messages, waits for 240 list rows and their bodies, then records 200 Down and 200 Up events at 33 ms and 15 ms. The sampler checks the selected ID, reader body ID, sanitized HTML frame readiness, frame loss, browser heap, server CPU/RSS, and host load. The initial run showed that the old 48-body cache and list-only warmup could not keep the reader warm through the sweep. I expanded the bounded cache, added one four-worker prefetch queue, warmed the first 240 bodies, and fixed repeat focus and virtual-scroll state races. Final Chromium profiles have zero selected/body mismatch and incomplete frames across all three layouts. WebKit still records incomplete frames under the overloaded local host; these results do not verify the every-step requirement. ### Before / after | Profile | Before: mismatch / incomplete per pass | After Chromium | After WebKit | Final local 1-minute load average | | --- | ---: | ---: | ---: | ---: | | Columns · 1440 | 114 / 128 average (8 passes) | 0 / 0 (4 passes) | 74 / 74 average (4 passes) | 22–33 | | Split · 1440 | No pre-change layout existed | 0 / 0 (4 passes) | 186.5 / 186.5 average (4 passes) | 26–34 | | Morph · 1440 and 820 | No pre-change layout existed | 0 / 0 (8 passes) | 23.8 / 23.8 average (8 passes) | 27–33 | The before Columns profile averaged 337.9 key steps without a complete paint per pass. In the final run, Chromium still had missed intermediate animation-frame samples under load, but every sampled frame had the selected body ready. WebKit’s p95 frame intervals ranged from 289 to 690 ms. The local perf VM connection returned `No route to host`, so these profiles did not use `bench/hdd-emu.sh`; the 50 ms warm target remains unverified. The current profile JSON is in `artifacts/blaze/` in the worktree. Warm observed paint latency, selected/body mismatch totals and p95 frame interval are summarized in #640. The final profiles all reached 250 rows from the 10,000-message test fixture. Gate output: ```text svelte-check found 0 errors and 0 warnings Test Files 150 passed (150) Tests 1015 passed (1015) Mail layouts e2e passed; screenshots are in /home/kayg/Developer/calternal-wt/maillayouts/artifacts/mail-layouts ✓ built in 58.72s Wrote site to "build" ✔ done ``` Head: `68bd0b4999e4c21dc613f6d9b995b5d319653012`.
Author
Owner

Frame-paced production-build follow-up for Settings. The opted-in arrow navigation now queues each intended section and applies one selection per rendered frame; focus and aria-current move together. Pointer picks remain immediate. The harness predicts each target from the prior key intent and attributes a complete frame in step order, with the opposite direction as a boundary so a later revisit cannot hide a miss. Each pass sent 200 Down and 200 Up keydowns at the requested cadence. Chromium was pinned to CPUs 0–1. This local host was busy: load average before the run was 18.53 / 16.86 / 16.45 (1/5/15 minutes).

Pass Cadence Mismatch frames Incomplete frames Unpainted steps Longest task p95 frame interval
cold 33 0 24 16 152 ms 116.7 ms
warm 33 0 0 0 65 ms 116.6 ms
cold 15 0 27 16 98 ms 150 ms
warm 15 0 0 0 56 ms 200 ms

Warm runs now paint all 40 distinct section changes at both cadences, with no mismatches or incomplete frames. The local p95 frame interval and >50 ms tasks remain load-sensitive and exceed the target in this run. Cold first-traversal has incomplete frames because the sections have no per-User cache yet; warm is the #641 merge target. SLOW-only timing is not conclusive on this host.

Per-step complete-paint latency table for the warm runs (ms from the key intent to its complete frame):

Direction Section 33 ms 15 ms
Down Apps 87.0 235.6
Down Maintenance 213.6 424.1
Down Appearance 330.2 522.8
Down Editor 404.3 654.4
Down Notifications 500.0 870.0
Down Calendars 610.2 960.9
Down Mail 672.5 1070.7
Down AI 736.7 1169.9
Down Photos 846.1 1284.1
Down Files 940.3 1415.0
Down Plugins 1007.9 1495.9
Down Admin · Users 1085.4 1611.5
Down Admin · Invitations 1147.6 1827.3
Down Admin · Sign-in 1205.3 1941.4
Down Admin · Configuration 1302.0 2061.9
Down Admin · Backups 1365.4 2194.2
Down Admin · Maintenance 1507.0 2326.6
Down Admin · Apps 1587.3 2459.3
Down Admin · Plugins 1680.8 2589.5
Down Admin · System 1754.8 2755.4
Up Admin · Plugins 28.5 25.1
Up Admin · Apps 105.5 23.0
Up Admin · Maintenance 197.0 271.3
Up Admin · Backups 277.6 332.2
Up Admin · Configuration 474.6 429.9
Up Admin · Sign-in 532.4 559.5
Up Admin · Invitations 612.0 717.0
Up Admin · Users 692.5 829.8
Up Plugins 742.4 939.9
Up Files 932.8 1083.2
Up Photos 996.5 1363.1
Up AI 1083.6 1528.2
Up Mail 1163.5 1690.2
Up Calendars 1243.5 1872.1
Up Notifications 1340.6 2003.6
Up Editor 1419.3 2130.0
Up Appearance 1499.9 2372.3
Up Maintenance 1577.9 2498.9
Up Apps 1658.7 2664.1
Up Account 1695.5 2906.1
Frame-paced production-build follow-up for Settings. The opted-in arrow navigation now queues each intended section and applies one selection per rendered frame; focus and aria-current move together. Pointer picks remain immediate. The harness predicts each target from the prior key intent and attributes a complete frame in step order, with the opposite direction as a boundary so a later revisit cannot hide a miss. Each pass sent 200 Down and 200 Up keydowns at the requested cadence. Chromium was pinned to CPUs 0–1. This local host was busy: load average before the run was 18.53 / 16.86 / 16.45 (1/5/15 minutes). | Pass | Cadence | Mismatch frames | Incomplete frames | Unpainted steps | Longest task | p95 frame interval | | --- | ---: | ---: | ---: | ---: | ---: | ---: | | cold | 33 | 0 | 24 | 16 | 152 ms | 116.7 ms | | warm | 33 | 0 | 0 | 0 | 65 ms | 116.6 ms | | cold | 15 | 0 | 27 | 16 | 98 ms | 150 ms | | warm | 15 | 0 | 0 | 0 | 56 ms | 200 ms | Warm runs now paint all 40 distinct section changes at both cadences, with no mismatches or incomplete frames. The local p95 frame interval and >50 ms tasks remain load-sensitive and exceed the target in this run. Cold first-traversal has incomplete frames because the sections have no per-User cache yet; warm is the #641 merge target. SLOW-only timing is not conclusive on this host. Per-step complete-paint latency table for the warm runs (ms from the key intent to its complete frame): | Direction | Section | 33 ms | 15 ms | | --- | --- | ---: | ---: | | Down | Apps | 87.0 | 235.6 | | Down | Maintenance | 213.6 | 424.1 | | Down | Appearance | 330.2 | 522.8 | | Down | Editor | 404.3 | 654.4 | | Down | Notifications | 500.0 | 870.0 | | Down | Calendars | 610.2 | 960.9 | | Down | Mail | 672.5 | 1070.7 | | Down | AI | 736.7 | 1169.9 | | Down | Photos | 846.1 | 1284.1 | | Down | Files | 940.3 | 1415.0 | | Down | Plugins | 1007.9 | 1495.9 | | Down | Admin · Users | 1085.4 | 1611.5 | | Down | Admin · Invitations | 1147.6 | 1827.3 | | Down | Admin · Sign-in | 1205.3 | 1941.4 | | Down | Admin · Configuration | 1302.0 | 2061.9 | | Down | Admin · Backups | 1365.4 | 2194.2 | | Down | Admin · Maintenance | 1507.0 | 2326.6 | | Down | Admin · Apps | 1587.3 | 2459.3 | | Down | Admin · Plugins | 1680.8 | 2589.5 | | Down | Admin · System | 1754.8 | 2755.4 | | Up | Admin · Plugins | 28.5 | 25.1 | | Up | Admin · Apps | 105.5 | 23.0 | | Up | Admin · Maintenance | 197.0 | 271.3 | | Up | Admin · Backups | 277.6 | 332.2 | | Up | Admin · Configuration | 474.6 | 429.9 | | Up | Admin · Sign-in | 532.4 | 559.5 | | Up | Admin · Invitations | 612.0 | 717.0 | | Up | Admin · Users | 692.5 | 829.8 | | Up | Plugins | 742.4 | 939.9 | | Up | Files | 932.8 | 1083.2 | | Up | Photos | 996.5 | 1363.1 | | Up | AI | 1083.6 | 1528.2 | | Up | Mail | 1163.5 | 1690.2 | | Up | Calendars | 1243.5 | 1872.1 | | Up | Notifications | 1340.6 | 2003.6 | | Up | Editor | 1419.3 | 2130.0 | | Up | Appearance | 1499.9 | 2372.3 | | Up | Maintenance | 1577.9 | 2498.9 | | Up | Apps | 1658.7 | 2664.1 | | Up | Account | 1695.5 | 2906.1 |
Author
Owner

Progress (2026-10-02, commit 00bfbd93d): the Photos surface initially showed its real empty state while startup reconciliation was still building the disk fixture projection. The API later exposed all 520 items. I updated the shared harness to poll the real Photos bucket count and reload after it becomes nonzero, so it measures the indexed fixture rather than an early empty response.

Focused local Chromium run (200 positions each way, 1440×900, 33 ms; full Photos fixture has 520 items): cold/warm had 8/10 incomplete frames and 39/27 steps never sampled complete. Main-thread key-handler p95 was 0.1 ms in both; longest browser task was 4,745 ms. Peak server RSS was 460.1/460.2 MiB. The local load average at start was 16.53, 14.39, 14.59. The profile failed its paint gate. The perf VM remains unreachable by SSH, so this is local evidence under a heavily loaded shared host, not the requested locked VM measurement.

Progress (2026-10-02, commit 00bfbd93d): the Photos surface initially showed its real empty state while startup reconciliation was still building the disk fixture projection. The API later exposed all 520 items. I updated the shared harness to poll the real Photos bucket count and reload after it becomes nonzero, so it measures the indexed fixture rather than an early empty response. Focused local Chromium run (200 positions each way, 1440×900, 33 ms; full Photos fixture has 520 items): cold/warm had 8/10 incomplete frames and 39/27 steps never sampled complete. Main-thread key-handler p95 was 0.1 ms in both; longest browser task was 4,745 ms. Peak server RSS was 460.1/460.2 MiB. The local load average at start was 16.53, 14.39, 14.59. The profile failed its paint gate. The perf VM remains unreachable by SSH, so this is local evidence under a heavily loaded shared host, not the requested locked VM measurement.
Author
Owner

Starting the WebKit blaze follow-up on job/maillayouts at 68bd0b4999e4c21dc613f6d9b995b5d319653012, based on origin/dev 687ff703136e71e89f8dfba139e93cd0788b25c1. I will run the shared production-build harness locally for Columns, Split and Morph in Chromium and WebKit, then report the per-engine mismatch/incomplete-frame and timing table.

Starting the WebKit blaze follow-up on `job/maillayouts` at `68bd0b4999e4c21dc613f6d9b995b5d319653012`, based on `origin/dev` `687ff703136e71e89f8dfba139e93cd0788b25c1`. I will run the shared production-build harness locally for Columns, Split and Morph in Chromium and WebKit, then report the per-engine mismatch/incomplete-frame and timing table.
Author
Owner

Cold traversal follow-up

The production profile after the first frame-queue change has zero missed/incomplete steps on both warm passes, but the cold pass still advances before section content finishes. At 1440×900 on the local host (load 18.53, 16.86, 16.45 before the profile), the 33 ms cold pass had 24 incomplete frames and 16 steps with no complete paint; the 15 ms cold pass had 27 incomplete frames and the same 16 missed steps. The first cold misses were account → apps and apps → maintenance; later misses covered Notifications, Calendars, Mail, AI, Photos, Files, Plugins and Admin sections. The full per-step samples are in artifacts/blaze/settings-round2-frame-paced.json in the worktree.

I am changing the queue to wait for the caller's selected route content to become complete and paint before it applies the next queued arrow step. The list still updates the selected route immediately; this wait only prevents a later repeat from replacing a section before its real content has painted.

### Cold traversal follow-up The production profile after the first frame-queue change has zero missed/incomplete steps on both warm passes, but the cold pass still advances before section content finishes. At 1440×900 on the local host (load `18.53, 16.86, 16.45` before the profile), the 33 ms cold pass had 24 incomplete frames and 16 steps with no complete paint; the 15 ms cold pass had 27 incomplete frames and the same 16 missed steps. The first cold misses were `account → apps` and `apps → maintenance`; later misses covered Notifications, Calendars, Mail, AI, Photos, Files, Plugins and Admin sections. The full per-step samples are in `artifacts/blaze/settings-round2-frame-paced.json` in the worktree. I am changing the queue to wait for the caller's selected route content to become complete and paint before it applies the next queued arrow step. The list still updates the selected route immediately; this wait only prevents a later repeat from replacing a section before its real content has painted.
Author
Owner

Columns WebKit blaze finding (local shared build host): Chromium reported 0 mismatch and 0 incomplete frames for both cadences. WebKit reported:

phase cadence mismatch frames incomplete frames
cold 33 ms 216 216
warm 33 ms 258 258
cold 15 ms 220 220
warm 15 ms 333 333

The 33 ms WebKit passes also had 19,996–22,308 ms p95 paint latency on the busy local host, so I will report timings with load context. I am tracing whether repeated srcdoc navigation while the first frame is loading leaves WebKit on a stale body.

Columns WebKit blaze finding (local shared build host): Chromium reported 0 mismatch and 0 incomplete frames for both cadences. WebKit reported: | phase | cadence | mismatch frames | incomplete frames | |---|---:|---:|---:| | cold | 33 ms | 216 | 216 | | warm | 33 ms | 258 | 258 | | cold | 15 ms | 220 | 220 | | warm | 15 ms | 333 | 333 | The 33 ms WebKit passes also had 19,996–22,308 ms p95 paint latency on the busy local host, so I will report timings with load context. I am tracing whether repeated `srcdoc` navigation while the first frame is loading leaves WebKit on a stale body.
Author
Owner

The WebKit body is not a srcdoc swap on warm selection: fitFrame reuses the document, but its live update path writes innerHTML and immediately calls getBoundingClientRect() to fit the iframe. That forces a synchronous layout inside the WebKit key handler. I am changing live updates to let the existing ResizeObserver measure the new root before paint; only a newly loaded document needs the initial synchronous fit. The long-to-short E2E will wait for the observer-driven shrink, and the blaze will be rerun.

The WebKit body is not a `srcdoc` swap on warm selection: `fitFrame` reuses the document, but its live update path writes `innerHTML` and immediately calls `getBoundingClientRect()` to fit the iframe. That forces a synchronous layout inside the WebKit key handler. I am changing live updates to let the existing ResizeObserver measure the new root before paint; only a newly loaded document needs the initial synchronous fit. The long-to-short E2E will wait for the observer-driven shrink, and the blaze will be rerun.
Author
Owner

Tabs profile and screenshot update (#641)

The Tabs selector is now resolved against the real mode TabList. A focused production-build run (local host, Chromium 1440×900, 33 ms, 200 mode-shortcut steps, cold + warm) completed collection but failed the zero-incomplete gate:

phase mismatch frames incomplete frames steps never painted complete p50 / p95 paint latency max key handler longest task
cold 42 282 75 1,138.1 / 5,787.8 ms 0.1 ms 597 ms
warm 10 77 46 1,404.7 / 9,961.7 ms 0.1 ms 597 ms

Local load average at run start was 24.4, 19.35, 18.94. This is local-only evidence; the perf VM remains unreachable, so it is not a stable baseline. The frames show warm route transitions taking longer than the key interval, so the profile cannot establish a product regression without a quiet host run.

The screenshot-only path had not installed the userStorage test seam before calling setTheme. I fixed the harness and captured Tabs from the production SPA at 390, 820 and 1440 px in light and dark. Files, Photos and Money captures remain in progress. Screenshots stay in ignored artifacts and are not committed.

Tabs profile and screenshot update (#641) The Tabs selector is now resolved against the real mode TabList. A focused production-build run (local host, Chromium 1440×900, 33 ms, 200 mode-shortcut steps, cold + warm) completed collection but failed the zero-incomplete gate: | phase | mismatch frames | incomplete frames | steps never painted complete | p50 / p95 paint latency | max key handler | longest task | | --- | ---: | ---: | ---: | ---: | ---: | ---: | | cold | 42 | 282 | 75 | 1,138.1 / 5,787.8 ms | 0.1 ms | 597 ms | | warm | 10 | 77 | 46 | 1,404.7 / 9,961.7 ms | 0.1 ms | 597 ms | Local load average at run start was 24.4, 19.35, 18.94. This is local-only evidence; the perf VM remains unreachable, so it is not a stable baseline. The frames show warm route transitions taking longer than the key interval, so the profile cannot establish a product regression without a quiet host run. The screenshot-only path had not installed the userStorage test seam before calling setTheme. I fixed the harness and captured Tabs from the production SPA at 390, 820 and 1440 px in light and dark. Files, Photos and Money captures remain in progress. Screenshots stay in ignored artifacts and are not committed.
Author
Owner

Round 2 cold-profile finding: applyKeyboardStep() clears its RAF handle before it awaits the route callback. While that callback is still waiting for the selected Settings pane to complete, each repeated key sees a null handle and schedules another consumer. Those consumers run concurrently, advance routes while earlier panes still show aria-busy/Loading, and cancel the earlier paint waits. The captured 1440×900 Chromium trace at 33 ms showed 218 of 400 steps without a complete frame; the 15 ms pass showed 218 as well. The frame sequence advanced through Apps, Maintenance, Appearance, Editor, Notifications, Calendars, Mail, AI, Photos, Files, Plugins, and Admin while each visible pane was still busy. I am adding an explicit in-flight guard so one callback owns the queue until its complete paint resolves.

Round 2 cold-profile finding: `applyKeyboardStep()` clears its RAF handle before it awaits the route callback. While that callback is still waiting for the selected Settings pane to complete, each repeated key sees a null handle and schedules another consumer. Those consumers run concurrently, advance routes while earlier panes still show `aria-busy`/Loading, and cancel the earlier paint waits. The captured 1440×900 Chromium trace at 33 ms showed 218 of 400 steps without a complete frame; the 15 ms pass showed 218 as well. The frame sequence advanced through Apps, Maintenance, Appearance, Editor, Notifications, Calendars, Mail, AI, Photos, Files, Plugins, and Admin while each visible pane was still busy. I am adding an explicit in-flight guard so one callback owns the queue until its complete paint resolves.
Author
Owner

The isolated frame update path no longer forces a parent layout read after each warm body update, and the long-to-short and image-heavy E2E checks pass with ResizeObserver sizing. That change alone did not clear the WebKit blaze: Columns still reports reader/selection mismatches (150–329 frames per pass before the focus guard; 227–436 after the guard). Chromium reports zero mismatches and zero incomplete frames in both runs. The mismatch traces show WebKit's [aria-current=true] row disappearing while all 250 rows remain in the list, with the reader stuck on a prior message. I am instrumenting the virtual range, scroll position and active row to identify why WebKit loses the selected row before applying another fix.

The isolated frame update path no longer forces a parent layout read after each warm body update, and the long-to-short and image-heavy E2E checks pass with ResizeObserver sizing. That change alone did not clear the WebKit blaze: Columns still reports reader/selection mismatches (150–329 frames per pass before the focus guard; 227–436 after the guard). Chromium reports zero mismatches and zero incomplete frames in both runs. The mismatch traces show WebKit's `[aria-current=true]` row disappearing while all 250 rows remain in the list, with the reader stuck on a prior message. I am instrumenting the virtual range, scroll position and active row to identify why WebKit loses the selected row before applying another fix.
Author
Owner

Finished #641 on job/blaze-surfaces at head 30f644436262f125d7ecbb5c246d85a0eb5fe52c.

Built:

  • Files and Photos now prefetch full images and neighbour previews through bounded decoded-image LRUs. ImageView waits for img.decode() before replacing the preview. Files also prefetches text and PDF neighbours through bounded caches.
  • Money now coalesces and caches register data, warms three accounts in travel direction, and queues held arrow navigation with a 512-target cap.
  • Tabs now register shortcuts for modes 1–9 and start route navigation immediately while cache refresh runs in the background.
  • Fixed the shared screenshot harness to initialize the userStorage seam, and to open the Money sidebar only below the app's 768 px breakpoint.

All 24 production-build screenshots are attached to this issue: Files, Photos, Money and Tabs at 390, 820 and 1440 px in light and dark.

Performance results are focused local runs only: Chromium, 1440×900, 33 ms cadence, cold and warm. No before baseline exists in docs/perf/baseline.json. The perf VM at root@10.69.69.63 was unreachable (No route to host), so these numbers are not a quiet-host comparison.

Surface Phase Incomplete / unpainted p50 / p95 paint ms CPU % RSS peak KiB Longest task ms
Files (1,200 mixed rows; 200 steps each way) cold 57 / 348 47.4 / 19,977.9 87.6 290,012 334
Files warm 115 / 387 23.9 / 56.7 19.5 313,876 1,690
Photos (520 items; 200 steps each way) cold 8 / 39 63.7 / 76,950.4 80.2 471,128 4,745
Photos warm 10 / 27 68.6 / 27,528.2 83.5 471,276 4,745
Money (220 accounts; 220 steps each way) cold 2 / 62 6,692.7 / 10,604.9 40.3 234,160 1,630
Money warm 3 / 74 18,274.5 / 26,751.5 20.4 234,244 1,630
Tabs (200 shortcuts) cold 282 / 75 1,138.1 / 5,787.8 21.0 210,984 597
Tabs warm 77 / 46 1,404.7 / 9,961.7 6.3 211,996 597

The blaze zero-incomplete gate did not pass these local profiles. Load averages at run start were 9.6–24.4 across the runs, and the frame cadence was also delayed by long tasks. This does not replace the requested Chromium/WebKit × desktop/phone × cold/warm matrix on the perf VM. No valid before/after comparison is available.

Web gate output (verbatim):

$ node scripts/check-user-storage.mjs && node scripts/check-type-tokens.mjs && node scripts/check-motion-tokens.mjs && svelte-kit sync && svelte-check --tsconfig ./tsconfig.json
User browser caches use userStorage; only documented device/public-link exceptions remain.
Text sizes and UI shape values use shared role tokens.
UI transitions and animation options use shared motion tokens or documented exceptions.
Loading svelte-check in workspace: /home/kayg/Developer/calternal-wt/blaze-surfaces/apps/web
Getting Svelte diagnostics...
svelte-check found 0 errors and 0 warnings
$ vitest run
Test Files  152 passed (152)
Tests  1020 passed (1020)
Start at  01:15:06
Duration  57.79s (transform 46%, environment 20%, import 18%, tests 12%, setup 4%)

Decisions not specified in DESIGN: use 8 decoded images, 5 PDF documents, and 8 text entries/8 MiB per viewer cache; retain 12 Money registers and 2 month views; prefetch four previewable viewer neighbours and three Money accounts; cap queued Money navigation at 512 targets. Extend the existing registered Mod+Alt+1–9 bindings for all tray modes.

Known gaps: the full browser/viewport/cadence matrix and a before baseline need a perf-VM run. The focused local profiles still show incomplete frames and long tasks; no claim of a zero-incomplete pass is made. No Rust crate or server route changed, so Rust or adversarial API gates were not applicable. cargo clean removed 7,237 files (4.6 GiB), and apps/web/build was removed after captures. Worktree is clean.

Finished #641 on `job/blaze-surfaces` at head `30f644436262f125d7ecbb5c246d85a0eb5fe52c`. Built: - Files and Photos now prefetch full images and neighbour previews through bounded decoded-image LRUs. ImageView waits for `img.decode()` before replacing the preview. Files also prefetches text and PDF neighbours through bounded caches. - Money now coalesces and caches register data, warms three accounts in travel direction, and queues held arrow navigation with a 512-target cap. - Tabs now register shortcuts for modes 1–9 and start route navigation immediately while cache refresh runs in the background. - Fixed the shared screenshot harness to initialize the userStorage seam, and to open the Money sidebar only below the app's 768 px breakpoint. All 24 production-build screenshots are attached to this issue: Files, Photos, Money and Tabs at 390, 820 and 1440 px in light and dark. Performance results are focused local runs only: Chromium, 1440×900, 33 ms cadence, cold and warm. No before baseline exists in `docs/perf/baseline.json`. The perf VM at `root@10.69.69.63` was unreachable (`No route to host`), so these numbers are not a quiet-host comparison. | Surface | Phase | Incomplete / unpainted | p50 / p95 paint ms | CPU % | RSS peak KiB | Longest task ms | | --- | --- | ---: | ---: | ---: | ---: | ---: | | Files (1,200 mixed rows; 200 steps each way) | cold | 57 / 348 | 47.4 / 19,977.9 | 87.6 | 290,012 | 334 | | Files | warm | 115 / 387 | 23.9 / 56.7 | 19.5 | 313,876 | 1,690 | | Photos (520 items; 200 steps each way) | cold | 8 / 39 | 63.7 / 76,950.4 | 80.2 | 471,128 | 4,745 | | Photos | warm | 10 / 27 | 68.6 / 27,528.2 | 83.5 | 471,276 | 4,745 | | Money (220 accounts; 220 steps each way) | cold | 2 / 62 | 6,692.7 / 10,604.9 | 40.3 | 234,160 | 1,630 | | Money | warm | 3 / 74 | 18,274.5 / 26,751.5 | 20.4 | 234,244 | 1,630 | | Tabs (200 shortcuts) | cold | 282 / 75 | 1,138.1 / 5,787.8 | 21.0 | 210,984 | 597 | | Tabs | warm | 77 / 46 | 1,404.7 / 9,961.7 | 6.3 | 211,996 | 597 | The blaze zero-incomplete gate did not pass these local profiles. Load averages at run start were 9.6–24.4 across the runs, and the frame cadence was also delayed by long tasks. This does not replace the requested Chromium/WebKit × desktop/phone × cold/warm matrix on the perf VM. No valid before/after comparison is available. Web gate output (verbatim): ```text $ node scripts/check-user-storage.mjs && node scripts/check-type-tokens.mjs && node scripts/check-motion-tokens.mjs && svelte-kit sync && svelte-check --tsconfig ./tsconfig.json User browser caches use userStorage; only documented device/public-link exceptions remain. Text sizes and UI shape values use shared role tokens. UI transitions and animation options use shared motion tokens or documented exceptions. Loading svelte-check in workspace: /home/kayg/Developer/calternal-wt/blaze-surfaces/apps/web Getting Svelte diagnostics... svelte-check found 0 errors and 0 warnings ``` ```text $ vitest run Test Files 152 passed (152) Tests 1020 passed (1020) Start at 01:15:06 Duration 57.79s (transform 46%, environment 20%, import 18%, tests 12%, setup 4%) ``` Decisions not specified in DESIGN: use 8 decoded images, 5 PDF documents, and 8 text entries/8 MiB per viewer cache; retain 12 Money registers and 2 month views; prefetch four previewable viewer neighbours and three Money accounts; cap queued Money navigation at 512 targets. Extend the existing registered Mod+Alt+1–9 bindings for all tray modes. Known gaps: the full browser/viewport/cadence matrix and a before baseline need a perf-VM run. The focused local profiles still show incomplete frames and long tasks; no claim of a zero-incomplete pass is made. No Rust crate or server route changed, so Rust or adversarial API gates were not applicable. `cargo clean` removed 7,237 files (4.6 GiB), and `apps/web/build` was removed after captures. Worktree is clean.
Author
Owner

Correction to the final gate excerpt above: these result lines preserve Vitest's output exactly, including indentation:

svelte-check found 0 errors and 0 warnings
 Test Files  152 passed (152)
      Tests  1020 passed (1020)
   Start at  01:15:06
   Duration  57.79s (transform 46%, environment 20%, import 18%, tests 12%, setup 4%)
Correction to the final gate excerpt above: these result lines preserve Vitest's output exactly, including indentation: ```text svelte-check found 0 errors and 0 warnings Test Files 152 passed (152) Tests 1020 passed (1020) Start at 01:15:06 Duration 57.79s (transform 46%, environment 20%, import 18%, tests 12%, setup 4%) ```
Author
Owner

Columns now passes WebKit correctness after repeated-selection route updates were deferred until key release. The opt-in BLAZE_TRACE_VIRTUAL=1 trace recorded WebKit's selected index, reader ID, mounted row range, active row, scroll position and pathname. Before the fix, the selected index advanced while the path/body stayed on an earlier thread and the mounted range stopped moving. After the fix, all four WebKit Columns passes report 0 mismatches and 0 incomplete frames (cold/warm at 33 ms and 15 ms). The local diagnostic run was under shared-host load average 20–25 and tracing was enabled, so I will use separate uninstrumented runs for the report's timing table.

Columns now passes WebKit correctness after repeated-selection route updates were deferred until key release. The opt-in `BLAZE_TRACE_VIRTUAL=1` trace recorded WebKit's selected index, reader ID, mounted row range, active row, scroll position and pathname. Before the fix, the selected index advanced while the path/body stayed on an earlier thread and the mounted range stopped moving. After the fix, all four WebKit Columns passes report 0 mismatches and 0 incomplete frames (cold/warm at 33 ms and 15 ms). The local diagnostic run was under shared-host load average 20–25 and tracing was enabled, so I will use separate uninstrumented runs for the report's timing table.
Author
Owner

#641 production traversal — local

Host calternal-dev, load average 21.48, 19.99, 17.97; Chromium, 1440×900, pinned to CPUs 0–1; 200 ArrowDown then 200 ArrowUp events per pass.

Phase Repeat ms Keydowns Mismatch frames Incomplete frames Unpainted steps Longest task ms p95 frame ms
cold 33 400 0 38 0 1089 483.3
warm 33 400 0 0 0 197 433.3
cold 15 400 0 39 0 203 350
warm 15 400 0 0 0 112 333.3

All 400 key events were accounted for in each pass. The 20 route changes in each direction received their own complete painted frame. At each end of the list, boundary repeats kept the same complete section selected and did not enqueue a new route.

Per-route transition latency (input to its complete frame)

Step Key Selected section Cold 33 ms Warm 33 ms Cold 15 ms Warm 15 ms
1 ArrowDown apps 622.2 ms 297.2 ms 1227.6 ms 299.2 ms
2 ArrowDown maintenance 1523.7 ms 936.1 ms 1599.2 ms 703.3 ms
3 ArrowDown appearance 1812.7 ms 1720.8 ms 1895.8 ms 1178.6 ms
4 ArrowDown editor 2573.6 ms 2384.8 ms 2916.5 ms 1594.0 ms
5 ArrowDown notifications 4185.7 ms 3041.9 ms 3635.4 ms 2057.5 ms
6 ArrowDown calendars 5960.6 ms 3988.5 ms 4498.2 ms 2497.3 ms
7 ArrowDown mail 6791.4 ms 4575.4 ms 5094.6 ms 2989.8 ms
8 ArrowDown ai 8064.2 ms 5063.2 ms 5660.7 ms 3351.5 ms
9 ArrowDown photos 8356.9 ms 5606.0 ms 6205.5 ms 3865.3 ms
10 ArrowDown files 11270.4 ms 5953.3 ms 6837.1 ms 4397.5 ms
11 ArrowDown plugins 12147.3 ms 6613.8 ms 7719.6 ms 4805.1 ms
12 ArrowDown admin/users 12958.5 ms 6942.3 ms 8351.1 ms 5336.5 ms
13 ArrowDown admin/invitations 13836.9 ms 7986.2 ms 8929.2 ms 5549.1 ms
14 ArrowDown admin/sign-in 14401.8 ms 8493.6 ms 9692.2 ms 5948.1 ms
15 ArrowDown admin/configuration 15061.5 ms 9158.2 ms 10088.6 ms 6512.4 ms
16 ArrowDown admin/backups 16699.5 ms 9630.8 ms 10554.5 ms 6805.0 ms
17 ArrowDown admin/maintenance 17412.7 ms 10044.9 ms 11066.5 ms 7616.1 ms
18 ArrowDown admin/apps 18235.2 ms 11071.1 ms 11896.1 ms 8346.0 ms
19 ArrowDown admin/plugins 18857.6 ms 11589.3 ms 12113.2 ms 8924.7 ms
20 ArrowDown admin/system 19705.4 ms 12344.8 ms 13010.4 ms 9451.5 ms
201 ArrowUp admin/plugins 9857.9 ms 5150.7 ms 9276.5 ms 5733.8 ms
202 ArrowUp admin/apps 10353.4 ms 6164.8 ms 10122.7 ms 6396.5 ms
203 ArrowUp admin/maintenance 11126.0 ms 6614.8 ms 10814.7 ms 6871.1 ms
204 ArrowUp admin/backups 12092.8 ms 7638.4 ms 11395.0 ms 7419.2 ms
205 ArrowUp admin/configuration 13017.0 ms 8385.4 ms 12013.1 ms 7897.4 ms
206 ArrowUp admin/sign-in 13339.0 ms 9345.4 ms 12720.6 ms 8796.5 ms
207 ArrowUp admin/invitations 14074.2 ms 10613.6 ms 13433.1 ms 9211.2 ms
208 ArrowUp admin/users 14961.3 ms 11131.7 ms 14180.6 ms 9624.7 ms
209 ArrowUp plugins 15844.7 ms 12225.0 ms 14746.8 ms 10004.2 ms
210 ArrowUp files 16487.9 ms 12435.7 ms 15096.0 ms 10282.3 ms
211 ArrowUp photos 17507.1 ms 13229.8 ms 15628.1 ms 11013.9 ms
212 ArrowUp ai 17933.0 ms 14041.9 ms 15921.9 ms 11917.3 ms
213 ArrowUp mail 18693.4 ms 15171.2 ms 16183.9 ms 12527.4 ms
214 ArrowUp calendars 20353.6 ms 15618.9 ms 16509.8 ms 13170.6 ms
215 ArrowUp notifications 21227.6 ms 16074.6 ms 16875.0 ms 13864.8 ms
216 ArrowUp editor 22009.5 ms 16379.3 ms 17317.3 ms 14419.0 ms
217 ArrowUp appearance 22662.6 ms 16889.7 ms 17833.2 ms 14875.8 ms
218 ArrowUp maintenance 23424.0 ms 17282.8 ms 18344.3 ms 15277.7 ms
219 ArrowUp apps 24534.5 ms 18107.1 ms 18926.9 ms 16067.2 ms
220 ArrowUp account 25081.7 ms 19161.9 ms 19392.5 ms 16599.5 ms

Raw per-key records: settings-round2-final.json

The queue defect was concurrent consumers: a repeated key scheduled another consumer after the RAF handle cleared while the prior route callback still awaited content paint. An explicit in-flight guard now serializes the queue. Fix commit: b3fbeab57; current branch head: 98bcdcdc2.

Load caused slow frames in this local run; the non-SLOW correctness checks were clean: no selected/content mismatch and no unpainted route transition.

## #641 production traversal — local Host `calternal-dev`, load average `21.48, 19.99, 17.97`; Chromium, 1440×900, pinned to CPUs 0–1; 200 ArrowDown then 200 ArrowUp events per pass. | Phase | Repeat ms | Keydowns | Mismatch frames | Incomplete frames | Unpainted steps | Longest task ms | p95 frame ms | | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | | cold | 33 | 400 | 0 | 38 | 0 | 1089 | 483.3 | | warm | 33 | 400 | 0 | 0 | 0 | 197 | 433.3 | | cold | 15 | 400 | 0 | 39 | 0 | 203 | 350 | | warm | 15 | 400 | 0 | 0 | 0 | 112 | 333.3 | All 400 key events were accounted for in each pass. The 20 route changes in each direction received their own complete painted frame. At each end of the list, boundary repeats kept the same complete section selected and did not enqueue a new route. ### Per-route transition latency (input to its complete frame) | Step | Key | Selected section | Cold 33 ms | Warm 33 ms | Cold 15 ms | Warm 15 ms | | ---: | --- | --- | ---: | ---: | ---: | ---: | | 1 | ArrowDown | `apps` | 622.2 ms | 297.2 ms | 1227.6 ms | 299.2 ms | | 2 | ArrowDown | `maintenance` | 1523.7 ms | 936.1 ms | 1599.2 ms | 703.3 ms | | 3 | ArrowDown | `appearance` | 1812.7 ms | 1720.8 ms | 1895.8 ms | 1178.6 ms | | 4 | ArrowDown | `editor` | 2573.6 ms | 2384.8 ms | 2916.5 ms | 1594.0 ms | | 5 | ArrowDown | `notifications` | 4185.7 ms | 3041.9 ms | 3635.4 ms | 2057.5 ms | | 6 | ArrowDown | `calendars` | 5960.6 ms | 3988.5 ms | 4498.2 ms | 2497.3 ms | | 7 | ArrowDown | `mail` | 6791.4 ms | 4575.4 ms | 5094.6 ms | 2989.8 ms | | 8 | ArrowDown | `ai` | 8064.2 ms | 5063.2 ms | 5660.7 ms | 3351.5 ms | | 9 | ArrowDown | `photos` | 8356.9 ms | 5606.0 ms | 6205.5 ms | 3865.3 ms | | 10 | ArrowDown | `files` | 11270.4 ms | 5953.3 ms | 6837.1 ms | 4397.5 ms | | 11 | ArrowDown | `plugins` | 12147.3 ms | 6613.8 ms | 7719.6 ms | 4805.1 ms | | 12 | ArrowDown | `admin/users` | 12958.5 ms | 6942.3 ms | 8351.1 ms | 5336.5 ms | | 13 | ArrowDown | `admin/invitations` | 13836.9 ms | 7986.2 ms | 8929.2 ms | 5549.1 ms | | 14 | ArrowDown | `admin/sign-in` | 14401.8 ms | 8493.6 ms | 9692.2 ms | 5948.1 ms | | 15 | ArrowDown | `admin/configuration` | 15061.5 ms | 9158.2 ms | 10088.6 ms | 6512.4 ms | | 16 | ArrowDown | `admin/backups` | 16699.5 ms | 9630.8 ms | 10554.5 ms | 6805.0 ms | | 17 | ArrowDown | `admin/maintenance` | 17412.7 ms | 10044.9 ms | 11066.5 ms | 7616.1 ms | | 18 | ArrowDown | `admin/apps` | 18235.2 ms | 11071.1 ms | 11896.1 ms | 8346.0 ms | | 19 | ArrowDown | `admin/plugins` | 18857.6 ms | 11589.3 ms | 12113.2 ms | 8924.7 ms | | 20 | ArrowDown | `admin/system` | 19705.4 ms | 12344.8 ms | 13010.4 ms | 9451.5 ms | | 201 | ArrowUp | `admin/plugins` | 9857.9 ms | 5150.7 ms | 9276.5 ms | 5733.8 ms | | 202 | ArrowUp | `admin/apps` | 10353.4 ms | 6164.8 ms | 10122.7 ms | 6396.5 ms | | 203 | ArrowUp | `admin/maintenance` | 11126.0 ms | 6614.8 ms | 10814.7 ms | 6871.1 ms | | 204 | ArrowUp | `admin/backups` | 12092.8 ms | 7638.4 ms | 11395.0 ms | 7419.2 ms | | 205 | ArrowUp | `admin/configuration` | 13017.0 ms | 8385.4 ms | 12013.1 ms | 7897.4 ms | | 206 | ArrowUp | `admin/sign-in` | 13339.0 ms | 9345.4 ms | 12720.6 ms | 8796.5 ms | | 207 | ArrowUp | `admin/invitations` | 14074.2 ms | 10613.6 ms | 13433.1 ms | 9211.2 ms | | 208 | ArrowUp | `admin/users` | 14961.3 ms | 11131.7 ms | 14180.6 ms | 9624.7 ms | | 209 | ArrowUp | `plugins` | 15844.7 ms | 12225.0 ms | 14746.8 ms | 10004.2 ms | | 210 | ArrowUp | `files` | 16487.9 ms | 12435.7 ms | 15096.0 ms | 10282.3 ms | | 211 | ArrowUp | `photos` | 17507.1 ms | 13229.8 ms | 15628.1 ms | 11013.9 ms | | 212 | ArrowUp | `ai` | 17933.0 ms | 14041.9 ms | 15921.9 ms | 11917.3 ms | | 213 | ArrowUp | `mail` | 18693.4 ms | 15171.2 ms | 16183.9 ms | 12527.4 ms | | 214 | ArrowUp | `calendars` | 20353.6 ms | 15618.9 ms | 16509.8 ms | 13170.6 ms | | 215 | ArrowUp | `notifications` | 21227.6 ms | 16074.6 ms | 16875.0 ms | 13864.8 ms | | 216 | ArrowUp | `editor` | 22009.5 ms | 16379.3 ms | 17317.3 ms | 14419.0 ms | | 217 | ArrowUp | `appearance` | 22662.6 ms | 16889.7 ms | 17833.2 ms | 14875.8 ms | | 218 | ArrowUp | `maintenance` | 23424.0 ms | 17282.8 ms | 18344.3 ms | 15277.7 ms | | 219 | ArrowUp | `apps` | 24534.5 ms | 18107.1 ms | 18926.9 ms | 16067.2 ms | | 220 | ArrowUp | `account` | 25081.7 ms | 19161.9 ms | 19392.5 ms | 16599.5 ms | Raw per-key records: [settings-round2-final.json](https://git.kayg.org/attachments/41ea7efc-2512-4b4c-aa6c-cd1002794350) The queue defect was concurrent consumers: a repeated key scheduled another consumer after the RAF handle cleared while the prior route callback still awaited content paint. An explicit in-flight guard now serializes the queue. Fix commit: `b3fbeab57`; current branch head: `98bcdcdc2`. Load caused slow frames in this local run; the non-SLOW correctness checks were clean: no selected/content mismatch and no unpainted route transition.
Author
Owner

Round 2 #641 WebKit correction and final profile report

Root cause: selectRowAt called SvelteKit replaceState once per repeated Arrow key. On WebKit, the shallow route and reader update lagged behind the selected row. The optional BLAZE_TRACE_VIRTUAL=1 sampler recorded the selected index, reader ID, active row, mounted range, scroll position and pathname; the pathname/body stayed on an older thread while selection moved. Mail now follows every key repeat immediately and writes the stable thread URL once on key release. The frame bridge also avoids a synchronous layout read after same-document body writes; the existing frame ResizeObserver handles live body and image growth.

Correctness compared with the previous docs/perf/baseline.json profile (commit 88cb3c63d):

Layout / viewport Engine Mismatch / incomplete frames (current; prior max) Unpainted key steps (current; prior range)
Columns 1440×900 Chromium 0 / 0; prior 0 / 0 67–202; prior 181–264
Columns 1440×900 WebKit 0 / 0; prior 170 / 170 262–311; prior 250–306
Split 1440×900 Chromium 0 / 0; prior 0 / 0 176–227; prior 129–221
Split 1440×900 WebKit 0 / 0; prior 250 / 250 206–305; prior 271–356
Morph 1440×900 Chromium 0 / 0; prior 0 / 0 218–295; prior 198–279
Morph 820×1000 Chromium 0 / 0; prior 0 / 0 112–208; prior 166–235
Morph 1440×900 WebKit 0 / 0; prior 32 / 32 282–324; prior 283–353
Morph 820×1000 WebKit 0 / 0; prior 36 / 36 270–313; prior 269–330

Performance comparison (ranges across cold/warm 33 ms and 15 ms passes):

Layout / viewport Engine p50 paint ms (current; prior) p95 paint ms (current; prior) Server CPU mean ms (current; prior) Server RSS peak MiB (current; prior)
Columns 1440×900 Chromium 16.8–30.3; 35.2–49.2 10,552–17,444; 19,557–31,511 2,095; 1,373 223; 225
Columns 1440×900 WebKit 36–51; 59–89 14,578–25,972; 25,934–32,719 1,560; 1,270 228; 226
Split 1440×900 Chromium 25.3–43.3; 28.9–36.8 13,497–21,340; 16,831–22,316 1,260; 1,430 226; 228
Split 1440×900 WebKit 22–44; 42–71 9,342–19,813; 11,648–24,134 1,718; 1,225 220; 229
Morph 1440×900 Chromium 12.2–40.4; 18.3–3,924.8 8,677–23,090; 16,534–55,526 1,605; 1,465 227; 229
Morph 820×1000 Chromium 13.9–5,396.7; 31.9–39,268.1 10,883–30,645; 18,991–93,810 2,063; 1,515 221; 227
Morph 1440×900 WebKit 36–86; 40–66 16,715–31,386; 13,335–32,136 1,625; 1,090 229; 226
Morph 820×1000 WebKit 34–62; 44–92 14,949–26,479; 8,182–26,158 1,350; 920 227; 228

These are local measurements on the shared host, not a clean performance baseline. The current 1-minute load averages ranged from 13.6 to 33.1; docs/perf/baseline.json also marks the previous local Mail timing as unverified due to host load. The runner sends 400 Down and 400 Up inputs while loading 250 visible rows, so unpainted includes steps beyond a painted selection. WebKit does not report Long Tasks in this runner.

Profiles: artifacts/blaze/mail-columns-round2-final.json, artifacts/blaze/mail-split-round2-final.json, and artifacts/blaze/mail-morph-round2-final.json (local, ignored; not committed). All measured passes used a 10,000-message projection and 240 warmed bodies.

Round 2 #641 WebKit correction and final profile report Root cause: `selectRowAt` called SvelteKit `replaceState` once per repeated Arrow key. On WebKit, the shallow route and reader update lagged behind the selected row. The optional `BLAZE_TRACE_VIRTUAL=1` sampler recorded the selected index, reader ID, active row, mounted range, scroll position and pathname; the pathname/body stayed on an older thread while selection moved. Mail now follows every key repeat immediately and writes the stable thread URL once on key release. The frame bridge also avoids a synchronous layout read after same-document body writes; the existing frame ResizeObserver handles live body and image growth. Correctness compared with the previous `docs/perf/baseline.json` profile (commit 88cb3c63d): | Layout / viewport | Engine | Mismatch / incomplete frames (current; prior max) | Unpainted key steps (current; prior range) | |---|---|---:|---:| | Columns 1440×900 | Chromium | 0 / 0; prior 0 / 0 | 67–202; prior 181–264 | | Columns 1440×900 | WebKit | 0 / 0; prior 170 / 170 | 262–311; prior 250–306 | | Split 1440×900 | Chromium | 0 / 0; prior 0 / 0 | 176–227; prior 129–221 | | Split 1440×900 | WebKit | 0 / 0; prior 250 / 250 | 206–305; prior 271–356 | | Morph 1440×900 | Chromium | 0 / 0; prior 0 / 0 | 218–295; prior 198–279 | | Morph 820×1000 | Chromium | 0 / 0; prior 0 / 0 | 112–208; prior 166–235 | | Morph 1440×900 | WebKit | 0 / 0; prior 32 / 32 | 282–324; prior 283–353 | | Morph 820×1000 | WebKit | 0 / 0; prior 36 / 36 | 270–313; prior 269–330 | Performance comparison (ranges across cold/warm 33 ms and 15 ms passes): | Layout / viewport | Engine | p50 paint ms (current; prior) | p95 paint ms (current; prior) | Server CPU mean ms (current; prior) | Server RSS peak MiB (current; prior) | |---|---|---:|---:|---:|---:| | Columns 1440×900 | Chromium | 16.8–30.3; 35.2–49.2 | 10,552–17,444; 19,557–31,511 | 2,095; 1,373 | 223; 225 | | Columns 1440×900 | WebKit | 36–51; 59–89 | 14,578–25,972; 25,934–32,719 | 1,560; 1,270 | 228; 226 | | Split 1440×900 | Chromium | 25.3–43.3; 28.9–36.8 | 13,497–21,340; 16,831–22,316 | 1,260; 1,430 | 226; 228 | | Split 1440×900 | WebKit | 22–44; 42–71 | 9,342–19,813; 11,648–24,134 | 1,718; 1,225 | 220; 229 | | Morph 1440×900 | Chromium | 12.2–40.4; 18.3–3,924.8 | 8,677–23,090; 16,534–55,526 | 1,605; 1,465 | 227; 229 | | Morph 820×1000 | Chromium | 13.9–5,396.7; 31.9–39,268.1 | 10,883–30,645; 18,991–93,810 | 2,063; 1,515 | 221; 227 | | Morph 1440×900 | WebKit | 36–86; 40–66 | 16,715–31,386; 13,335–32,136 | 1,625; 1,090 | 229; 226 | | Morph 820×1000 | WebKit | 34–62; 44–92 | 14,949–26,479; 8,182–26,158 | 1,350; 920 | 227; 228 | These are local measurements on the shared host, not a clean performance baseline. The current 1-minute load averages ranged from 13.6 to 33.1; `docs/perf/baseline.json` also marks the previous local Mail timing as unverified due to host load. The runner sends 400 Down and 400 Up inputs while loading 250 visible rows, so `unpainted` includes steps beyond a painted selection. WebKit does not report Long Tasks in this runner. Profiles: `artifacts/blaze/mail-columns-round2-final.json`, `artifacts/blaze/mail-split-round2-final.json`, and `artifacts/blaze/mail-morph-round2-final.json` (local, ignored; not committed). All measured passes used a 10,000-message projection and 240 warmed bodies.
Author
Owner

Started #641 on job/blaze-surfaces at 30f644436262f125d7ecbb5c246d85a0eb5fe52c; baseline ref origin/dev is 687ff703136e71e89f8dfba139e93cd0788b25c. The worktree is clean. I have read CLAUDE.md, CONTEXT.md, and docs/DESIGN.md. I am validating the shared blaze harness and setting up interleaved baseline runs before making code changes.

Started #641 on `job/blaze-surfaces` at `30f644436262f125d7ecbb5c246d85a0eb5fe52c`; baseline ref `origin/dev` is `687ff703136e71e89f8dfba139e93cd0788b25c`. The worktree is clean. I have read `CLAUDE.md`, `CONTEXT.md`, and `docs/DESIGN.md`. I am validating the shared blaze harness and setting up interleaved baseline runs before making code changes.
Author
Owner

Final head after the documentation audit: 357e2a12c422d8fb948d5c890b338611951856fa.

Final verification after that documentation-only commit:

  • bun run check: svelte-check found 0 errors and 0 warnings
  • git diff --check: passed
  • cargo clean: Removed 7238 files, 4.6GiB total
  • Removed apps/web/build, apps/web/.svelte-kit, and target/tmp after verification.

The production measurements and attached screenshots in the report above remain unchanged. Known gap: the Settings warm-open target (≤100 ms p95 / about 50 ms main-thread work) is still not met in the local pinned run; the perf VM was offline and the local host load was high. No Rust source changed.

Final head after the documentation audit: `357e2a12c422d8fb948d5c890b338611951856fa`. Final verification after that documentation-only commit: - `bun run check`: `svelte-check found 0 errors and 0 warnings` - `git diff --check`: passed - `cargo clean`: `Removed 7238 files, 4.6GiB total` - Removed `apps/web/build`, `apps/web/.svelte-kit`, and `target/tmp` after verification. The production measurements and attached screenshots in the report above remain unchanged. Known gap: the Settings warm-open target (≤100 ms p95 / about 50 ms main-thread work) is still not met in the local pinned run; the perf VM was offline and the local host load was high. No Rust source changed.
Author
Owner

Baseline harness finding: the first origin/dev Photos run failed before key sampling. waitForSurfaceComplete timed out after 60 s because the runner queried .ql[data-viewer-key] and .ql-stage[data-content-key], which do not exist in the origin/dev QuickLook markup. That revision exposes the item name in .ql-title h2 and .image-view[aria-label]; its image decoder is .image-view img.decoder. No baseline timing was recorded. I am updating the shared sampler to read these real accessible values and decode state in both revisions, then I will rebuild the same baseline SPA and restart A/B/A/B.

Baseline harness finding: the first `origin/dev` Photos run failed before key sampling. `waitForSurfaceComplete` timed out after 60 s because the runner queried `.ql[data-viewer-key]` and `.ql-stage[data-content-key]`, which do not exist in the `origin/dev` QuickLook markup. That revision exposes the item name in `.ql-title h2` and `.image-view[aria-label]`; its image decoder is `.image-view img.decoder`. No baseline timing was recorded. I am updating the shared sampler to read these real accessible values and decode state in both revisions, then I will rebuild the same baseline SPA and restart A/B/A/B.
Author
Owner

origin/dev baseline A1 is measured with the same production build mode and 520-photo fixture. The perf VM probe still returns No route to host, so this is local. Profile: Chromium, 1440×900, 33 ms, 520 Down + 520 Up per cold/warm pass. The runner wrote artifacts/blaze/ab/photos-A1.json and exited non-zero on the expected blaze gate.

Phase Mismatch Incomplete Unpainted p50 paint ms p95 paint ms p95 key ms Max key ms CPU % RSS peak KiB Longest task ms
cold 0 1607 1031 224.2 47617.4 0.1 5.2 54.9 475668 4136
warm 0 1682 1029 187.8 93958.8 0.1 5.9 6.0 490296 801

This confirms the profile is not a 16 ms key-handler bottleneck on origin/dev; its content readiness and full traversal still fail badly. This is one run only. I am continuing the required interleaved comparison before changing app code.

`origin/dev` baseline A1 is measured with the same production build mode and 520-photo fixture. The perf VM probe still returns `No route to host`, so this is local. Profile: Chromium, 1440×900, 33 ms, 520 Down + 520 Up per cold/warm pass. The runner wrote `artifacts/blaze/ab/photos-A1.json` and exited non-zero on the expected blaze gate. | Phase | Mismatch | Incomplete | Unpainted | p50 paint ms | p95 paint ms | p95 key ms | Max key ms | CPU % | RSS peak KiB | Longest task ms | | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | | cold | 0 | 1607 | 1031 | 224.2 | 47617.4 | 0.1 | 5.2 | 54.9 | 475668 | 4136 | | warm | 0 | 1682 | 1029 | 187.8 | 93958.8 | 0.1 | 5.9 | 6.0 | 490296 | 801 | This confirms the profile is not a 16 ms key-handler bottleneck on `origin/dev`; its content readiness and full traversal still fail badly. This is one run only. I am continuing the required interleaved comparison before changing app code.
Author
Owner

Working on the #641 Settings slice in the #642 Round 3 branch job/blaze-settings, based on origin/dev at 687ff703136e71e89f8dfba139e93cd0788b25c1 (current head 357e2a12c422d8fb948d5c890b338611951856fa). I am checking whether queued route waits let repeated ArrowDown/ArrowUp inputs outrun the sidebar highlight; I will post the 15 ms frame table with the fix results.

Working on the #641 Settings slice in the #642 Round 3 branch `job/blaze-settings`, based on `origin/dev` at `687ff703136e71e89f8dfba139e93cd0788b25c1` (current head `357e2a12c422d8fb948d5c890b338611951856fa`). I am checking whether queued route waits let repeated ArrowDown/ArrowUp inputs outrun the sidebar highlight; I will post the 15 ms frame table with the fix results.
Author
Owner

Photos A/B round 1 (local, Chromium 1440×900, 33 ms, 520 steps each direction, production SPA, same fixture). Warm results: origin/dev A1 → this branch B1: incomplete frames 1682 → 39; unpainted steps 1029 → 115; p95 paint 93,958.8 ms → 485.8 ms; longest task 801 ms → 15,076 ms. The local 1/5/15-minute load averages were 29.13/22.86/17.95 for A1 and 26.64/28.11/24.05 for B1. Cold B1 also recorded 275,091.7 ms p95 paint and a 7,122 ms task, versus A1's 47,617.4 ms and 4,136 ms. These are single-run shared-host numbers: the complete-frame rate improved on B1, while its long task and cold tail are a possible regression that needs the remaining alternating runs and the requested trace before any app change. JSONs: photos-A1.json, photos-B1.json under ignored artifacts/blaze/ab/.

Photos A/B round 1 (local, Chromium 1440×900, 33 ms, 520 steps each direction, production SPA, same fixture). Warm results: `origin/dev` A1 → this branch B1: incomplete frames 1682 → 39; unpainted steps 1029 → 115; p95 paint 93,958.8 ms → 485.8 ms; longest task 801 ms → 15,076 ms. The local 1/5/15-minute load averages were 29.13/22.86/17.95 for A1 and 26.64/28.11/24.05 for B1. Cold B1 also recorded 275,091.7 ms p95 paint and a 7,122 ms task, versus A1's 47,617.4 ms and 4,136 ms. These are single-run shared-host numbers: the complete-frame rate improved on B1, while its long task and cold tail are a possible regression that needs the remaining alternating runs and the requested trace before any app change. JSONs: `photos-A1.json`, `photos-B1.json` under ignored `artifacts/blaze/ab/`.
Author
Owner

Photos A2 is complete on the same local Chromium 1440x900 / 33 ms / 520-item fixture. Load average was 21.53–23.86 cold and 24.01–26.70 warm.

Revision/run Phase Incomplete frames Unpainted steps p95 paint Longest task
origin/dev A1 cold 1,607 1,031 47,617.4 ms 4,136 ms
origin/dev A1 warm 1,682 1,029 93,958.8 ms 801 ms
branch B1 cold 17 156 275,091.7 ms 7,122 ms
branch B1 warm 39 115 485.8 ms 15,076 ms
origin/dev A2 cold 1,793 1,029 147,526.8 ms 633 ms
origin/dev A2 warm 1,904 1,021 224,512.8 ms 835 ms

There are no identity mismatches. The branch warm paint tail is far lower than both dev samples, but B1 has a 15.076 s warm task versus 0.801–0.835 s on dev. Cold p95 varies greatly between dev repeats, so it is not stable enough yet to classify as a branch regression. These are single-host A/B1/A2 observations; continuing alternating runs and tracing the warm main-thread work before changing app code.

Photos A2 is complete on the same local Chromium 1440x900 / 33 ms / 520-item fixture. Load average was 21.53–23.86 cold and 24.01–26.70 warm. | Revision/run | Phase | Incomplete frames | Unpainted steps | p95 paint | Longest task | | --- | --- | ---: | ---: | ---: | ---: | | origin/dev A1 | cold | 1,607 | 1,031 | 47,617.4 ms | 4,136 ms | | origin/dev A1 | warm | 1,682 | 1,029 | 93,958.8 ms | 801 ms | | branch B1 | cold | 17 | 156 | 275,091.7 ms | 7,122 ms | | branch B1 | warm | 39 | 115 | 485.8 ms | 15,076 ms | | origin/dev A2 | cold | 1,793 | 1,029 | 147,526.8 ms | 633 ms | | origin/dev A2 | warm | 1,904 | 1,021 | 224,512.8 ms | 835 ms | There are no identity mismatches. The branch warm paint tail is far lower than both dev samples, but B1 has a 15.076 s warm task versus 0.801–0.835 s on dev. Cold p95 varies greatly between dev repeats, so it is not stable enough yet to classify as a branch regression. These are single-host A/B1/A2 observations; continuing alternating runs and tracing the warm main-thread work before changing app code.
Author
Owner

Photos B2 (branch), same local Chromium 1440x900 / 33 ms / 520-item fixture: load average was 24.08 cold and 19.36 warm. Cold: 9 incomplete frames, 174 unpainted steps, p95 paint 232,534.8 ms, longest task 6,808 ms. Warm: 2 incomplete frames, 148 unpainted steps, p95 paint 130,977.1 ms, longest task 6,808 ms. No identity mismatches; max key handler slice was 7.4 ms cold / 2.1 ms warm.

This is still a large warm completeness improvement over dev A1/A2 (1,682/1,904 incomplete frames), but it misses the zero-incomplete/zero-unpainted and <=50 ms target. Warm p95 differs sharply from B1's 485.8 ms, so the tail is not yet repeatable. B1 and B2 both show multi-second long tasks (15,076 ms and 6,808 ms) versus 801–835 ms on A1/A2. Continuing the third pair and then tracing before touching app code.

Photos B2 (branch), same local Chromium 1440x900 / 33 ms / 520-item fixture: load average was 24.08 cold and 19.36 warm. Cold: 9 incomplete frames, 174 unpainted steps, p95 paint 232,534.8 ms, longest task 6,808 ms. Warm: 2 incomplete frames, 148 unpainted steps, p95 paint 130,977.1 ms, longest task 6,808 ms. No identity mismatches; max key handler slice was 7.4 ms cold / 2.1 ms warm. This is still a large warm completeness improvement over dev A1/A2 (1,682/1,904 incomplete frames), but it misses the zero-incomplete/zero-unpainted and <=50 ms target. Warm p95 differs sharply from B1's 485.8 ms, so the tail is not yet repeatable. B1 and B2 both show multi-second long tasks (15,076 ms and 6,808 ms) versus 801–835 ms on A1/A2. Continuing the third pair and then tracing before touching app code.
Author
Owner

Photos A/B/A/B/A/B is complete. All six runs use Chromium 1440x900, 33 ms cadence, 520 fixture photos, 520 steps each direction per cold/warm pass, production web bundles and the same local server binary. Measurements are local; the perf VM is unreachable. There were no selected/content identity mismatches.

Run Rev Phase Incomplete frames Unpainted steps p95 paint Longest task Load avg
A1 dev cold 1,607 1,031 47,617.4 ms 4,136 ms 21.34 / 18.67 / 15.58
B1 branch cold 17 156 275,091.7 ms 7,122 ms 32.45 / 27.70 / 22.09
A2 dev cold 1,793 1,029 147,526.8 ms 633 ms 21.53 / 23.86 / 23.59
B2 branch cold 9 174 232,534.8 ms 6,808 ms 24.08 / 23.94 / 23.91
A3 dev cold 1,899 1,027 94,988.6 ms 1,106 ms 18.17 / 22.08 / 23.14
B3 branch cold 12 100 158,055.1 ms 12,852 ms 17.62 / 18.91 / 21.34
A1 dev warm 1,682 1,029 93,958.8 ms 801 ms 29.13 / 22.86 / 17.95
B1 branch warm 39 115 485.8 ms 15,076 ms 26.64 / 28.11 / 24.05
A2 dev warm 1,904 1,021 224,512.8 ms 835 ms 26.70 / 24.95 / 24.01
B2 branch warm 2 148 130,977.1 ms 6,808 ms 19.36 / 22.70 / 23.58
A3 dev warm 1,940 1,012 7,441.7 ms 2,182 ms 18.90 / 20.49 / 22.36
B3 branch warm 0 17 149.8 ms 737 ms 21.72 / 19.87 / 21.13

Across three runs, dev has about 1,766 cold and 1,842 warm incomplete frames on average; the branch has 13 cold and 14 warm. The branch also reduces unpainted steps (warm mean 93 versus 1,021 on dev). Its median warm p95 paint is 485.8 ms versus 93,958.8 ms on dev. The measured regression is the branch's cold longest task in all three runs (6.8–12.9 s; dev 0.6–4.1 s), and two of three branch warm passes also exceed dev's 0.8–2.2 s range. Paint tails vary widely under host load. I will capture the requested warm Photos and Money traces and post attribution before changing application code.

Photos A/B/A/B/A/B is complete. All six runs use Chromium 1440x900, 33 ms cadence, 520 fixture photos, 520 steps each direction per cold/warm pass, production web bundles and the same local server binary. Measurements are local; the perf VM is unreachable. There were no selected/content identity mismatches. | Run | Rev | Phase | Incomplete frames | Unpainted steps | p95 paint | Longest task | Load avg | | --- | --- | --- | ---: | ---: | ---: | ---: | --- | | A1 | dev | cold | 1,607 | 1,031 | 47,617.4 ms | 4,136 ms | 21.34 / 18.67 / 15.58 | | B1 | branch | cold | 17 | 156 | 275,091.7 ms | 7,122 ms | 32.45 / 27.70 / 22.09 | | A2 | dev | cold | 1,793 | 1,029 | 147,526.8 ms | 633 ms | 21.53 / 23.86 / 23.59 | | B2 | branch | cold | 9 | 174 | 232,534.8 ms | 6,808 ms | 24.08 / 23.94 / 23.91 | | A3 | dev | cold | 1,899 | 1,027 | 94,988.6 ms | 1,106 ms | 18.17 / 22.08 / 23.14 | | B3 | branch | cold | 12 | 100 | 158,055.1 ms | 12,852 ms | 17.62 / 18.91 / 21.34 | | A1 | dev | warm | 1,682 | 1,029 | 93,958.8 ms | 801 ms | 29.13 / 22.86 / 17.95 | | B1 | branch | warm | 39 | 115 | 485.8 ms | 15,076 ms | 26.64 / 28.11 / 24.05 | | A2 | dev | warm | 1,904 | 1,021 | 224,512.8 ms | 835 ms | 26.70 / 24.95 / 24.01 | | B2 | branch | warm | 2 | 148 | 130,977.1 ms | 6,808 ms | 19.36 / 22.70 / 23.58 | | A3 | dev | warm | 1,940 | 1,012 | 7,441.7 ms | 2,182 ms | 18.90 / 20.49 / 22.36 | | B3 | branch | warm | 0 | 17 | 149.8 ms | 737 ms | 21.72 / 19.87 / 21.13 | Across three runs, dev has about 1,766 cold and 1,842 warm incomplete frames on average; the branch has 13 cold and 14 warm. The branch also reduces unpainted steps (warm mean 93 versus 1,021 on dev). Its median warm p95 paint is 485.8 ms versus 93,958.8 ms on dev. The measured regression is the branch's cold longest task in all three runs (6.8–12.9 s; dev 0.6–4.1 s), and two of three branch warm passes also exceed dev's 0.8–2.2 s range. Paint tails vary widely under host load. I will capture the requested warm Photos and Money traces and post attribution before changing application code.
Author
Owner

Chrome warm traces captured before application changes, using the production build with external source maps, Chromium 1440x900, 33 ms cadence, and local API-backed fixtures. Trace files: artifacts/blaze/traces-mapped/photos-warm.{profile,timeline}.json and artifacts/blaze/traces-mapped/money-warm.{profile,timeline}.json. Trace timing is attribution-only because profiling adds overhead.

Photos: 520 items; 1 ms V8 sampling; 182,537 samples / 296.6 s sampled.

Top self-time source Samples Self time
systemTimeZone() — packages/ui/src/time.ts:302 61,532 125.9 s
V8 program time 68,081 68.7 s
formatLongDate() — packages/ui/src/time.ts:232 8,552 17.6 s
downloadUrl() — apps/web/src/lib/files/api.ts:308 7,399 14.6 s
timeOptions() — packages/ui/src/time.ts:339 3,840 7.8 s
captureLabel() — apps/web/src/lib/photos/format.ts:35 2,999 6.4 s

PhotoTimeline calls captureLabel for tile accessibility labels. The sampled hot path repeatedly creates Intl.DateTimeFormat() to resolve the same system time zone. The timeline also records 110.7 s inclusive ImageDecodeTask time (largest 2.02 s); style 0.52 s; layout 1.09 s; main-thread Paint 1.07 s; RasterTask 17.6 s; and GC 18.5 s. The longest traced RunTask was 2.72 s. This points to timezone lookup as the main attributable JS cost, with image decode work also substantial; style/layout/paint are small by comparison.

Money: 220 accounts; 1 ms V8 sampling; 28,806 samples / 28.2 s sampled.

Top self-time source Samples Self time
V8 program time 18,675 16.0 s
V8 idle 5,697 7.8 s
native getBoundingClientRect 370 531 ms
native focus 607 477 ms
native fetch 133 255 ms
moveAccount() — apps/web/src/lib/components/SidebarLinks.svelte:91 38 45 ms
MoneySidebar rows() — apps/web/src/lib/components/money/MoneySidebar.svelte:40 41 41 ms

Money timeline totals: style 297 ms; layout 532 ms; paint/raster 8.79 s; GC 655 ms; longest traced RunTask 314 ms. The blaze Long Task observer recorded 617 ms in this traced pass. The profile does not attribute that task to one large application function; native focus and rectangle work are the largest identified per-event costs, and the route page/layout functions are each under 90 ms sampled self time. The current moveAccount still scans the entire sidebar list for each key, and drainAccountMoves still processes intermediate routes serially, so I will remove that unbounded per-step work as required. The trace does not establish the queue as the single source of the longest task.

Chrome warm traces captured before application changes, using the production build with external source maps, Chromium 1440x900, 33 ms cadence, and local API-backed fixtures. Trace files: `artifacts/blaze/traces-mapped/photos-warm.{profile,timeline}.json` and `artifacts/blaze/traces-mapped/money-warm.{profile,timeline}.json`. Trace timing is attribution-only because profiling adds overhead. **Photos: 520 items; 1 ms V8 sampling; 182,537 samples / 296.6 s sampled.** | Top self-time source | Samples | Self time | | --- | ---: | ---: | | `systemTimeZone()` — `packages/ui/src/time.ts:302` | 61,532 | 125.9 s | | V8 program time | 68,081 | 68.7 s | | `formatLongDate()` — `packages/ui/src/time.ts:232` | 8,552 | 17.6 s | | `downloadUrl()` — `apps/web/src/lib/files/api.ts:308` | 7,399 | 14.6 s | | `timeOptions()` — `packages/ui/src/time.ts:339` | 3,840 | 7.8 s | | `captureLabel()` — `apps/web/src/lib/photos/format.ts:35` | 2,999 | 6.4 s | `PhotoTimeline` calls `captureLabel` for tile accessibility labels. The sampled hot path repeatedly creates `Intl.DateTimeFormat()` to resolve the same system time zone. The timeline also records 110.7 s inclusive ImageDecodeTask time (largest 2.02 s); style 0.52 s; layout 1.09 s; main-thread Paint 1.07 s; RasterTask 17.6 s; and GC 18.5 s. The longest traced RunTask was 2.72 s. This points to timezone lookup as the main attributable JS cost, with image decode work also substantial; style/layout/paint are small by comparison. **Money: 220 accounts; 1 ms V8 sampling; 28,806 samples / 28.2 s sampled.** | Top self-time source | Samples | Self time | | --- | ---: | ---: | | V8 program time | 18,675 | 16.0 s | | V8 idle | 5,697 | 7.8 s | | native `getBoundingClientRect` | 370 | 531 ms | | native `focus` | 607 | 477 ms | | native `fetch` | 133 | 255 ms | | `moveAccount()` — `apps/web/src/lib/components/SidebarLinks.svelte:91` | 38 | 45 ms | | MoneySidebar `rows()` — `apps/web/src/lib/components/money/MoneySidebar.svelte:40` | 41 | 41 ms | Money timeline totals: style 297 ms; layout 532 ms; paint/raster 8.79 s; GC 655 ms; longest traced RunTask 314 ms. The blaze Long Task observer recorded 617 ms in this traced pass. The profile does not attribute that task to one large application function; native focus and rectangle work are the largest identified per-event costs, and the route page/layout functions are each under 90 ms sampled self time. The current `moveAccount` still scans the entire sidebar list for each key, and `drainAccountMoves` still processes intermediate routes serially, so I will remove that unbounded per-step work as required. The trace does not establish the queue as the single source of the longest task.
Author
Owner

Money A1 (origin/dev), local Chromium 1440x900 / 33 ms, 220-account fixture: cold had 368 mismatch and incomplete frames, 440 unpainted steps, no paint latency, and a 174 ms longest task (load average 7.49 / 13.39 / 16.85). Warm had 456 mismatch and incomplete frames, 440 unpainted steps, no paint latency, and a 174 ms longest task (load average 6.20 / 12.81 / 16.60).

The dev build does not move Money account selection with held ArrowDown/ArrowUp, so it never paints the expected account route. The branch comparison will include the new navigation behavior as well as its performance cost; these A1 values are not a same-behavior latency baseline. Continuing the alternating runs.

Money A1 (origin/dev), local Chromium 1440x900 / 33 ms, 220-account fixture: cold had 368 mismatch and incomplete frames, 440 unpainted steps, no paint latency, and a 174 ms longest task (load average 7.49 / 13.39 / 16.85). Warm had 456 mismatch and incomplete frames, 440 unpainted steps, no paint latency, and a 174 ms longest task (load average 6.20 / 12.81 / 16.60). The dev build does not move Money account selection with held ArrowDown/ArrowUp, so it never paints the expected account route. The branch comparison will include the new navigation behavior as well as its performance cost; these A1 values are not a same-behavior latency baseline. Continuing the alternating runs.
Author
Owner

Money A/B/A/B/A/B is complete. All six runs use local Chromium 1440x900, 33 ms cadence, the real API-backed 220-account/220-transaction fixture, production bundles and the same server binary. Load averages are recorded in each JSON; the perf VM is unreachable.

Run Rev Warm mismatches Warm incomplete Warm unpainted Warm p95 paint Longest task
A1 dev 456 456 440 n/a 174 ms
B1 branch 0 33 130 20,820.3 ms 697 ms
A2 dev 229 229 440 n/a 339 ms
B2 branch 0 40 152 54,570.1 ms 3,547 ms
A3 dev 325 325 440 n/a 784 ms
B3 branch 0 36 154 37,691.4 ms 989 ms

Dev does not switch Money accounts with held arrows, so all 440 steps are unpainted and its p95 is unavailable. The branch implements navigation with zero identity mismatches and reduces warm incomplete frames from 229–456 to 33–40, but it still misses 130–154 steps and leaves 36 incomplete frames on average. Branch warm p95 paint is 20.8–54.6 s. Its longest tasks were 697 ms, 3,547 ms and 989 ms (median 989 ms), versus 174 ms, 339 ms and 784 ms on dev (median 339 ms). The branch queue/navigation path remains over the <=50 ms target and must be fixed; comparison timing has the expected behavior difference.

Money A/B/A/B/A/B is complete. All six runs use local Chromium 1440x900, 33 ms cadence, the real API-backed 220-account/220-transaction fixture, production bundles and the same server binary. Load averages are recorded in each JSON; the perf VM is unreachable. | Run | Rev | Warm mismatches | Warm incomplete | Warm unpainted | Warm p95 paint | Longest task | | --- | --- | ---: | ---: | ---: | ---: | ---: | | A1 | dev | 456 | 456 | 440 | n/a | 174 ms | | B1 | branch | 0 | 33 | 130 | 20,820.3 ms | 697 ms | | A2 | dev | 229 | 229 | 440 | n/a | 339 ms | | B2 | branch | 0 | 40 | 152 | 54,570.1 ms | 3,547 ms | | A3 | dev | 325 | 325 | 440 | n/a | 784 ms | | B3 | branch | 0 | 36 | 154 | 37,691.4 ms | 989 ms | Dev does not switch Money accounts with held arrows, so all 440 steps are unpainted and its p95 is unavailable. The branch implements navigation with zero identity mismatches and reduces warm incomplete frames from 229–456 to 33–40, but it still misses 130–154 steps and leaves 36 incomplete frames on average. Branch warm p95 paint is 20.8–54.6 s. Its longest tasks were 697 ms, 3,547 ms and 989 ms (median 989 ms), versus 174 ms, 339 ms and 784 ms on dev (median 339 ms). The branch queue/navigation path remains over the <=50 ms target and must be fixed; comparison timing has the expected behavior difference.
Author
Owner

The first pinned Chromium repeat run failed before measurement with TypeError: Cannot read properties of undefined (reading 'indexOf') in summarizePass at bench/blaze.mjs:355. waitForPaintedSteps() returned samples and steps but omitted the sampled ids array that the summary indexes. I added ids to both the polling and final snapshots. The run was pinned to CPUs 0-1; host load was 30.99, 26.37, 22.97. I am rerunning once after this harness fix.

The first pinned Chromium repeat run failed before measurement with `TypeError: Cannot read properties of undefined (reading 'indexOf')` in `summarizePass` at `bench/blaze.mjs:355`. `waitForPaintedSteps()` returned samples and steps but omitted the sampled `ids` array that the summary indexes. I added `ids` to both the polling and final snapshots. The run was pinned to CPUs 0-1; host load was 30.99, 26.37, 22.97. I am rerunning once after this harness fix.
Author
Owner

Files A1/B1 first pair, local Chromium 1440x900 / 33 ms, full 1,200-entry fixture, production bundles. Neither run had selected/content identity mismatches.

Run Rev Phase Incomplete frames Unpainted steps p95 paint Longest task Load avg
A1 dev cold 1,993 1,694 16.4 ms 632 ms 27.61 / 25.12 / 22.11
B1 branch cold 1,259 961 17,119.1 ms 4,520 ms 28.43 / 27.59 / 23.91
A1 dev warm 546 996 21.1 ms 740 ms 26.36 / 25.17 / 22.51
B1 branch warm 324 929 264.4 ms 4,520 ms 31.84 / 30.01 / 25.38

The branch improves completeness in this pair, but its cold and warm paint tails and longest task regress. This is a single pair under high local load; continuing A2/B2/A3/B3 before classifying the result or changing app code.

Files A1/B1 first pair, local Chromium 1440x900 / 33 ms, full 1,200-entry fixture, production bundles. Neither run had selected/content identity mismatches. | Run | Rev | Phase | Incomplete frames | Unpainted steps | p95 paint | Longest task | Load avg | | --- | --- | --- | ---: | ---: | ---: | ---: | --- | | A1 | dev | cold | 1,993 | 1,694 | 16.4 ms | 632 ms | 27.61 / 25.12 / 22.11 | | B1 | branch | cold | 1,259 | 961 | 17,119.1 ms | 4,520 ms | 28.43 / 27.59 / 23.91 | | A1 | dev | warm | 546 | 996 | 21.1 ms | 740 ms | 26.36 / 25.17 / 22.51 | | B1 | branch | warm | 324 | 929 | 264.4 ms | 4,520 ms | 31.84 / 30.01 / 25.38 | The branch improves completeness in this pair, but its cold and warm paint tails and longest task regress. This is a single pair under high local load; continuing A2/B2/A3/B3 before classifying the result or changing app code.
Author
Owner

The corrected pinned Chromium probe completed data collection but failed the requested per-frame assertion. At 15 ms cadence on CPUs 0-1 (host load at launch: 32.30, 27.71, 23.71), the warm pass had 2 highlight-lag and 2 content-intent-lag frames out of 218 sampled warm frames. At 4119.3 ms, after key intent 203 expected admin/maintenance (index 17), both selected and content had fallen back to admin/backups (index 16); at 4352.6 ms, intent 208 expected admin/users (index 12), both had fallen back to plugins (index 11). The pane and highlight stayed in sync (mismatchFrames=0). This is consistent with overlapping SvelteKit goto calls landing out of order. I am changing the route writer to allow one in-flight goto and queue only the latest URL target; local highlight and pane selection still update in the key task.

The corrected pinned Chromium probe completed data collection but failed the requested per-frame assertion. At 15 ms cadence on CPUs 0-1 (host load at launch: 32.30, 27.71, 23.71), the warm pass had 2 highlight-lag and 2 content-intent-lag frames out of 218 sampled warm frames. At 4119.3 ms, after key intent 203 expected `admin/maintenance` (index 17), both selected and content had fallen back to `admin/backups` (index 16); at 4352.6 ms, intent 208 expected `admin/users` (index 12), both had fallen back to `plugins` (index 11). The pane and highlight stayed in sync (`mismatchFrames=0`). This is consistent with overlapping SvelteKit `goto` calls landing out of order. I am changing the route writer to allow one in-flight `goto` and queue only the latest URL target; local highlight and pane selection still update in the key task.
Author
Owner

The follow-up run after serializing route writes exposed a frame-probe clock flaw. At high load, installSampler() stored the requestAnimationFrame scheduled timestamp, but it read the DOM when the callback ran. A queued key task can run between those points: the frame sample then contains the new selected/content IDs while keysPressed excludes the key because its performance.now() time is newer than the old rAF timestamp. Example: in the warm pass the sample at 270.5 ms showed maintenance while only key 1 (apps) was counted; key 2 was recorded at 282.3 ms. The probe reported 11 warm mismatches from that combination. I am changing frame time to performance.now() inside the callback so the expected intent and DOM snapshot share one clock. This run is not valid proof of a product lag.

The follow-up run after serializing route writes exposed a frame-probe clock flaw. At high load, `installSampler()` stored the `requestAnimationFrame` scheduled timestamp, but it read the DOM when the callback ran. A queued key task can run between those points: the frame sample then contains the new selected/content IDs while `keysPressed` excludes the key because its `performance.now()` time is newer than the old rAF timestamp. Example: in the warm pass the sample at 270.5 ms showed `maintenance` while only key 1 (`apps`) was counted; key 2 was recorded at 282.3 ms. The probe reported 11 warm mismatches from that combination. I am changing frame time to `performance.now()` inside the callback so the expected intent and DOM snapshot share one clock. This run is not valid proof of a product lag.
Author
Owner

After the callback-clock correction, the 15 ms Chromium run (CPUs 0-1) reports zero warm-frame highlight/content lag across 447 frames, zero incomplete warm frames, p95 frame time 37.5 ms, and 44.1% Chromium CPU for the warm burst. The full 200 Down + 200 Up warm intent table is in artifacts/blaze/settings-e2e-641.json. The cold pass still failed: 3,157 of 3,158 sampled frames lagged, with a 1,504 ms longest task; the rail remained at Apps while its document-level input sampler recorded keys. I am adding focus state to the local frame evidence to distinguish a focus handoff from cold pane mounting. The cold result does not alter the warm-pass frame proof.

After the callback-clock correction, the 15 ms Chromium run (CPUs 0-1) reports zero warm-frame highlight/content lag across 447 frames, zero incomplete warm frames, p95 frame time 37.5 ms, and 44.1% Chromium CPU for the warm burst. The full 200 Down + 200 Up warm intent table is in `artifacts/blaze/settings-e2e-641.json`. The cold pass still failed: 3,157 of 3,158 sampled frames lagged, with a 1,504 ms longest task; the rail remained at Apps while its document-level input sampler recorded keys. I am adding focus state to the local frame evidence to distinguish a focus handoff from cold pane mounting. The cold result does not alter the warm-pass frame proof.
Author
Owner

Files A/B/A/B/A/B is complete. All runs use local Chromium 1440x900, 33 ms, full 1,200-entry mixed fixture, production bundles and the same server binary. No run had selected/content identity mismatches. Load average during measured phases ranged from 17.61 to 31.84 across the one- and five-minute readings; all numbers are local because the perf VM is unreachable.

Run Rev Phase Incomplete frames Unpainted steps p95 paint Longest task
A1 dev cold 1,993 1,694 16.4 ms 632 ms
B1 branch cold 1,259 961 17,119.1 ms 4,520 ms
A2 dev cold 285 993 203.3 ms 712 ms
B2 branch cold 484 910 8,065.6 ms 1,966 ms
A3 dev cold 1,879 995 26,127.8 ms 1,160 ms
B3 branch cold 211 893 32.8 ms 450 ms
A1 dev warm 546 996 21.1 ms 740 ms
B1 branch warm 324 929 264.4 ms 4,520 ms
A2 dev warm 449 992 117.4 ms 3,021 ms
B2 branch warm 185 902 35.6 ms 1,966 ms
A3 dev warm 277 994 16.1 ms 1,160 ms
B3 branch warm 197 916 55.2 ms 450 ms

Across the warm runs, branch incomplete frames average 235 versus 424 on dev, and unpainted steps average 916 versus 994. Branch median warm p95 paint is 55.2 ms versus 21.1 ms on dev; median longest task is 1,966 ms versus 1,160 ms. The branch therefore improves completeness while regressing the warm paint tail and task length across these three samples. The p95 and longest-task results vary with host load, so I will use traces and post-fix measurements to isolate the remaining work.

Files A/B/A/B/A/B is complete. All runs use local Chromium 1440x900, 33 ms, full 1,200-entry mixed fixture, production bundles and the same server binary. No run had selected/content identity mismatches. Load average during measured phases ranged from 17.61 to 31.84 across the one- and five-minute readings; all numbers are local because the perf VM is unreachable. | Run | Rev | Phase | Incomplete frames | Unpainted steps | p95 paint | Longest task | | --- | --- | --- | ---: | ---: | ---: | ---: | | A1 | dev | cold | 1,993 | 1,694 | 16.4 ms | 632 ms | | B1 | branch | cold | 1,259 | 961 | 17,119.1 ms | 4,520 ms | | A2 | dev | cold | 285 | 993 | 203.3 ms | 712 ms | | B2 | branch | cold | 484 | 910 | 8,065.6 ms | 1,966 ms | | A3 | dev | cold | 1,879 | 995 | 26,127.8 ms | 1,160 ms | | B3 | branch | cold | 211 | 893 | 32.8 ms | 450 ms | | A1 | dev | warm | 546 | 996 | 21.1 ms | 740 ms | | B1 | branch | warm | 324 | 929 | 264.4 ms | 4,520 ms | | A2 | dev | warm | 449 | 992 | 117.4 ms | 3,021 ms | | B2 | branch | warm | 185 | 902 | 35.6 ms | 1,966 ms | | A3 | dev | warm | 277 | 994 | 16.1 ms | 1,160 ms | | B3 | branch | warm | 197 | 916 | 55.2 ms | 450 ms | Across the warm runs, branch incomplete frames average 235 versus 424 on dev, and unpainted steps average 916 versus 994. Branch median warm p95 paint is 55.2 ms versus 21.1 ms on dev; median longest task is 1,966 ms versus 1,160 ms. The branch therefore improves completeness while regressing the warm paint tail and task length across these three samples. The p95 and longest-task results vary with host load, so I will use traces and post-fix measurements to isolate the remaining work.
Author
Owner

The latest raw frame data explains the cold-only mismatch. On the first cold key, the event target and active element were the Account rail row; subsequent events had no data-sidebar-id on the target or active element, and sampled frames had no focused rail ID. The warm pass kept focus and still had zero lag. OverlaySurface deliberately sets initial focus in a later frame (focusTrap with deferFocus), so the cold probe was sending keys before that focus handoff had settled. I added a two-rAF wait after the surface reports complete, then the probe focuses the first rail row. This aligns the cold input start with the real post-open focus state.

The latest raw frame data explains the cold-only mismatch. On the first cold key, the event target and active element were the Account rail row; subsequent events had no `data-sidebar-id` on the target or active element, and sampled frames had no focused rail ID. The warm pass kept focus and still had zero lag. OverlaySurface deliberately sets initial focus in a later frame (`focusTrap` with `deferFocus`), so the cold probe was sending keys before that focus handoff had settled. I added a two-rAF wait after the surface reports complete, then the probe focuses the first rail row. This aligns the cold input start with the real post-open focus state.
Author
Owner

Tabs cross-revision A/B/A/B/A/B comparison is complete with the fixed runner. Fixture and profile are unchanged: Chromium 1440x900, 33 ms cadence, 200 key steps, local production build, same server binary and fixtures. Medians across three runs per revision:

phase metric origin/dev branch
cold incomplete frames 105 93
cold steps never painted complete 77 32
cold p95 paint latency 5299.9 ms 4088.2 ms
cold longest task 224 ms 524 ms
warm incomplete frames 79 85
warm steps never painted complete 72 46
warm p95 paint latency 1893.2 ms 4760.1 ms
warm longest task 224 ms 545 ms

The branch paints more held-key steps eventually, but warm incomplete frames are slightly higher, warm p95 paint is about 2.5x the dev median, and both branch task medians exceed 50 ms. Neither revision meets the zero-incomplete, no-task-over-50-ms target. The A/B runner records 200 keydown handler samples per pass with max handler time at or below 9.2 ms; the long tasks occur outside the handler itself. Run artifacts are in the local ignored artifacts/blaze/ab/tabs-{A1,B1,A2,B2,A3,B3}-fixed.json files.

Tabs cross-revision A/B/A/B/A/B comparison is complete with the fixed runner. Fixture and profile are unchanged: Chromium 1440x900, 33 ms cadence, 200 key steps, local production build, same server binary and fixtures. Medians across three runs per revision: | phase | metric | origin/dev | branch | |---|---:|---:|---:| | cold | incomplete frames | 105 | 93 | | cold | steps never painted complete | 77 | 32 | | cold | p95 paint latency | 5299.9 ms | 4088.2 ms | | cold | longest task | 224 ms | 524 ms | | warm | incomplete frames | 79 | 85 | | warm | steps never painted complete | 72 | 46 | | warm | p95 paint latency | 1893.2 ms | 4760.1 ms | | warm | longest task | 224 ms | 545 ms | The branch paints more held-key steps eventually, but warm incomplete frames are slightly higher, warm p95 paint is about 2.5x the dev median, and both branch task medians exceed 50 ms. Neither revision meets the zero-incomplete, no-task-over-50-ms target. The A/B runner records 200 keydown handler samples per pass with max handler time at or below 9.2 ms; the long tasks occur outside the handler itself. Run artifacts are in the local ignored `artifacts/blaze/ab/tabs-{A1,B1,A2,B2,A3,B3}-fixed.json` files.
Author
Owner

Money retest after the first Round 2 fix (single local Chromium 1440x900, 33 ms pass; full 220-account fixture):

phase incomplete frames steps never painted complete p95 paint longest task
cold 262 415 14,365.2 ms 161 ms
warm 314 410 16,915.4 ms 186 ms

Compared with the earlier three-run branch median (warm p95 20.8–54.6 s; longest task median 989 ms; incomplete 33–40; unpainted 130–154), the task duration and paint tail improved, but complete-paint counts regressed. Source review found that the first fix cancels stale prefetch transaction reads, while selected route reads still have no AbortSignal. The page increments a request token to suppress stale UI publication, but that does not stop the underlying getTransactions request. I am testing cancellation of superseded active reads while keeping a promoted neighbor request alive.

Money retest after the first Round 2 fix (single local Chromium 1440x900, 33 ms pass; full 220-account fixture): | phase | incomplete frames | steps never painted complete | p95 paint | longest task | |---|---:|---:|---:|---:| | cold | 262 | 415 | 14,365.2 ms | 161 ms | | warm | 314 | 410 | 16,915.4 ms | 186 ms | Compared with the earlier three-run branch median (warm p95 20.8–54.6 s; longest task median 989 ms; incomplete 33–40; unpainted 130–154), the task duration and paint tail improved, but complete-paint counts regressed. Source review found that the first fix cancels stale *prefetch* transaction reads, while selected route reads still have no AbortSignal. The page increments a request token to suppress stale UI publication, but that does not stop the underlying getTransactions request. I am testing cancellation of superseded active reads while keeping a promoted neighbor request alive.
Author
Owner

Money retest after aborting superseded active register requests (single local Chromium 1440x900, 33 ms cadence; fixture creates all 220 accounts and transactions through the API):

phase incomplete frames steps never painted complete p95 paint longest task
cold 315 418 10,186.6 ms 1,397 ms
warm 284 421 18,361.8 ms 1,397 ms

This still misses the zero-incomplete / zero-unpainted / ≤50 ms target. Compared with the immediately prior one-run retest, warm p95 and longest task increased (16,915.4 → 18,361.8 ms; 186 → 1,397 ms), while incomplete frames fell slightly (314 → 284). The active-read cancellation did not restore complete paint; I will use the captured warm trace and route data to identify the task source before another change. Results are local because the perf VM is unreachable.

Money retest after aborting superseded active register requests (single local Chromium 1440x900, 33 ms cadence; fixture creates all 220 accounts and transactions through the API): | phase | incomplete frames | steps never painted complete | p95 paint | longest task | |---|---:|---:|---:|---:| | cold | 315 | 418 | 10,186.6 ms | 1,397 ms | | warm | 284 | 421 | 18,361.8 ms | 1,397 ms | This still misses the zero-incomplete / zero-unpainted / ≤50 ms target. Compared with the immediately prior one-run retest, warm p95 and longest task increased (16,915.4 → 18,361.8 ms; 186 → 1,397 ms), while incomplete frames fell slightly (314 → 284). The active-read cancellation did not restore complete paint; I will use the captured warm trace and route data to identify the task source before another change. Results are local because the perf VM is unreachable.
Author
Owner

Money follow-up trace (Chromium 1440 px, 33 ms cadence; local fixture). This traced warm pass had 100 incomplete frames, 430 unpainted steps, p95 paint 7,950 ms, and a 222 ms longest task. Key input had 0.1 ms p95 / 6.8 ms max. Of 26,337 CPU samples, 12,708 ms were (program) and 9,688 ms idle; the largest named browser self-time entries were querySelector (360 ms), focus (132 ms), resolve (111 ms), URL (100 ms), fetch (74 ms), and getBoundingClientRect (61 ms). Timeline totals: style/layout 232 ms, paint 6,028 ms (30 ms max), raster 259 ms, function calls 2,432 ms (46.9 ms max), and max RunTask 205 ms. This capture does not show a long account-navigation handler as the cause of the paint tail. Results still vary substantially across fixture runs; the paint-completion gap needs the next bounded-prefetch measurement.

Money follow-up trace (Chromium 1440 px, 33 ms cadence; local fixture). This traced warm pass had 100 incomplete frames, 430 unpainted steps, p95 paint 7,950 ms, and a 222 ms longest task. Key input had 0.1 ms p95 / 6.8 ms max. Of 26,337 CPU samples, 12,708 ms were `(program)` and 9,688 ms idle; the largest named browser self-time entries were querySelector (360 ms), focus (132 ms), resolve (111 ms), URL (100 ms), fetch (74 ms), and getBoundingClientRect (61 ms). Timeline totals: style/layout 232 ms, paint 6,028 ms (30 ms max), raster 259 ms, function calls 2,432 ms (46.9 ms max), and max RunTask 205 ms. This capture does not show a long account-navigation handler as the cause of the paint tail. Results still vary substantially across fixture runs; the paint-completion gap needs the next bounded-prefetch measurement.
Author
Owner

Money lookahead experiment (local Chromium 1440x900, 33 ms, full 220-account fixture): increasing the prefetch window from 3 to 11 produced 266 cold / 281 warm incomplete frames, 423 unpainted steps in both phases, warm p95 paint 24,942.3 ms, and a 3,973 ms longest task. The preceding full 3-neighbor run had 284 warm incomplete frames, 421 unpainted steps, p95 18,361.8 ms, and a 1,397 ms longest task. The wider window starts more real API work but does not improve the number of steps that paint; it increases the measured warm tail. I am restoring the three-neighbor cap and adding a test for that bound.

Money lookahead experiment (local Chromium 1440x900, 33 ms, full 220-account fixture): increasing the prefetch window from 3 to 11 produced 266 cold / 281 warm incomplete frames, 423 unpainted steps in both phases, warm p95 paint 24,942.3 ms, and a 3,973 ms longest task. The preceding full 3-neighbor run had 284 warm incomplete frames, 421 unpainted steps, p95 18,361.8 ms, and a 1,397 ms longest task. The wider window starts more real API work but does not improve the number of steps that paint; it increases the measured warm tail. I am restoring the three-neighbor cap and adding a test for that bound.
Author
Owner

Photos retest after caching the system time zone (local Chromium 1440x900, 33 ms, 520-image fixture): warm recorded 61 incomplete frames and 350 steps never painted complete; p95 paint was 97,037 ms and the longest task was 3,582 ms. Key handling stayed at 0.1 ms p95 / 7.5 ms max. The local server used about 54% CPU warm and peaked near 473 MiB RSS. This does not meet the warm target. The original trace's systemTimeZone() hotspot is removed by the cache, so I am capturing another warm trace to identify the remaining decode or rendering work before changing the viewer.

Photos retest after caching the system time zone (local Chromium 1440x900, 33 ms, 520-image fixture): warm recorded 61 incomplete frames and 350 steps never painted complete; p95 paint was 97,037 ms and the longest task was 3,582 ms. Key handling stayed at 0.1 ms p95 / 7.5 ms max. The local server used about 54% CPU warm and peaked near 473 MiB RSS. This does not meet the warm target. The original trace's `systemTimeZone()` hotspot is removed by the cache, so I am capturing another warm trace to identify the remaining decode or rendering work before changing the viewer.
Author
Owner

Photos warm source-mapped trace after the time-zone cache (Chromium 1440 px, 220 steps; trace overhead applies): 45.7 s of sampled self time was non-idle browser JS. The largest mapped app functions were formatLongDate in packages/ui/src/time.ts (7.74 s), timeOptions there (5.25 s), downloadUrl in apps/web/src/lib/files/api.ts (5.36 s), formatTime in packages/ui/src/time.ts (5.25 s in its own aggregate), and captureLabel in apps/web/src/lib/photos/format.ts (0.86 s). The Photos viewer maps all tiles to ViewerItem objects when its selected index changes; each pass formats capture labels and rebuilds download and thumbnail URLs across the full collection. This is O(collection size) per key and accounts for the repeated format/URL work. ImageDecodeTask totaled 19.35 s with a 755.5 ms maximum; Decode Image totaled 15.02 s (185.9 ms max); Decode LazyPixelRef totaled 16.39 s (436.1 ms max); RasterTask totaled 9.21 s (682.6 ms max); Paint 0.95 s (59.8 ms max); Layout 0.40 s (22.1 ms max); UpdateLayoutTree 0.29 s (18.9 ms max); GC sampled self was 1.20 s. I am removing the full-list rebuild per selection and will measure again.

Photos warm source-mapped trace after the time-zone cache (Chromium 1440 px, 220 steps; trace overhead applies): 45.7 s of sampled self time was non-idle browser JS. The largest mapped app functions were `formatLongDate` in `packages/ui/src/time.ts` (7.74 s), `timeOptions` there (5.25 s), `downloadUrl` in `apps/web/src/lib/files/api.ts` (5.36 s), `formatTime` in `packages/ui/src/time.ts` (5.25 s in its own aggregate), and `captureLabel` in `apps/web/src/lib/photos/format.ts` (0.86 s). The Photos viewer maps all tiles to `ViewerItem` objects when its selected index changes; each pass formats capture labels and rebuilds download and thumbnail URLs across the full collection. This is O(collection size) per key and accounts for the repeated format/URL work. ImageDecodeTask totaled 19.35 s with a 755.5 ms maximum; Decode Image totaled 15.02 s (185.9 ms max); Decode LazyPixelRef totaled 16.39 s (436.1 ms max); RasterTask totaled 9.21 s (682.6 ms max); Paint 0.95 s (59.8 ms max); Layout 0.40 s (22.1 ms max); UpdateLayoutTree 0.29 s (18.9 ms max); GC sampled self was 1.20 s. I am removing the full-list rebuild per selection and will measure again.
Author
Owner

Correction to the mapped Photos profile table above: formatTime() self time was 1.75 s (at packages/ui/src/time.ts:363); timeOptions() self time was 5.25 s (at packages/ui/src/time.ts:345).

Correction to the mapped Photos profile table above: `formatTime()` self time was 1.75 s (at `packages/ui/src/time.ts:363`); `timeOptions()` self time was 5.25 s (at `packages/ui/src/time.ts:345`).
Author
Owner

Photos retest after stabilizing the viewer-item collection (local Chromium 1440x900, 33 ms, full 520-photo fixture): warm recorded 260 incomplete frames, 890 unpainted steps, p95 paint 59,822.3 ms, 2,764 ms longest task, and key-handler max 5.7 ms. The preceding full fixture run had 61 incomplete frames / 350 unpainted steps, p95 97,037 ms and 3,582 ms longest task. The current warm server CPU was 67.7% versus 53.7% in the preceding pass, so a single local pair is too noisy to claim the incomplete count improved or regressed. Both runs miss the target. I will confirm the source-mapped JS profile after the collection change.

Photos retest after stabilizing the viewer-item collection (local Chromium 1440x900, 33 ms, full 520-photo fixture): warm recorded 260 incomplete frames, 890 unpainted steps, p95 paint 59,822.3 ms, 2,764 ms longest task, and key-handler max 5.7 ms. The preceding full fixture run had 61 incomplete frames / 350 unpainted steps, p95 97,037 ms and 3,582 ms longest task. The current warm server CPU was 67.7% versus 53.7% in the preceding pass, so a single local pair is too noisy to claim the incomplete count improved or regressed. Both runs miss the target. I will confirm the source-mapped JS profile after the collection change.
Author
Owner

Photos source-mapped follow-up trace after the collection cache (Chromium 1440 px, 220 steps; local, trace overhead applies): downloadUrl() self time fell from 5.36 s to 18.6 ms; formatLongDate() from 7.74 s to 14.2 ms; timeOptions() from 5.25 s to 16.3 ms; sampled GC from 1.20 s to 229 ms. The prior O(collection) URL and date work is gone from each selected-item update. The updated warm trace measured 80 incomplete frames, 314 unpainted steps, p95 paint 24,087 ms, and a 301 ms longest task; key-handler max was 11.4 ms.

The remaining long task aligns with one full layout of 654 render objects (288.8 ms, dirtyObjects: 1) immediately before the first recorded keydown. The timeline totals show image decoding on renderer worker threads: ImageDecodeTask 14.07 s total / 725 ms max, Decode Image 11.10 s / 133.5 ms max, Decode LazyPixelRef 11.56 s / 139.6 ms max, RasterTask 4.13 s / 516.7 ms max. Main-thread Paint totaled 400 ms (21.1 ms max), Layout 618 ms (288.8 ms max), and UpdateLayoutTree 202 ms (17.3 ms max). The remaining profile is not per-key JavaScript formatting or URL construction. I am retaining these as local-host limitations to recheck in the requested browser/viewport matrix; the warm target remains unmet.

Photos source-mapped follow-up trace after the collection cache (Chromium 1440 px, 220 steps; local, trace overhead applies): `downloadUrl()` self time fell from 5.36 s to 18.6 ms; `formatLongDate()` from 7.74 s to 14.2 ms; `timeOptions()` from 5.25 s to 16.3 ms; sampled GC from 1.20 s to 229 ms. The prior O(collection) URL and date work is gone from each selected-item update. The updated warm trace measured 80 incomplete frames, 314 unpainted steps, p95 paint 24,087 ms, and a 301 ms longest task; key-handler max was 11.4 ms. The remaining long task aligns with one full layout of 654 render objects (288.8 ms, `dirtyObjects: 1`) immediately before the first recorded keydown. The timeline totals show image decoding on renderer worker threads: ImageDecodeTask 14.07 s total / 725 ms max, Decode Image 11.10 s / 133.5 ms max, Decode LazyPixelRef 11.56 s / 139.6 ms max, RasterTask 4.13 s / 516.7 ms max. Main-thread Paint totaled 400 ms (21.1 ms max), Layout 618 ms (288.8 ms max), and UpdateLayoutTree 202 ms (17.3 ms max). The remaining profile is not per-key JavaScript formatting or URL construction. I am retaining these as local-host limitations to recheck in the requested browser/viewport matrix; the warm target remains unmet.
Author
Owner

Tabs A/B validity finding: the branch adds registered mode shortcuts 6–9 in +layout.svelte and TabBar.svelte; the compared dev revision registers only shortcuts 1–5. The runner sent digits 1–9 in both revisions, so dev ignored four of every nine mode requests. The earlier Tabs p95/task comparison therefore does not hold a same-workload baseline. I am adding a runner count option and will repeat the interleaved A/B runs over the five shortcuts supported by both revisions; the branch’s full nine-mode target will remain a separate current-build check.

Tabs A/B validity finding: the branch adds registered mode shortcuts 6–9 in `+layout.svelte` and `TabBar.svelte`; the compared dev revision registers only shortcuts 1–5. The runner sent digits 1–9 in both revisions, so dev ignored four of every nine mode requests. The earlier Tabs p95/task comparison therefore does not hold a same-workload baseline. I am adding a runner count option and will repeat the interleaved A/B runs over the five shortcuts supported by both revisions; the branch’s full nine-mode target will remain a separate current-build check.
Author
Owner

Corrected Tabs baseline (same workload)

The earlier Tabs A/B comparison pressed nine keys on both revisions, but origin/dev only registers five shortcuts. That comparison was not equivalent. I reran an interleaved A/B/A/B/A/B profile at Chromium 1440×900, 33 ms, 200 steps, with BLAZE_TAB_MODE_COUNT=5 on both revisions and the same local server, build mode, and fixture.

Run Revision Warm incomplete frames Warm unpainted steps Warm p95 paint Warm longest task
A1 origin/dev 85 13 4,685 ms 1,336 ms
B1 branch 66 57 3,462.2 ms 241 ms
A2 origin/dev 64 26 7,037 ms 313 ms
B2 branch 60 60 6,625.7 ms 521 ms
A3 origin/dev 69 63 3,725.3 ms 363 ms
B3 branch 100 6 3,093.6 ms 785 ms
Median origin/dev 69 26 4,685 ms 363 ms
Median branch 66 57 3,462 ms 521 ms

The branch median does not show a regression in incomplete frames or p95 for this common five-key workload. Unpainted steps and longest-task median are worse on the branch, with substantial run-to-run spread. Both revisions miss the target of zero incomplete frames and no task over 50 ms. These local results do not support attributing the original apparent regression to the branch. The branch’s separate nine-shortcut trace remains useful for current-build attribution; I will report it separately from this baseline comparison.

## Corrected Tabs baseline (same workload) The earlier Tabs A/B comparison pressed nine keys on both revisions, but `origin/dev` only registers five shortcuts. That comparison was not equivalent. I reran an interleaved A/B/A/B/A/B profile at Chromium 1440×900, 33 ms, 200 steps, with `BLAZE_TAB_MODE_COUNT=5` on both revisions and the same local server, build mode, and fixture. | Run | Revision | Warm incomplete frames | Warm unpainted steps | Warm p95 paint | Warm longest task | |---|---|---:|---:|---:|---:| | A1 | origin/dev | 85 | 13 | 4,685 ms | 1,336 ms | | B1 | branch | 66 | 57 | 3,462.2 ms | 241 ms | | A2 | origin/dev | 64 | 26 | 7,037 ms | 313 ms | | B2 | branch | 60 | 60 | 6,625.7 ms | 521 ms | | A3 | origin/dev | 69 | 63 | 3,725.3 ms | 363 ms | | B3 | branch | 100 | 6 | 3,093.6 ms | 785 ms | | Median | origin/dev | 69 | 26 | 4,685 ms | 363 ms | | Median | branch | 66 | 57 | 3,462 ms | 521 ms | The branch median does not show a regression in incomplete frames or p95 for this common five-key workload. Unpainted steps and longest-task median are worse on the branch, with substantial run-to-run spread. Both revisions miss the target of zero incomplete frames and no task over 50 ms. These local results do not support attributing the original apparent regression to the branch. The branch’s separate nine-shortcut trace remains useful for current-build attribution; I will report it separately from this baseline comparison.
Author
Owner

#641 closeout for job/blaze-settings, head bf3ad5f29d1108c34e87c0e9d20b1dbe67ffc5c3.

The benchmark keeps rail selection synchronous with repeated keys, renders the content for the latest selection per frame, and serializes only route writes. The corrected sampler reads both key intent and DOM state on the same callback clock.

At 15 ms key repeat in Chromium pinned to CPUs 0–1, the warm traversal sampled 398 frames: highlight lag 0, content lag 0, mismatch 0 and incomplete frames 0. The cold traversal still had 21 incomplete frames. The detailed run table:

Phase Viewport Cadence Mismatch Highlight lag Content lag Incomplete Unpainted p95 step p95 frame Chromium CPU Server RSS
Cold 1440×900 15 ms 0 0 0 21 218 56.1 ms 91.4 ms 56.1% 153332 KiB
Warm 1440×900 15 ms 0 0 0 0 217 88.1 ms 65.2 ms 36.3% 217524 KiB

At this cadence, keys can arrive faster than display frames; the frame assertion stayed aligned to the latest intent on every warm frame. The cold first traversal does not meet the no-incomplete-frame target.

Shared gate output and the Notes suite exception are recorded in the final #642 comment. The production captures for Settings are attached there at phone, tablet and desktop widths in light and dark.

#641 closeout for `job/blaze-settings`, head `bf3ad5f29d1108c34e87c0e9d20b1dbe67ffc5c3`. The benchmark keeps rail selection synchronous with repeated keys, renders the content for the latest selection per frame, and serializes only route writes. The corrected sampler reads both key intent and DOM state on the same callback clock. At 15 ms key repeat in Chromium pinned to CPUs 0–1, the warm traversal sampled 398 frames: highlight lag 0, content lag 0, mismatch 0 and incomplete frames 0. The cold traversal still had 21 incomplete frames. The detailed run table: | Phase | Viewport | Cadence | Mismatch | Highlight lag | Content lag | Incomplete | Unpainted | p95 step | p95 frame | Chromium CPU | Server RSS | | --- | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | | Cold | 1440×900 | 15 ms | 0 | 0 | 0 | 21 | 218 | 56.1 ms | 91.4 ms | 56.1% | 153332 KiB | | Warm | 1440×900 | 15 ms | 0 | 0 | 0 | 0 | 217 | 88.1 ms | 65.2 ms | 36.3% | 217524 KiB | At this cadence, keys can arrive faster than display frames; the frame assertion stayed aligned to the latest intent on every warm frame. The cold first traversal does not meet the no-incomplete-frame target. Shared gate output and the Notes suite exception are recorded in the final #642 comment. The production captures for Settings are attached there at phone, tablet and desktop widths in light and dark.
Author
Owner

Photos post-fix cross-browser profile

Local production build, real 520-photo fixture, 33 ms cadence, 1,040 key events per cold/warm pass (520 forward and 520 backward). Screenshots were captured from the same build and fixture at 390, 820 and 1440 px in Light and Dark. The full JSON is in artifacts/blaze/photos-final.json in the worktree.

Browser / viewport Warm incomplete frames Warm steps never painted complete Warm p95 paint Warm max key handler Warm longest Long Task
Chromium 1440×900 321 809 90,848.7 ms 12.2 ms 5,939 ms
Chromium 390×844 316 743 62,085 ms 3.1 ms 2,204 ms
WebKit 1440×900 4 704 46,967 ms 1 ms not exposed by WebKit Long Tasks API
WebKit 390×844 0 562 49,128 ms 1 ms not exposed by WebKit Long Tasks API

Selected and content IDs matched in all warm frames. The failure is image completeness/paint latency, not selection identity. WebKit's zero incomplete frames at phone width does not mean every step completed: 562 of 1,040 steps had no sampled complete frame for their target. Chromium also exceeded the 50 ms long-task goal by a wide margin.

The existing source-mapped Chromium trace after the collection cache showed downloadUrl() at 18.6 ms, formatLongDate() at 14.2 ms, and timeOptions() at 16.3 ms total, down from 5.36 s, 7.74 s, and 5.25 s. It also showed 360 file-download requests, repeated per-photo tag lookups, 14.07 s of renderer-thread image decoding, and 4.13 s raster work. I am investigating the repeated selection side work and bounded image-cache churn before changing this path.

## Photos post-fix cross-browser profile Local production build, real 520-photo fixture, 33 ms cadence, 1,040 key events per cold/warm pass (520 forward and 520 backward). Screenshots were captured from the same build and fixture at 390, 820 and 1440 px in Light and Dark. The full JSON is in `artifacts/blaze/photos-final.json` in the worktree. | Browser / viewport | Warm incomplete frames | Warm steps never painted complete | Warm p95 paint | Warm max key handler | Warm longest Long Task | |---|---:|---:|---:|---:|---:| | Chromium 1440×900 | 321 | 809 | 90,848.7 ms | 12.2 ms | 5,939 ms | | Chromium 390×844 | 316 | 743 | 62,085 ms | 3.1 ms | 2,204 ms | | WebKit 1440×900 | 4 | 704 | 46,967 ms | 1 ms | not exposed by WebKit Long Tasks API | | WebKit 390×844 | 0 | 562 | 49,128 ms | 1 ms | not exposed by WebKit Long Tasks API | Selected and content IDs matched in all warm frames. The failure is image completeness/paint latency, not selection identity. WebKit's zero incomplete frames at phone width does not mean every step completed: 562 of 1,040 steps had no sampled complete frame for their target. Chromium also exceeded the 50 ms long-task goal by a wide margin. The existing source-mapped Chromium trace after the collection cache showed `downloadUrl()` at 18.6 ms, `formatLongDate()` at 14.2 ms, and `timeOptions()` at 16.3 ms total, down from 5.36 s, 7.74 s, and 5.25 s. It also showed 360 file-download requests, repeated per-photo tag lookups, 14.07 s of renderer-thread image decoding, and 4.13 s raster work. I am investigating the repeated selection side work and bounded image-cache churn before changing this path.
Author
Owner

Photos matrix after Info-request and frame-coalescing changes

Local production build; 520 real photos; Chromium and WebKit; 1440×900 and 390×844; 33 ms and 15 ms; 1,040 key events per cold/warm pass. Selected/content IDs matched in every warm row. The test still misses the selected-image completeness target. Full data: artifacts/blaze/photos-final-v2.json; screenshots: artifacts/blaze/screenshots/photos/.

Browser / viewport / cadence Warm incomplete frames Warm steps never painted complete Warm p95 paint Warm max key handler Warm longest task
Chromium 1440 / 33 ms 234 994 49,619.1 ms 3.5 ms 1,947 ms
Chromium 1440 / 15 ms 232 1,034 4,197.1 ms 2.4 ms 195 ms
Chromium 390 / 33 ms 371 1,001 50,133.4 ms 7.7 ms 1,177 ms
Chromium 390 / 15 ms 327 1,029 2,953.8 ms 2.8 ms 467 ms
WebKit 1440 / 33 ms 9 896 48,110 ms 1 ms not exposed
WebKit 1440 / 15 ms 7 933 29,699 ms 2 ms not exposed
WebKit 390 / 33 ms 0 702 38,261 ms 3 ms not exposed
WebKit 390 / 15 ms 5 839 21,117 ms 1 ms not exposed

The frame-coalescing test confirms only the latest queued target is selected once per frame. The current step counter still asks whether each individual key event, including one superseded before the next frame, received a complete paint. I will update the sampler to report those superseded inputs separately and limit paint latency to the same or next frame. The measured warm incomplete frames and long tasks remain real failures; I am continuing the requested Money, Files and Tabs runs before changing the Photos image prefetch window.

## Photos matrix after Info-request and frame-coalescing changes Local production build; 520 real photos; Chromium and WebKit; 1440×900 and 390×844; 33 ms and 15 ms; 1,040 key events per cold/warm pass. Selected/content IDs matched in every warm row. The test still misses the selected-image completeness target. Full data: `artifacts/blaze/photos-final-v2.json`; screenshots: `artifacts/blaze/screenshots/photos/`. | Browser / viewport / cadence | Warm incomplete frames | Warm steps never painted complete | Warm p95 paint | Warm max key handler | Warm longest task | |---|---:|---:|---:|---:|---:| | Chromium 1440 / 33 ms | 234 | 994 | 49,619.1 ms | 3.5 ms | 1,947 ms | | Chromium 1440 / 15 ms | 232 | 1,034 | 4,197.1 ms | 2.4 ms | 195 ms | | Chromium 390 / 33 ms | 371 | 1,001 | 50,133.4 ms | 7.7 ms | 1,177 ms | | Chromium 390 / 15 ms | 327 | 1,029 | 2,953.8 ms | 2.8 ms | 467 ms | | WebKit 1440 / 33 ms | 9 | 896 | 48,110 ms | 1 ms | not exposed | | WebKit 1440 / 15 ms | 7 | 933 | 29,699 ms | 2 ms | not exposed | | WebKit 390 / 33 ms | 0 | 702 | 38,261 ms | 3 ms | not exposed | | WebKit 390 / 15 ms | 5 | 839 | 21,117 ms | 1 ms | not exposed | The frame-coalescing test confirms only the latest queued target is selected once per frame. The current step counter still asks whether each individual key event, including one superseded before the next frame, received a complete paint. I will update the sampler to report those superseded inputs separately and limit paint latency to the same or next frame. The measured warm incomplete frames and long tasks remain real failures; I am continuing the requested Money, Files and Tabs runs before changing the Photos image prefetch window.
Author
Owner

The previous matrix identified a measurement defect before interpreting its results: Money compared the displayed register with the stale aria-current route marker instead of the focused account, and the runner counted every intermediate key event as a required paint even when a surface coalesces events within one animation frame. The sampler now records the focused Money account, marks earlier same-frame destinations as superseded, and requires the final destination to become complete within the next sampled frame. bunx vitest run src/lib/navigation/blazeMetrics.test.ts passes (3 tests). I am rerunning the affected profiles with this corrected sampler; the earlier step-level paint counts are superseded.

The previous matrix identified a measurement defect before interpreting its results: Money compared the displayed register with the stale `aria-current` route marker instead of the focused account, and the runner counted every intermediate key event as a required paint even when a surface coalesces events within one animation frame. The sampler now records the focused Money account, marks earlier same-frame destinations as superseded, and requires the final destination to become complete within the next sampled frame. `bunx vitest run src/lib/navigation/blazeMetrics.test.ts` passes (3 tests). I am rerunning the affected profiles with this corrected sampler; the earlier step-level paint counts are superseded.
Author
Owner

The corrected full Photos matrix reproduces a warm-path failure. Warm Chromium results (1440/390, 33/15 ms) report 202–370 incomplete frames, 239–627 latest destinations missing the 33 ms paint deadline, and 76–1,132 ms longest tasks. Warm WebKit reports 1–22 incomplete frames and 196–445 late destinations; WebKit does not expose Long Tasks API values in this runner. Selection/content mismatch frames are 0. The source-mapped warm trace from the attribution pass identifies image decode/raster as the dominant blocking work (Decode Image max 133.5 ms, Raster max 516.7 ms); I am fixing the hidden image work/prefetch path before another matrix.

The corrected full Photos matrix reproduces a warm-path failure. Warm Chromium results (1440/390, 33/15 ms) report 202–370 incomplete frames, 239–627 latest destinations missing the 33 ms paint deadline, and 76–1,132 ms longest tasks. Warm WebKit reports 1–22 incomplete frames and 196–445 late destinations; WebKit does not expose Long Tasks API values in this runner. Selection/content mismatch frames are 0. The source-mapped warm trace from the attribution pass identifies image decode/raster as the dominant blocking work (Decode Image max 133.5 ms, Raster max 516.7 ms); I am fixing the hidden image work/prefetch path before another matrix.
Author
Owner

The current production build now records 0 incomplete warm frames in all eight Photos browser/viewport/cadence combinations. Warm Chromium key-handler maxima are 0.1–6.8 ms and the runner recorded no warm Long Tasks; warm WebKit Long Tasks are unavailable in its engine. The next-target paint deadline still reports 168–743 late targets, and local headless frame-sample p95 is 98–240 ms while host load averaged 12.8–14.6. This run meets the explicit warm incomplete-frame goal, but the late-target and frame-cadence results need attribution before I report the Photos path as complete.

The current production build now records 0 incomplete warm frames in all eight Photos browser/viewport/cadence combinations. Warm Chromium key-handler maxima are 0.1–6.8 ms and the runner recorded no warm Long Tasks; warm WebKit Long Tasks are unavailable in its engine. The next-target paint deadline still reports 168–743 late targets, and local headless frame-sample p95 is 98–240 ms while host load averaged 12.8–14.6. This run meets the explicit warm incomplete-frame goal, but the late-target and frame-cadence results need attribution before I report the Photos path as complete.
Author
Owner

Fresh warm Photos trace: Chromium 1440 px, 520 items, 33 ms cadence. The profile shows little named app JavaScript per step. Largest named self-time entries were History.replaceState 390.4 ms, V8 garbage collector 151.9 ms, querySelector 79.1 ms, Svelte render.js 33.7 ms, Svelte batch.js 26.5 ms, and Svelte DOM event dispatch 22.1 ms. The app bundle frames without matching source maps were each below 22 ms; the benchmark's own number() sampler was 13.9 ms total.

Timeline totals (max event in parentheses): ImageDecodeTask 5.1 ms (0.13 ms); UpdateLayoutTree 266.4 ms (9.36 ms); Layout 307.9 ms (5.41 ms); Paint 411.6 ms (9.62 ms); RasterTask 1,003.2 ms (6.16 ms). Sampled GC self time was 151.9 ms. The longest traced RunTask was 241 ms on VizCompositorThread, not the page main thread; its thread CPU time was 95 ms. The untraced warm matrix recorded no Long Tasks entries in Chromium and 0 incomplete frames, but p95 requestAnimationFrame sample intervals were 98–240 ms and 168–743 latest targets missed the 33 ms deadline. I am recording this frame-pacing gap as a local browser/compositor limitation because the profile points away from synchronous decode, app JavaScript, style, layout, paint, or raster as the cause.

Fresh warm Photos trace: Chromium 1440 px, 520 items, 33 ms cadence. The profile shows little named app JavaScript per step. Largest named self-time entries were History.replaceState 390.4 ms, V8 garbage collector 151.9 ms, querySelector 79.1 ms, Svelte render.js 33.7 ms, Svelte batch.js 26.5 ms, and Svelte DOM event dispatch 22.1 ms. The app bundle frames without matching source maps were each below 22 ms; the benchmark's own `number()` sampler was 13.9 ms total. Timeline totals (max event in parentheses): ImageDecodeTask 5.1 ms (0.13 ms); UpdateLayoutTree 266.4 ms (9.36 ms); Layout 307.9 ms (5.41 ms); Paint 411.6 ms (9.62 ms); RasterTask 1,003.2 ms (6.16 ms). Sampled GC self time was 151.9 ms. The longest traced RunTask was 241 ms on VizCompositorThread, not the page main thread; its thread CPU time was 95 ms. The untraced warm matrix recorded no Long Tasks entries in Chromium and 0 incomplete frames, but p95 requestAnimationFrame sample intervals were 98–240 ms and 168–743 latest targets missed the 33 ms deadline. I am recording this frame-pacing gap as a local browser/compositor limitation because the profile points away from synchronous decode, app JavaScript, style, layout, paint, or raster as the cause.
Author
Owner

The corrected Money matrix seeded all 220 accounts and completed its Chromium cases, but a later WebKit case timed out in the runner's 120-second focus/selection/register settle wait. The run failed before writing the summary JSON. I am adding a final-state snapshot of the focused account, route-selected account and displayed register so I can isolate the exact browser/viewport case before interpreting the timing data.

The corrected Money matrix seeded all 220 accounts and completed its Chromium cases, but a later WebKit case timed out in the runner's 120-second focus/selection/register settle wait. The run failed before writing the summary JSON. I am adding a final-state snapshot of the focused account, route-selected account and displayed register so I can isolate the exact browser/viewport case before interpreting the timing data.
Author
Owner

Money held-key follow-up: a targeted WebKit 390×844, 33 ms run completed the 220-account seed and route settle, then recorded 718 warm mismatch frames and 719 incomplete frames across 400 arrow events. The first sampled frame had focus on blaze-account-001 while both aria-current and the register still showed blaze-account-000; warm key-handler work stayed at 0–1 ms, and WebKit did not expose Long Tasks. This points to route/register paint lag rather than a slow key handler. The current code rebuilds account-link records from the route on each account change, so I am removing that 220-row dependency and keeping current-page state to two attribute updates. Result: artifacts/blaze/money-webkit-phone-diag.json.

Money held-key follow-up: a targeted WebKit 390×844, 33 ms run completed the 220-account seed and route settle, then recorded 718 warm mismatch frames and 719 incomplete frames across 400 arrow events. The first sampled frame had focus on `blaze-account-001` while both `aria-current` and the register still showed `blaze-account-000`; warm key-handler work stayed at 0–1 ms, and WebKit did not expose Long Tasks. This points to route/register paint lag rather than a slow key handler. The current code rebuilds account-link records from the route on each account change, so I am removing that 220-row dependency and keeping current-page state to two attribute updates. Result: `artifacts/blaze/money-webkit-phone-diag.json`.
Author
Owner

Money serialization prevented the previous 120 s settle timeout, but did not meet the paint target. WebKit 390×844 at 33 ms finished with a complete final route; the warm pass still had 424 incomplete frames and 417 focus/content mismatches across 400 ArrowDown/ArrowUp events, with p95 frame interval 116 ms and max key-handler work 1 ms. The trace samples show focus one account ahead of the register during most of the pass. I am moving the existing three-register prefetch window to the latest focused target so its data can start loading before route navigation. Result: artifacts/blaze/money-serialized-webkit-phone.json; host load average was 11.78 / 15.39 / 18.67.

Money serialization prevented the previous 120 s settle timeout, but did not meet the paint target. WebKit 390×844 at 33 ms finished with a complete final route; the warm pass still had 424 incomplete frames and 417 focus/content mismatches across 400 ArrowDown/ArrowUp events, with p95 frame interval 116 ms and max key-handler work 1 ms. The trace samples show focus one account ahead of the register during most of the pass. I am moving the existing three-register prefetch window to the latest focused target so its data can start loading before route navigation. Result: `artifacts/blaze/money-serialized-webkit-phone.json`; host load average was 11.78 / 15.39 / 18.67.
Author
Owner

Focused-prefetch Money Chromium 1440×900, 33 ms (local; load average 27.50 / 25.02 / 23.63): warm recorded 303 incomplete frames, 302 focus/content mismatches, 301 unpainted destinations, p95 frame interval 133.4 ms, max key-handler work 6.3 ms, and longest Long Task 101 ms. Cold was 267 incomplete / 266 mismatches / 82 ms longest task. The prefetch window improved WebKit phone warm incomplete frames from 424 to 308, but neither profile meets the zero-incomplete or 50 ms task target. Result: artifacts/blaze/money-focused-prefetch-chromium-desktop.json. I am capturing a post-change warm trace to attribute the 101 ms task before selecting the next fix.

Focused-prefetch Money Chromium 1440×900, 33 ms (local; load average 27.50 / 25.02 / 23.63): warm recorded 303 incomplete frames, 302 focus/content mismatches, 301 unpainted destinations, p95 frame interval 133.4 ms, max key-handler work 6.3 ms, and longest Long Task 101 ms. Cold was 267 incomplete / 266 mismatches / 82 ms longest task. The prefetch window improved WebKit phone warm incomplete frames from 424 to 308, but neither profile meets the zero-incomplete or 50 ms task target. Result: `artifacts/blaze/money-focused-prefetch-chromium-desktop.json`. I am capturing a post-change warm trace to attribute the 101 ms task before selecting the next fix.
Author
Owner

Post-change Money profiling update for #641 (local host load was high: 22.8–27.5).

Warm 200-step runs after latest-route coalescing and focused three-register prefetch:

Browser / viewport Incomplete frames Mismatches p95 frame Longest task Max key handler
Chromium 1440×900 303 302 133.4 ms 101 ms 6.3 ms
WebKit 390×844 308 305 160 ms Long Tasks API unavailable 1 ms

The traced Chromium 1440×900 pass reported 362 incomplete frames (trace overhead), 56 ms longest task. Timeline totals: main-thread RunTask 11.16 s; FunctionCall 3.32 s (max 23.8 ms); UpdateLayoutTree 309.7 ms (max 8 ms); Layout 747.6 ms (max 11.3 ms); Paint 4.38 s (max 21.9 ms); RasterTask 320 ms (max 6.2 ms). Two inspected long tasks were paint/layerization-heavy; one included a keydown dispatch of 15.8 ms and a minified app call of 11.6 ms. CPU sampling showed focus self-time 593 ms, fetch 293 ms, querySelector 197 ms, GC 31.6 ms; app frames were minified in the captured trace. The result still misses the zero-incomplete and 50 ms targets. Results: artifacts/blaze/money-focused-prefetch-chromium-desktop.json, artifacts/blaze/money-focused-prefetch-webkit-phone.json, and artifacts/blaze/traces-money-post-prefetch/.

Post-change Money profiling update for #641 (local host load was high: 22.8–27.5). Warm 200-step runs after latest-route coalescing and focused three-register prefetch: | Browser / viewport | Incomplete frames | Mismatches | p95 frame | Longest task | Max key handler | | --- | ---: | ---: | ---: | ---: | ---: | | Chromium 1440×900 | 303 | 302 | 133.4 ms | 101 ms | 6.3 ms | | WebKit 390×844 | 308 | 305 | 160 ms | Long Tasks API unavailable | 1 ms | The traced Chromium 1440×900 pass reported 362 incomplete frames (trace overhead), 56 ms longest task. Timeline totals: main-thread RunTask 11.16 s; FunctionCall 3.32 s (max 23.8 ms); UpdateLayoutTree 309.7 ms (max 8 ms); Layout 747.6 ms (max 11.3 ms); Paint 4.38 s (max 21.9 ms); RasterTask 320 ms (max 6.2 ms). Two inspected long tasks were paint/layerization-heavy; one included a keydown dispatch of 15.8 ms and a minified app call of 11.6 ms. CPU sampling showed `focus` self-time 593 ms, `fetch` 293 ms, `querySelector` 197 ms, GC 31.6 ms; app frames were minified in the captured trace. The result still misses the zero-incomplete and 50 ms targets. Results: `artifacts/blaze/money-focused-prefetch-chromium-desktop.json`, `artifacts/blaze/money-focused-prefetch-webkit-phone.json`, and `artifacts/blaze/traces-money-post-prefetch/`.
Author
Owner

Full post-change Money matrix (local host, 33 ms cadence, 200 key steps per direction) finished on both browsers and viewports. All eight cold/warm passes failed the complete-paint assertion. Warm results:

Browser / viewport Incomplete frames Mismatch frames Never-painted destinations Longest task Max key handler
Chromium 1440×900 258 258 256 78 ms 0.2 ms
Chromium 390×844 571 568 399 71 ms 0.1 ms
WebKit 1440×900 126 126 126 Long Tasks API unavailable 1 ms
WebKit 390×844 340 331 331 Long Tasks API unavailable 1 ms

Frame samples show the focused account updates at keydown, but the register content remains on the prior account until the async transaction data arrives. The handler stays under 1 ms in every warm case; the remaining gap is register paint readiness. Full JSON: artifacts/blaze/money-full-final.json. Host load was shared and high; I am treating these as local diagnostic results, not a clean-host performance claim.

Full post-change Money matrix (local host, 33 ms cadence, 200 key steps per direction) finished on both browsers and viewports. All eight cold/warm passes failed the complete-paint assertion. Warm results: | Browser / viewport | Incomplete frames | Mismatch frames | Never-painted destinations | Longest task | Max key handler | | --- | ---: | ---: | ---: | ---: | ---: | | Chromium 1440×900 | 258 | 258 | 256 | 78 ms | 0.2 ms | | Chromium 390×844 | 571 | 568 | 399 | 71 ms | 0.1 ms | | WebKit 1440×900 | 126 | 126 | 126 | Long Tasks API unavailable | 1 ms | | WebKit 390×844 | 340 | 331 | 331 | Long Tasks API unavailable | 1 ms | Frame samples show the focused account updates at keydown, but the register content remains on the prior account until the async transaction data arrives. The handler stays under 1 ms in every warm case; the remaining gap is register paint readiness. Full JSON: `artifacts/blaze/money-full-final.json`. Host load was shared and high; I am treating these as local diagnostic results, not a clean-host performance claim.
Author
Owner

Finished the current #641 pass.

Branch: job/blaze-surfaces
Head: f94fc31442e85cbfd9d605ba2f4c45ab501ba08f (includes the required merge of origin/dev).

Built: coalesced held-key selection work to the latest target per frame; kept Files selection indexes and Money sidebar row state stable; bounded Quick Look image, PDF, and text preview work; deferred full image decode off the navigation path; bounded Money register prefetch and stale-read cancellation; and extended Blaze metrics for superseded destinations, Money timeout state, and trace summaries.

Files: apps/web/src/lib/files/FilesBrowser.svelte, apps/web/src/lib/files/selection.ts, apps/web/src/lib/photos/PhotoViewer.svelte, apps/web/src/lib/money/{api.ts,store.svelte.ts}, apps/web/src/lib/components/{SidebarLinks.svelte,money/MoneySidebar.svelte}, apps/web/src/routes/money/[budget]/accounts/[[account]]/+page.svelte, packages/ui/src/{imageCache.ts,navigation/latestFrame.ts}, packages/ui/src/components/viewer/{QuickLook.svelte,ImageView.svelte,PdfView.svelte,TextView.svelte,pdfPreviewCache.ts,textPreviewCache.ts}, packages/ui/src/components/TabBar.svelte, bench/blaze.{mjs,md}, and related unit tests and harness updates.

Performance: the A/B/A/B origin/dev comparison and the required pre-change Photos and Money trace attribution were posted earlier in this issue. The latest Photos profile had zero warm incomplete frames in all 8 browser/viewport/cadence cases; Chromium reported no warm Long Task entries and WebKit does not expose that API. The final Money matrix did not meet the target. Warm results were Chromium 1440×900: 258 incomplete frames, 256 never-painted destinations, 78 ms longest task; Chromium 390×844: 571, 399, 71 ms; WebKit 1440×900: 126, 126, Long Tasks API unavailable; WebKit 390×844: 340, 331, Long Tasks API unavailable. Warm key handlers were at or below 1 ms. Frames show focused account identity updates immediately, while register content trails until its async data arrives. Raw result: artifacts/blaze/money-full-final.json. The full Files and Tabs matrices were not rerun after the merge; their older local runs remain in artifacts/blaze/ and show failures.

Production screenshots from the merged build cover all 4 surfaces × 3 widths × 2 themes under artifacts/blaze/screenshots/{files,money,photos,tabs}/. They remain in the worktree and were not committed. fj v0.6.0 exposes no issue-asset upload command, so the screenshots are not attached to this comment.

Gates:

bun run build: passed; Vite emitted existing MODULE_LEVEL_DIRECTIVE warnings.

svelte-check found 404 errors and 0 warnings in 3 files
error: script "check" exited with code 1

 Test Files  162 passed (162)
      Tests  1081 passed (1081)
   Start at  11:31:09
   Duration  67.82s (transform 47%, environment 20%, import 17%, tests 12%, setup 4%)

Environment  |component| jsdom was created 52 times · 72.68s total, 28% of tracked time
             create it once per worker with pool: "vmThreads" (keeps per-file isolation) or isolate: false (shares it across files)
             learn more: https://vitest.dev/guide/improving-performance#test-environments

cargo clean: Removed 7255 files, 4.7GiB total

The check diagnostics are in apps/web/e2e/csp.mjs, apps/web/e2e/harness.mjs, and bench/blaze.mjs; the same typing debt was present before the final merge. bun run test passed.

Known gaps: Money still exceeds both the incomplete-frame and 50 ms warm-task targets in Chromium; WebKit Long Task measurements are unavailable. Files and Tabs need a current post-merge matrix. Screenshot attachments remain outstanding.

Decisions not specified by DESIGN: the Money directional prefetch window remains three registers. An earlier wider window increased server work without improving complete paints, but the final local matrix still misses the paint target. The local captures also ran under shared-host load, so clean-host performance remains unverified.

Finished the current #641 pass. **Branch:** `job/blaze-surfaces` **Head:** `f94fc31442e85cbfd9d605ba2f4c45ab501ba08f` (includes the required merge of `origin/dev`). **Built:** coalesced held-key selection work to the latest target per frame; kept Files selection indexes and Money sidebar row state stable; bounded Quick Look image, PDF, and text preview work; deferred full image decode off the navigation path; bounded Money register prefetch and stale-read cancellation; and extended Blaze metrics for superseded destinations, Money timeout state, and trace summaries. **Files:** `apps/web/src/lib/files/FilesBrowser.svelte`, `apps/web/src/lib/files/selection.ts`, `apps/web/src/lib/photos/PhotoViewer.svelte`, `apps/web/src/lib/money/{api.ts,store.svelte.ts}`, `apps/web/src/lib/components/{SidebarLinks.svelte,money/MoneySidebar.svelte}`, `apps/web/src/routes/money/[budget]/accounts/[[account]]/+page.svelte`, `packages/ui/src/{imageCache.ts,navigation/latestFrame.ts}`, `packages/ui/src/components/viewer/{QuickLook.svelte,ImageView.svelte,PdfView.svelte,TextView.svelte,pdfPreviewCache.ts,textPreviewCache.ts}`, `packages/ui/src/components/TabBar.svelte`, `bench/blaze.{mjs,md}`, and related unit tests and harness updates. **Performance:** the A/B/A/B `origin/dev` comparison and the required pre-change Photos and Money trace attribution were posted earlier in this issue. The latest Photos profile had zero warm incomplete frames in all 8 browser/viewport/cadence cases; Chromium reported no warm Long Task entries and WebKit does not expose that API. The final Money matrix did not meet the target. Warm results were Chromium 1440×900: 258 incomplete frames, 256 never-painted destinations, 78 ms longest task; Chromium 390×844: 571, 399, 71 ms; WebKit 1440×900: 126, 126, Long Tasks API unavailable; WebKit 390×844: 340, 331, Long Tasks API unavailable. Warm key handlers were at or below 1 ms. Frames show focused account identity updates immediately, while register content trails until its async data arrives. Raw result: `artifacts/blaze/money-full-final.json`. The full Files and Tabs matrices were not rerun after the merge; their older local runs remain in `artifacts/blaze/` and show failures. Production screenshots from the merged build cover all 4 surfaces × 3 widths × 2 themes under `artifacts/blaze/screenshots/{files,money,photos,tabs}/`. They remain in the worktree and were not committed. `fj v0.6.0` exposes no issue-asset upload command, so the screenshots are not attached to this comment. **Gates:** ```text bun run build: passed; Vite emitted existing MODULE_LEVEL_DIRECTIVE warnings. svelte-check found 404 errors and 0 warnings in 3 files error: script "check" exited with code 1 Test Files 162 passed (162) Tests 1081 passed (1081) Start at 11:31:09 Duration 67.82s (transform 47%, environment 20%, import 17%, tests 12%, setup 4%) Environment |component| jsdom was created 52 times · 72.68s total, 28% of tracked time create it once per worker with pool: "vmThreads" (keeps per-file isolation) or isolate: false (shares it across files) learn more: https://vitest.dev/guide/improving-performance#test-environments cargo clean: Removed 7255 files, 4.7GiB total ``` The check diagnostics are in `apps/web/e2e/csp.mjs`, `apps/web/e2e/harness.mjs`, and `bench/blaze.mjs`; the same typing debt was present before the final merge. `bun run test` passed. **Known gaps:** Money still exceeds both the incomplete-frame and 50 ms warm-task targets in Chromium; WebKit Long Task measurements are unavailable. Files and Tabs need a current post-merge matrix. Screenshot attachments remain outstanding. **Decisions not specified by DESIGN:** the Money directional prefetch window remains three registers. An earlier wider window increased server work without improving complete paints, but the final local matrix still misses the paint target. The local captures also ran under shared-host load, so clean-host performance remains unverified.
Author
Owner

Results

Surface dev c4a61e8 feature build Result
Settings (job/blaze-settings, bf3ad5f) Not measured Not measured No samples before cutoff
Mail columns (job/maillayouts, f9f360e) Not measured Not measured No samples before cutoff
Mail split (job/maillayouts, f9f360e) Not measured Not measured No samples before cutoff
Mail morph (job/maillayouts, f9f360e) Not measured Not measured No samples before cutoff

No latency, CPU or RSS samples exist, and no regression is inferred. The full browser and cadence matrices with three interleaved A/B runs did not fit in the perf-VM window.

Commands

For Settings: BLAZE_BROWSERS=chromium,webkit BLAZE_WIDTHS=1440x900,390x844 BLAZE_CADENCES=33,15 BLAZE_STEPS=200 bun bench/blaze.mjs --surface settings --json <output>.

For Mail: the same environment and command with --surface mail-columns, --surface mail-split and --surface mail-morph, using bench/blaze.mjs from job/maillayouts. Each run would use the saved dev or feature production bundle, an optimized release server, /root/hdd-emu.sh run-limited, TMPDIR=/srv/hdd-emu/tmp and flock -w 14400 /root/perf.lock.

## Results | Surface | dev `c4a61e8` | feature build | Result | | --- | --- | --- | --- | | Settings (`job/blaze-settings`, `bf3ad5f`) | Not measured | Not measured | No samples before cutoff | | Mail columns (`job/maillayouts`, `f9f360e`) | Not measured | Not measured | No samples before cutoff | | Mail split (`job/maillayouts`, `f9f360e`) | Not measured | Not measured | No samples before cutoff | | Mail morph (`job/maillayouts`, `f9f360e`) | Not measured | Not measured | No samples before cutoff | No latency, CPU or RSS samples exist, and no regression is inferred. The full browser and cadence matrices with three interleaved A/B runs did not fit in the perf-VM window. ## Commands For Settings: `BLAZE_BROWSERS=chromium,webkit BLAZE_WIDTHS=1440x900,390x844 BLAZE_CADENCES=33,15 BLAZE_STEPS=200 bun bench/blaze.mjs --surface settings --json <output>`. For Mail: the same environment and command with `--surface mail-columns`, `--surface mail-split` and `--surface mail-morph`, using `bench/blaze.mjs` from `job/maillayouts`. Each run would use the saved dev or feature production bundle, an optimized release server, `/root/hdd-emu.sh run-limited`, `TMPDIR=/srv/hdd-emu/tmp` and `flock -w 14400 /root/perf.lock`.
Author
Owner

Starting round 3 on job/blaze-surfaces. Branch HEAD is f94fc31442e85cbfd9d605ba2f4c45ab501ba08f; origin/dev base is c4a61e8cf090170f35b1bed3350d9de20c83ecd5.

I will compare the branch check output with origin/dev, run the Files and Tabs production blaze matrix on the locked perf VM, capture the requested macOS-rendered UI evidence, and record Money read-path findings on #687. No push, deploy, or merge.

Starting round 3 on `job/blaze-surfaces`. Branch HEAD is `f94fc31442e85cbfd9d605ba2f4c45ab501ba08f`; `origin/dev` base is `c4a61e8cf090170f35b1bed3350d9de20c83ecd5`. I will compare the branch check output with `origin/dev`, run the Files and Tabs production blaze matrix on the locked perf VM, capture the requested macOS-rendered UI evidence, and record Money read-path findings on #687. No push, deploy, or merge.
Author
Owner

Check-scope finding: the branch’s svelte-check reaches e2e/csp.mjs, e2e/harness.mjs and bench/blaze.mjs through apps/web/src/lib/navigation/blazeMetrics.test.ts importing bench/blaze.mjs. That import graph adds browser-only JavaScript to the Svelte check.

I ran bun run check at origin/dev (c4a61e8cf090170f35b1bed3350d9de20c83ecd5) in a detached worktree with the same installed workspace dependencies. Its output ends svelte-check found 4 errors and 0 warnings in 1 file; all four are nav.mode6–nav.mode9 in src/lib/shortcuts/registry.ts. It reports no diagnostics in the three browser/benchmark files. The current branch reports svelte-check found 404 errors and 0 warnings in 3 files. I will keep the existing metric assertions and move the pure helpers into typed app source, so this test no longer widens the check graph.

Check-scope finding: the branch’s `svelte-check` reaches `e2e/csp.mjs`, `e2e/harness.mjs` and `bench/blaze.mjs` through `apps/web/src/lib/navigation/blazeMetrics.test.ts` importing `bench/blaze.mjs`. That import graph adds browser-only JavaScript to the Svelte check. I ran `bun run check` at `origin/dev` (`c4a61e8cf090170f35b1bed3350d9de20c83ecd5`) in a detached worktree with the same installed workspace dependencies. Its output ends `svelte-check found 4 errors and 0 warnings in 1 file`; all four are `nav.mode6`–`nav.mode9` in `src/lib/shortcuts/registry.ts`. It reports no diagnostics in the three browser/benchmark files. The current branch reports `svelte-check found 404 errors and 0 warnings in 3 files`. I will keep the existing metric assertions and move the pure helpers into typed app source, so this test no longer widens the check graph.
Author
Owner

Round 3 web test finding: bun run test completed in 800.33 s with 13 failed files, 22 failed tests and 1,059 passed tests (1,081 total). All listed failures were test timeouts (mostly the 5 s default; one 15 s test), and Vitest also reported a worker termination in src/lib/components/analytics/BklitTooltipMaterial.test.ts. Transform time was 2,380.20 s tracked across workers (59%); the build host load average was 168.47 at completion. The focused changed-metrics test passes: 1 file, 3 tests. No existing expectations were changed. This run does not establish whether the unrelated component timeouts reproduce on a quieter host.

Round 3 web test finding: `bun run test` completed in 800.33 s with 13 failed files, 22 failed tests and 1,059 passed tests (1,081 total). All listed failures were test timeouts (mostly the 5 s default; one 15 s test), and Vitest also reported a worker termination in `src/lib/components/analytics/BklitTooltipMaterial.test.ts`. Transform time was 2,380.20 s tracked across workers (59%); the build host load average was 168.47 at completion. The focused changed-metrics test passes: 1 file, 3 tests. No existing expectations were changed. This run does not establish whether the unrelated component timeouts reproduce on a quieter host.
Author
Owner

Perf finding (branch A1, before the origin/dev comparison): Files used its 1,200-item fixture; the comparison traversal was 200 steps across Chromium/WebKit, phone/desktop and 33/15 ms. The lock-held load average was 2.70 1.31 0.60. The runner wrote files-A-A1.json but failed its completeness gate in all 16 phase/cases. Selected warm results: Chromium desktop/33 ms had 80 incomplete frames, 137 latest targets without a complete paint and a 339 ms longest task; Chromium phone/33 ms had 156 incomplete, 326 unpainted and a 653 ms longest task; Chromium phone/15 ms had 159 incomplete and 240 unpainted; WebKit phone/33 ms had 33 incomplete and 160 unpainted (WebKit does not expose this runner's long-task value). ID mismatches were zero. The B1 run on origin/dev is in progress, so these results are not yet labelled regressions.

Perf finding (branch A1, before the `origin/dev` comparison): Files used its 1,200-item fixture; the comparison traversal was 200 steps across Chromium/WebKit, phone/desktop and 33/15 ms. The lock-held load average was `2.70 1.31 0.60`. The runner wrote `files-A-A1.json` but failed its completeness gate in all 16 phase/cases. Selected warm results: Chromium desktop/33 ms had 80 incomplete frames, 137 latest targets without a complete paint and a 339 ms longest task; Chromium phone/33 ms had 156 incomplete, 326 unpainted and a 653 ms longest task; Chromium phone/15 ms had 159 incomplete and 240 unpainted; WebKit phone/33 ms had 33 incomplete and 160 unpainted (WebKit does not expose this runner's long-task value). ID mismatches were zero. The B1 run on `origin/dev` is in progress, so these results are not yet labelled regressions.
Author
Owner

Files A/B table (round 3, run pair 1). Both production bundles used the same current runner, 1,200-item mixed fixture on the #549 HDD emulator, and the same full browser/viewport/cadence matrix; this pair used 200 steps per pass. A = branch edb8861, B = origin/dev c4a61e8. Both runs held /root/perf.lock; load inside A/B locks was 2.70 1.31 0.60 / 4.22 6.73 3.87.

Warm values are incomplete frames / latest destinations without complete paint / p95 paint ms / longest task ms:

Case A: branch B: dev
Chromium desktop 33 ms 80 / 137 / 26.3 / 339 38 / 131 / 25.4 / 470
Chromium desktop 15 ms 52 / 88 / 24.6 / 271 53 / 89 / 11.2 / 81
Chromium phone 33 ms 156 / 326 / 32.9 / 653 144 / 301 / 32.1 / 159
Chromium phone 15 ms 159 / 240 / 33.3 / 132 133 / 213 / 32.2 / 61
WebKit desktop 33 ms 16 / 98 / n/a / n/a 17 / 84 / n/a / n/a
WebKit desktop 15 ms 14 / 68 / n/a / n/a 23 / 60 / n/a / n/a
WebKit phone 33 ms 33 / 160 / 30 / n/a 34 / 156 / n/a / n/a
WebKit phone 15 ms 28 / 123 / n/a / n/a 35 / 111 / n/a / n/a

ID mismatches were zero in every pass. These data show branch regressions in Chromium phone (both cadences), Chromium desktop at 15 ms, and smaller WebKit unpainted counts in several cases; I am tracing the Files code before deciding the smallest fix. Both revisions miss the zero-incomplete target. Full JSON is retained on the worktree/perf VM; screenshots remain separate.

Files A/B table (round 3, run pair 1). Both production bundles used the same current runner, 1,200-item mixed fixture on the #549 HDD emulator, and the same full browser/viewport/cadence matrix; this pair used 200 steps per pass. A = branch `edb8861`, B = `origin/dev` `c4a61e8`. Both runs held `/root/perf.lock`; load inside A/B locks was `2.70 1.31 0.60` / `4.22 6.73 3.87`. Warm values are `incomplete frames / latest destinations without complete paint / p95 paint ms / longest task ms`: | Case | A: branch | B: dev | |---|---:|---:| | Chromium desktop 33 ms | 80 / 137 / 26.3 / 339 | 38 / 131 / 25.4 / 470 | | Chromium desktop 15 ms | 52 / 88 / 24.6 / 271 | 53 / 89 / 11.2 / 81 | | Chromium phone 33 ms | 156 / 326 / 32.9 / 653 | 144 / 301 / 32.1 / 159 | | Chromium phone 15 ms | 159 / 240 / 33.3 / 132 | 133 / 213 / 32.2 / 61 | | WebKit desktop 33 ms | 16 / 98 / n/a / n/a | 17 / 84 / n/a / n/a | | WebKit desktop 15 ms | 14 / 68 / n/a / n/a | 23 / 60 / n/a / n/a | | WebKit phone 33 ms | 33 / 160 / 30 / n/a | 34 / 156 / n/a / n/a | | WebKit phone 15 ms | 28 / 123 / n/a / n/a | 35 / 111 / n/a / n/a | ID mismatches were zero in every pass. These data show branch regressions in Chromium phone (both cadences), Chromium desktop at 15 ms, and smaller WebKit unpainted counts in several cases; I am tracing the Files code before deciding the smallest fix. Both revisions miss the zero-incomplete target. Full JSON is retained on the worktree/perf VM; screenshots remain separate.
Author
Owner

Round 3 report — head f4ebbac5c731b5ed6f15a934fcb433056c54c6ca (branch job/blaze-surfaces). No push or deploy.

Built:

  • Moved Blaze frame metrics into apps/web/src/lib/navigation/blazeMetrics.ts, so bun run check no longer imports browser-only e2e/harness.mjs/e2e/csp.mjs/bench/blaze.mjs through the test. The origin/dev check at c4a61e8 had 4 nav.mode6–nav.mode9 errors in packages/ui/src/components/TabBar.svelte; it did not type-check those three .mjs files. Branch check passed with 0 errors and 0 warnings before the final Files change.
  • Made Blaze screenshots emulate navigator.platform=MacIntel and userAgentData.platform=macOS; the runner asserts both values. Files and Tabs captures cover 390, 820 and 1440 px in Light and Dark. They are in the ignored worktree at artifacts/blaze/screenshots/{files,tabs}/; they are not attached to this comment.
  • Removed speculative 1024 px warming for unopened Files rows. Thumbnail warming remains in Files; Quick Look owns decoded preview warming.
  • Recorded Money request/projection evidence on #687.

Files perf evidence: A1/B1 each ran the full browser, width and cadence matrix with the 1,200-item fixture and 200 steps per pass. There was one full-matrix pair, not three repeated full matrices. A1/B1 warm values below are incomplete / unpainted / p95 paint ms / longest task ms; zero ID mismatches in all cases:

Case Branch A1 origin/dev B1
Chromium desktop 33 ms 80 / 137 / 26.3 / 339 38 / 131 / 25.4 / 470
Chromium desktop 15 ms 52 / 88 / 24.6 / 271 53 / 89 / 11.2 / 81
Chromium phone 33 ms 156 / 326 / 32.9 / 653 144 / 301 / 32.1 / 159
Chromium phone 15 ms 159 / 240 / 33.3 / 132 133 / 213 / 32.2 / 61
WebKit desktop 33 ms 16 / 98 / n/a / n/a 17 / 84 / n/a / n/a
WebKit desktop 15 ms 14 / 68 / n/a / n/a 23 / 60 / n/a / n/a
WebKit phone 33 ms 33 / 160 / 30 / n/a 34 / 156 / n/a / n/a
WebKit phone 15 ms 28 / 123 / n/a / n/a 35 / 111 / n/a / n/a

A2/B2 repeated Chromium phone/33 ms after the Files change. A2/B2 warm was 129 / 338 / 32.2 / 102 versus 119 / 301 / 31.2 / 120. The change lowered the longest task but did not close the completeness regression; unpainted targets were higher. No third interleaved pair was completed. Both versions still miss the zero-incomplete target.

Tabs: A1 ran the matrix using five common modes. B1 skipped all 16 cases because origin/dev has no visible mode tray (surface list is not present at this width). The benchmark docs define that older revision as not applicable, so this is not a valid A/B baseline. A1 warm Chromium had 68/135 incomplete/unpainted at desktop 33 ms, 39/79 at desktop 15 ms, 175/102 at phone 33 ms and 170/139 at phone 15 ms. No repeated Tabs rounds were completed.

Gates and evidence:

  • bun run check before the final Files reduction: svelte-check found 0 errors and 0 warnings. The post-merge final check was started but did not finish before the 2.5 h round limit.
  • bun run test: Test Files 13 failed | 149 passed (162); Tests 22 failed | 1059 passed (1081); duration 800.33s. Failures were timeouts across unrelated files. Vitest reported 2,380.20 s of tracked transform time and a terminated worker. Host load average at completion was 168.47. Focused changed-metrics test passed: 1 file, 3 tests.
  • bun run build passed after the Files change: ✓ built in 8m 33s; Wrote site to "build".
  • node --check bench/blaze.mjs passed before the screenshot run. The macOS screenshot assertions passed.
  • No Rust files changed, so Rust gates were not run.

Known gaps: fewer than three interleaved measurements; Files still misses the complete-paint target and the A2/B2 comparison is mixed; Tabs has no origin/dev baseline; screenshot PNGs remain local and are not attached; the post-merge check, cargo clean, and removal of web build output did not complete within this round. No existing test expectation was changed.

Decisions where DESIGN was silent: use a 200-step traversal for the repeated A/B workload while retaining the full 1,200-item Files fixture, to fit the round limit; treat Tabs on origin/dev as not applicable when its mode tray is absent, matching bench/blaze.md.

Round 3 report — head `f4ebbac5c731b5ed6f15a934fcb433056c54c6ca` (branch `job/blaze-surfaces`). No push or deploy. Built: - Moved Blaze frame metrics into `apps/web/src/lib/navigation/blazeMetrics.ts`, so `bun run check` no longer imports browser-only `e2e/harness.mjs`/`e2e/csp.mjs`/`bench/blaze.mjs` through the test. The `origin/dev` check at `c4a61e8` had 4 `nav.mode6`–`nav.mode9` errors in `packages/ui/src/components/TabBar.svelte`; it did not type-check those three `.mjs` files. Branch check passed with 0 errors and 0 warnings before the final Files change. - Made Blaze screenshots emulate `navigator.platform=MacIntel` and `userAgentData.platform=macOS`; the runner asserts both values. Files and Tabs captures cover 390, 820 and 1440 px in Light and Dark. They are in the ignored worktree at `artifacts/blaze/screenshots/{files,tabs}/`; they are not attached to this comment. - Removed speculative 1024 px warming for unopened Files rows. Thumbnail warming remains in Files; Quick Look owns decoded preview warming. - Recorded Money request/projection evidence on #687. Files perf evidence: A1/B1 each ran the full browser, width and cadence matrix with the 1,200-item fixture and 200 steps per pass. There was one full-matrix pair, not three repeated full matrices. A1/B1 warm values below are `incomplete / unpainted / p95 paint ms / longest task ms`; zero ID mismatches in all cases: | Case | Branch A1 | origin/dev B1 | |---|---:|---:| | Chromium desktop 33 ms | 80 / 137 / 26.3 / 339 | 38 / 131 / 25.4 / 470 | | Chromium desktop 15 ms | 52 / 88 / 24.6 / 271 | 53 / 89 / 11.2 / 81 | | Chromium phone 33 ms | 156 / 326 / 32.9 / 653 | 144 / 301 / 32.1 / 159 | | Chromium phone 15 ms | 159 / 240 / 33.3 / 132 | 133 / 213 / 32.2 / 61 | | WebKit desktop 33 ms | 16 / 98 / n/a / n/a | 17 / 84 / n/a / n/a | | WebKit desktop 15 ms | 14 / 68 / n/a / n/a | 23 / 60 / n/a / n/a | | WebKit phone 33 ms | 33 / 160 / 30 / n/a | 34 / 156 / n/a / n/a | | WebKit phone 15 ms | 28 / 123 / n/a / n/a | 35 / 111 / n/a / n/a | A2/B2 repeated Chromium phone/33 ms after the Files change. A2/B2 warm was `129 / 338 / 32.2 / 102` versus `119 / 301 / 31.2 / 120`. The change lowered the longest task but did not close the completeness regression; unpainted targets were higher. No third interleaved pair was completed. Both versions still miss the zero-incomplete target. Tabs: A1 ran the matrix using five common modes. B1 skipped all 16 cases because `origin/dev` has no visible mode tray (`surface list is not present at this width`). The benchmark docs define that older revision as not applicable, so this is not a valid A/B baseline. A1 warm Chromium had 68/135 incomplete/unpainted at desktop 33 ms, 39/79 at desktop 15 ms, 175/102 at phone 33 ms and 170/139 at phone 15 ms. No repeated Tabs rounds were completed. Gates and evidence: - `bun run check` before the final Files reduction: `svelte-check found 0 errors and 0 warnings`. The post-merge final check was started but did not finish before the 2.5 h round limit. - `bun run test`: `Test Files 13 failed | 149 passed (162)`; `Tests 22 failed | 1059 passed (1081)`; duration `800.33s`. Failures were timeouts across unrelated files. Vitest reported 2,380.20 s of tracked transform time and a terminated worker. Host load average at completion was 168.47. Focused changed-metrics test passed: 1 file, 3 tests. - `bun run build` passed after the Files change: `✓ built in 8m 33s`; `Wrote site to "build"`. - `node --check bench/blaze.mjs` passed before the screenshot run. The macOS screenshot assertions passed. - No Rust files changed, so Rust gates were not run. Known gaps: fewer than three interleaved measurements; Files still misses the complete-paint target and the A2/B2 comparison is mixed; Tabs has no `origin/dev` baseline; screenshot PNGs remain local and are not attached; the post-merge check, `cargo clean`, and removal of web build output did not complete within this round. No existing test expectation was changed. Decisions where DESIGN was silent: use a 200-step traversal for the repeated A/B workload while retaining the full 1,200-item Files fixture, to fit the round limit; treat Tabs on `origin/dev` as not applicable when its mode tray is absent, matching `bench/blaze.md`.
Author
Owner

Starting the requested perf VM rerun for #641. Branch: job/blaze-surfaces. Head/base checkout: f4ebbac5c. Earlier build-host timings are invalid and will not be used. Compare production origin/dev and branch bundles with one fixed harness/server, three alternating A/B pairs for Settings, Mail, Notes and Files, cold and warm, under /root/perf.lock with HDD emulation. No deployment or push.

Starting the requested perf VM rerun for #641. Branch: job/blaze-surfaces. Head/base checkout: f4ebbac5c. Earlier build-host timings are invalid and will not be used. Compare production origin/dev and branch bundles with one fixed harness/server, three alternating A/B pairs for Settings, Mail, Notes and Files, cold and warm, under /root/perf.lock with HDD emulation. No deployment or push.
Author
Owner

Fresh verification after merging origin/dev c4faf184d into this branch:

$ node scripts/check-user-storage.mjs && node scripts/check-type-tokens.mjs && node scripts/check-motion-tokens.mjs && svelte-kit sync && svelte-check --tsconfig ./tsconfig.json
User browser caches use userStorage; only documented device/public-link exceptions remain.
Text sizes and UI shape values use shared role tokens.
UI transitions and animation options use shared motion tokens or documented exceptions.
Loading svelte-check in workspace: /home/kayg/Developer/calternal-wt/blaze-surfaces/apps/web
Getting Svelte diagnostics...

svelte-check found 0 errors and 0 warnings

$ vitest run "--maxWorkers=2"

 RUN  v5.0.1 /home/kayg/Developer/calternal-wt/blaze-surfaces/apps/web

Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Could not parse CSS stylesheet
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method
Not implemented: Window's scrollTo() method

 Test Files  162 passed (162)
      Tests  1081 passed (1081)
   Start at  20:11:44
   Duration  287.20s (transform 29%, import 27%, environment 26%, tests 14%, setup 4%)

Environment  |component| jsdom was created 52 times · 92.90s total, 32% of tracked time
             create it once per worker with pool: 'vmThreads' (keeps per-file isolation) or isolate: false (shares it across files)
             learn more: https://vitest.dev/guide/improving-performance#test-environments


Coverage finding: bench/blaze.mjs on this branch defines Settings, Files, Photos, Money and Tabs. It has no Mail or Notes adapter. MailView.svelte has no held-arrow message selection handler; NotesExplorer.svelte ArrowDown/Up moves focus without opening a Note. The available Settings and Files A/B matrices are being run on root@10.69.69.63, using three alternating origin/dev (A) / branch (B) pairs with a fixed release server and sampler, under the perf lock and HDD scope. The earlier build-host timings remain invalid.

Fresh verification after merging origin/dev c4faf184d into this branch: ```text $ node scripts/check-user-storage.mjs && node scripts/check-type-tokens.mjs && node scripts/check-motion-tokens.mjs && svelte-kit sync && svelte-check --tsconfig ./tsconfig.json User browser caches use userStorage; only documented device/public-link exceptions remain. Text sizes and UI shape values use shared role tokens. UI transitions and animation options use shared motion tokens or documented exceptions. Loading svelte-check in workspace: /home/kayg/Developer/calternal-wt/blaze-surfaces/apps/web Getting Svelte diagnostics... svelte-check found 0 errors and 0 warnings $ vitest run "--maxWorkers=2" RUN v5.0.1 /home/kayg/Developer/calternal-wt/blaze-surfaces/apps/web Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Could not parse CSS stylesheet Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Not implemented: Window's scrollTo() method Test Files 162 passed (162) Tests 1081 passed (1081) Start at 20:11:44 Duration 287.20s (transform 29%, import 27%, environment 26%, tests 14%, setup 4%) Environment |component| jsdom was created 52 times · 92.90s total, 32% of tracked time create it once per worker with pool: 'vmThreads' (keeps per-file isolation) or isolate: false (shares it across files) learn more: https://vitest.dev/guide/improving-performance#test-environments ``` Coverage finding: bench/blaze.mjs on this branch defines Settings, Files, Photos, Money and Tabs. It has no Mail or Notes adapter. MailView.svelte has no held-arrow message selection handler; NotesExplorer.svelte ArrowDown/Up moves focus without opening a Note. The available Settings and Files A/B matrices are being run on root@10.69.69.63, using three alternating origin/dev (A) / branch (B) pairs with a fixed release server and sampler, under the perf lock and HDD scope. The earlier build-host timings remain invalid.
Author
Owner

Measurement validity finding: Settings A1 and B1 each sent 400 key events per phase. Every non-skipped phase sampled only one selected identity, Account. FloatingSidebar.svelte explicitly has no arrow-key handling. Thus these runs prove a missing interaction; their latency cannot be called Settings navigation performance. Mail has no held-arrow message selection and NotesExplorer arrows only move tree focus. Those surfaces cannot supply a valid held-arrow comparison from these two production revisions.

Files probe finding: the old sampler rejected an already decoded visible preview while a hidden full-image decoder existed. Its PDF check also accepted an allocated canvas before rendering (including an explicit false readiness flag). I replaced those checks with one shared visible-element probe: decoded visible image, actual text with no loading status, or non-white raster pixels in the fixed PDF fixture text region. Focused regression verification: 2 test files, 7 tests passed. Both Files revisions will use this same corrected sampler. Canvas readback is benchmark instrumentation only.

Measurement validity finding: Settings A1 and B1 each sent 400 key events per phase. Every non-skipped phase sampled only one selected identity, Account. FloatingSidebar.svelte explicitly has no arrow-key handling. Thus these runs prove a missing interaction; their latency cannot be called Settings navigation performance. Mail has no held-arrow message selection and NotesExplorer arrows only move tree focus. Those surfaces cannot supply a valid held-arrow comparison from these two production revisions. Files probe finding: the old sampler rejected an already decoded visible preview while a hidden full-image decoder existed. Its PDF check also accepted an allocated canvas before rendering (including an explicit false readiness flag). I replaced those checks with one shared visible-element probe: decoded visible image, actual text with no loading status, or non-white raster pixels in the fixed PDF fixture text region. Focused regression verification: 2 test files, 7 tests passed. Both Files revisions will use this same corrected sampler. Canvas readback is benchmark instrumentation only.
Author
Owner

Measurement validity finding: holdKey awaited each Input.dispatchKeyEvent/page.keyboard.down acknowledgement, then slept 33 or 15 ms. A main-thread stall therefore reduced the delivered input rate and weakened the Blaze workload. The pre-fix Files comparisons are invalid for the requested fixed cadence; no branch performance verdict is drawn from them.

I extracted the repeat dispatcher into the typed metrics module. A regression test with 300 ms browser acknowledgements failed against the old implementation: only the initial key was sent by 100 ms, where the independently specified deadlines require keys at 0, 33, 66 and 99 ms. The fixed dispatcher sends on absolute deadlines, observes all promises immediately, drains them before key-up, and stops on transport failure. The sampler uses native event timestamps for queue delay and measures handler work separately. Focused verification passed: 3 files, 9 tests. A single corrected round has started on the locked perf VM, with fresh result names and three alternating A/B pairs for the available Settings and Files profiles. No app cache changes have been made from the invalid timing comparison.

Measurement validity finding: holdKey awaited each Input.dispatchKeyEvent/page.keyboard.down acknowledgement, then slept 33 or 15 ms. A main-thread stall therefore reduced the delivered input rate and weakened the Blaze workload. The pre-fix Files comparisons are invalid for the requested fixed cadence; no branch performance verdict is drawn from them. I extracted the repeat dispatcher into the typed metrics module. A regression test with 300 ms browser acknowledgements failed against the old implementation: only the initial key was sent by 100 ms, where the independently specified deadlines require keys at 0, 33, 66 and 99 ms. The fixed dispatcher sends on absolute deadlines, observes all promises immediately, drains them before key-up, and stops on transport failure. The sampler uses native event timestamps for queue delay and measures handler work separately. Focused verification passed: 3 files, 9 tests. A single corrected round has started on the locked perf VM, with fresh result names and three alternating A/B pairs for the available Settings and Files profiles. No app cache changes have been made from the invalid timing comparison.
Author
Owner

Corrected perf-VM round continues at head 8075f518e (sampler cadb776e2). Settings A1/B1/A2/B2/A3/B3 completed: all 48 desktop cold/warm phases remain on Account after 400 keys. This does not measure navigation. Files A1/B1/A2 are complete; remaining B2/A3/B3 use the same production bundles and fixed-deadline sampler under /root/perf.lock.

A focused branch Chromium desktop/33 ms trace captured a 427.6 ms main-thread task around layer work and a 319.2 ms task containing a 287.7 ms compositor commit. It does not attribute the regression to PDF prefetch. Trace timings are excluded from the A/B table. An isolated production comparison bundle removes only commit 4e58f2779's Quick Look frame queue; a locked VM run will test whether that delay explains the lower complete-paint count. No application change has been made from this hypothesis.

Full web verification already passed after the measurement fixes: svelte-check found 0 errors and 0 warnings; Test Files 164 passed (164), Tests 1087 passed (1087). The final report will quote the actual gate summaries and include all completed pairs.

Corrected perf-VM round continues at head 8075f518e (sampler cadb776e2). Settings A1/B1/A2/B2/A3/B3 completed: all 48 desktop cold/warm phases remain on Account after 400 keys. This does not measure navigation. Files A1/B1/A2 are complete; remaining B2/A3/B3 use the same production bundles and fixed-deadline sampler under /root/perf.lock. A focused branch Chromium desktop/33 ms trace captured a 427.6 ms main-thread task around layer work and a 319.2 ms task containing a 287.7 ms compositor commit. It does not attribute the regression to PDF prefetch. Trace timings are excluded from the A/B table. An isolated production comparison bundle removes only commit 4e58f2779's Quick Look frame queue; a locked VM run will test whether that delay explains the lower complete-paint count. No application change has been made from this hypothesis. Full web verification already passed after the measurement fixes: svelte-check found 0 errors and 0 warnings; Test Files 164 passed (164), Tests 1087 passed (1087). The final report will quote the actual gate summaries and include all completed pairs.
Author
Owner

Head 63529a7ed fixes a second measurement bias. The Chromium trace has sampler callback 211 at 60491936136 microseconds, followed by Quick Look update callback 212 at 60491936785 in the same animation frame. The old sampler records the previous item before queued navigation updates, even though the update runs before paint. It adds a false frame of latency to the queued revision.

A focused regression test reproduced the old order: observed [1], expected [2]. The fixed sampler moves its pending frame callback behind input-handler frame work on both revisions. Focused verification: Test Files 3 passed (3), Tests 10 passed (10). bun run check: svelte-check found 0 errors and 0 warnings. The requested full web test suite is running.

All earlier paint counters, including the queue and image-prewarming attribution probes, are excluded from the verdict. No application revert is justified by them. The final production A1/B1/A2/B2/A3/B3 round is running under /root/perf.lock in /root/blaze-641-20261002/results-ordered, using sampler 63529a7ed and the unchanged A/B bundles and optimized release server. Source files are frozen for this round.

Head 63529a7ed fixes a second measurement bias. The Chromium trace has sampler callback 211 at 60491936136 microseconds, followed by Quick Look update callback 212 at 60491936785 in the same animation frame. The old sampler records the previous item before queued navigation updates, even though the update runs before paint. It adds a false frame of latency to the queued revision. A focused regression test reproduced the old order: observed [1], expected [2]. The fixed sampler moves its pending frame callback behind input-handler frame work on both revisions. Focused verification: Test Files 3 passed (3), Tests 10 passed (10). bun run check: svelte-check found 0 errors and 0 warnings. The requested full web test suite is running. All earlier paint counters, including the queue and image-prewarming attribution probes, are excluded from the verdict. No application revert is justified by them. The final production A1/B1/A2/B2/A3/B3 round is running under /root/perf.lock in /root/blaze-641-20261002/results-ordered, using sampler 63529a7ed and the unchanged A/B bundles and optimized release server. Source files are frozen for this round.
Author
Owner

The final review reproduced native listener ordering: Chromium drains microtasks between capture and target listeners. The capture-based fix in 63529a7ed still sampled before queued UI work. Its Files timing counts, including results-ordered, are excluded. Settings identity coverage remains valid (19,200 desktop keys, Account unchanged).

Fixed native sampling at window bubble, stopped-pass listener/frame cleanup, and rejection of pre-dispatch samples. Native browser regressions pass in Chromium and WebKit. The pure stopped-work regression passed (6 readiness tests); frame-index and cadence regressions passed (6 tests). A focused two-worker run timed out starting one worker on the build host, so readiness was rerun with one worker. The requested full test and check commands are running.

A1/B1/A2/B2/A3/B3 Files round now runs under the perf VM lock in results-final with frozen corrected source and the same production bundles and server. Starting inside-lock load: 0.02 0.29 2.64. No application revert can be justified from the discarded counts.

The final review reproduced native listener ordering: Chromium drains microtasks between capture and target listeners. The capture-based fix in 63529a7ed still sampled before queued UI work. Its Files timing counts, including results-ordered, are excluded. Settings identity coverage remains valid (19,200 desktop keys, Account unchanged). Fixed native sampling at window bubble, stopped-pass listener/frame cleanup, and rejection of pre-dispatch samples. Native browser regressions pass in Chromium and WebKit. The pure stopped-work regression passed (6 readiness tests); frame-index and cadence regressions passed (6 tests). A focused two-worker run timed out starting one worker on the build host, so readiness was rerun with one worker. The requested full test and check commands are running. A1/B1/A2/B2/A3/B3 Files round now runs under the perf VM lock in results-final with frozen corrected source and the same production bundles and server. Starting inside-lock load: 0.02 0.29 2.64. No application revert can be justified from the discarded counts.
Author
Owner

Head 63e745a55. The native-order/lifecycle sampler fix is ff104ac0a. Later changes document invariants and add BLAZE_FIRST_PAIR=2/3 so a completed pair can be retained on resume.

Final bun run check: exit 0, svelte-check found 0 errors and 0 warnings.

Requested bun run test --maxWorkers=2: exit 1. Verbatim summary:

 Test Files  3 failed | 151 passed (154)
      Tests  3 failed | 970 passed (973)
     Errors  10 errors
   Start at  23:25:58
   Duration  1591.55s (transform 39%, import 33%, environment 17%, setup 6%, tests 5%)

All three failed source-scan tests report Error: Test timed out in 5000ms. Ten workers report Timeout waiting for worker to respond. No failed assertion is reported. The unchanged shortcut-registry source test passes alone. No expectations or timeout limits were changed. All 12 focused measurement tests and the native Chromium/WebKit ordering/lifecycle regressions passed. Repeat the full suite in the merge round; this is not a full-suite pass.

The first corrected matched Files pair is mixed: A completes 41/1204 eligible targets, B 118/1093. Chromium 15 ms cases regress in frame intervals. A focused queue-removal bundle is worse (phone/15 ms p95 intervals 749.9 ms cold, 900 ms warm); it does not support reverting 4e58f2779. An image-prefetch-disabled probe is queued. B2 incurred an HDD startup delay with idle CPU and one active measurement scope. Keep startup separate from measured input phases. The final table and remaining work follow at the job time limit.

Head 63e745a55. The native-order/lifecycle sampler fix is ff104ac0a. Later changes document invariants and add BLAZE_FIRST_PAIR=2/3 so a completed pair can be retained on resume. Final `bun run check`: exit 0, `svelte-check found 0 errors and 0 warnings`. Requested `bun run test --maxWorkers=2`: exit 1. Verbatim summary: ```text Test Files 3 failed | 151 passed (154) Tests 3 failed | 970 passed (973) Errors 10 errors Start at 23:25:58 Duration 1591.55s (transform 39%, import 33%, environment 17%, setup 6%, tests 5%) ``` All three failed source-scan tests report `Error: Test timed out in 5000ms.` Ten workers report `Timeout waiting for worker to respond`. No failed assertion is reported. The unchanged shortcut-registry source test passes alone. No expectations or timeout limits were changed. All 12 focused measurement tests and the native Chromium/WebKit ordering/lifecycle regressions passed. Repeat the full suite in the merge round; this is not a full-suite pass. The first corrected matched Files pair is mixed: A completes 41/1204 eligible targets, B 118/1093. Chromium 15 ms cases regress in frame intervals. A focused queue-removal bundle is worse (phone/15 ms p95 intervals 749.9 ms cold, 900 ms warm); it does not support reverting 4e58f2779. An image-prefetch-disabled probe is queued. B2 incurred an HDD startup delay with idle CPU and one active measurement scope. Keep startup separate from measured input phases. The final table and remaining work follow at the job time limit.
Author
Owner

Final report (partial). Stopped at the four-hour job limit. The requested three-pair comparison and regression fix are not complete.

Head: c97afaec36. Branch: job/blaze-surfaces. No push, deployment or owner merge. The authorized origin/dev merge was ffaad902f, with A pinned to c4faf184df. B's production bundle is ffaad902f; later commits change the measurement code and documentation only.

Built: fixed-deadline repeats and visible-content readiness, a serialized perf-VM comparison driver, native input ordering/lifecycle regressions, stopped-pass cleanup, a pre-dispatch sample guard, and a pair-resume option. The frozen sampler producer is ff104ac0a. An independent read-only code review found no remaining definite sampler defect. No product interaction was added or changed this round.

Files: bench/blaze.mjs, bench/blaze-ab.sh, bench/blaze-native.test.mjs, bench/blaze.md; apps/web/src/lib/navigation/blazeMetrics.ts, blazeMetrics.test.ts, blazeReadiness.test.ts, blazeCadence.test.ts; docs/perf/2026-10-02-blaze-surfaces-641.md. docs/DESIGN.md changed only through the authorized origin/dev merge.

Coverage: Settings has six identity runs (19,200 desktop keys), but Account remains selected in every phase; arrow navigation is absent in both revisions, so no navigation latency is available. Phone Settings is a drill-in list. Mail lacks held-arrow message selection and a profile. Notes arrows move tree focus without opening the focused Note and have no valid profile. Those interactions were not fabricated.

Two complete Files A/B pairs were measured on root@10.69.69.63 with /root/perf.lock held for every run. Same pinned optimized server, production bundles, sampler and HDD settings: direct I/O, 8 ms read/write delay, 200 IOPS; full 1,200-file fixture, 200 keys each direction. Cold uses a restarted server/context and dropped Linux page caches; warm reuses them. Server SHA-256: 5eb0ba4775542f517948c2b396ba436df44cd9e0a0cb2ad7fd41c3b7e10e9a9d. All earlier Files counters (results, results-fixed, results-ordered) are excluded due to driver or sampler defects. No build-host timing enters this table. Only this job's measurement scope was active during the spot check. Per-run load readings are retained in the raw logs.

Provisional table (two pairs, not three):

Matched production runs: files-A1, files-A2, files-B1, files-B2

Browser Width Repeat Phase A complete/eligible by pair B complete/eligible by pair A/B p95 frame median ms A/B longest task median ms
chromium 1440 33 cold 0/63, 5/78 7/72, 13/82 266.7 / 233.3 294.0 / 98.5
chromium 1440 33 warm 4/56, 3/88 3/80, 11/73 291.6 / 249.9 602.5 / 611.5
chromium 1440 15 cold 1/33, 3/38 1/32, 4/46 174.9 / 283.3 151.0 / 81.0
chromium 1440 15 warm 2/22, 7/39 2/30, 4/47 241.6 / 283.4 1822.0 / 1135.0
chromium 390 33 cold 2/114, 8/115 10/155, 12/166 174.9 / 133.3 859.0 / 1001.5
chromium 390 33 warm 17/106, 11/132 13/173, 36/190 175.0 / 108.2 526.5 / 133.5
chromium 390 15 cold 2/85, 10/100 1/50, 2/52 99.9 / 200.0 79.0 / 241.0
chromium 390 15 warm 11/88, 11/97 2/53, 4/53 158.3 / 216.7 495.0 / 463.5
webkit 1440 33 cold 0/68, 0/65 1/46, 1/56 392.0 / 380.5 unavailable / unavailable
webkit 1440 33 warm 0/65, 0/66 4/56, 3/57 337.5 / 367.5 unavailable / unavailable
webkit 1440 15 cold 0/66, 0/70 0/27, 2/27 361.5 / 307.5 unavailable / unavailable
webkit 1440 15 warm 0/61, 0/65 0/29, 2/24 347.5 / 328.5 unavailable / unavailable
webkit 390 33 cold 0/98, 0/97 24/98, 8/79 273.5 / 219.0 unavailable / unavailable
webkit 390 33 warm 2/100, 0/91 48/116, 4/73 210.5 / 231.5 unavailable / unavailable
webkit 390 15 cold 0/97, 0/92 0/38, 0/39 248.5 / 316.0 unavailable / unavailable
webkit 390 15 warm 0/82, 0/92 2/38, 0/36 218.5 / 247.0 unavailable / unavailable

Complete means selected content ready within 33.33 ms of native event creation. Superseded same-frame targets are excluded from eligible. Frame intervals measure DOM sampling, not GPU presentation. WebKit does not expose Long Tasks.

A: complete/eligible=99/2529, keys=12800
serverCpuPercent: median=26.85, max=83.5
serverRssAverageKiB: median=354368.0, max=455000
serverRssPeakKiB: median=392106.0, max=523292
jsHeapPeakBytes: median=17150000.0, max=23100000

B: complete/eligible=224/2193, keys=12800
serverCpuPercent: median=25.200000000000003, max=44.3
serverRssAverageKiB: median=232431.0, max=458020
serverRssPeakKiB: median=346010.0, max=458020
jsHeapPeakBytes: median=15200000.0, max=21700000

Verdict: mixed and incomplete. B completes 224/2193 eligible targets (10.2%), A 99/2529 (3.9%). The 33 ms cases generally improve. Chromium phone/15 ms regresses in both pairs: cold median p95 frame interval 99.9 → 200.0 ms; warm 158.3 → 216.7 ms. Cold complete targets fall 12 → 3, warm 22 → 6. Neither revision meets the zero-incomplete-target target.

Attribution probes used the same frozen final sampler and lock. C removes only Quick Look's frame queue (4e58f2779): phone/15 ms cold 2/33 complete, p95 749.9 ms; warm 5/21, 900 ms. D disables image prefetch while retaining that queue: cold 0/63, 183.3 ms; warm 3/67, 166.7 ms. C worsens frame intervals. D improves some intervals but does not restore the reference completion rate. These isolated probes do not establish one offending commit, so no unsupported application revert was made. Attribution and repair of the fast-repeat regression remain unfinished.

Final gates, verbatim excerpts:

$ node scripts/check-user-storage.mjs && node scripts/check-type-tokens.mjs && node scripts/check-motion-tokens.mjs && svelte-kit sync && svelte-check --tsconfig ./tsconfig.json
User browser caches use userStorage; only documented device/public-link exceptions remain.
Text sizes and UI shape values use shared role tokens.
UI transitions and animation options use shared motion tokens or documented exceptions.
Loading svelte-check in workspace: /home/kayg/Developer/calternal-wt/blaze-surfaces/apps/web
Getting Svelte diagnostics...

svelte-check found 0 errors and 0 warnings

bun run test --maxWorkers=2 exited 1:

 Test Files  3 failed | 151 passed (154)
      Tests  3 failed | 970 passed (973)
     Errors  10 errors
   Start at  23:25:58
   Duration  1591.55s (transform 39%, import 33%, environment 17%, setup 6%, tests 5%)

All three source-scan failures report Error: Test timed out in 5000ms. All ten worker errors report [vitest-pool-runner]: Timeout waiting for worker to respond. No assertion failure is reported. The unchanged shortcut-registry source test passes alone. Existing expectations and timeout limits were not changed. Full log: artifacts/blaze/perf-vm-20261002/test-final.log.

Focused metrics and cadence: two files, six tests passed. The initial focused attempt lost a readiness worker at startup; readiness then passed alone with one worker (one file, six tests). Native regression, exit 0:

chromium: native sampler ordering and lifecycle passed
webkit: native sampler ordering and lifecycle passed

cargo fmt --check: exit 0, no output. No Rust crate changed, so no Rust Clippy/test run. cargo clean: exit 0, Removed 1 file, 356B total. Local web build output, baseline tree and copied archives were deleted; ignored artifacts remain. The unfinished A3 run was stopped and excluded. The perf lock was released and the resume wrapper is staged on the VM.

Known gaps: pair 3; valid Settings/Mail/Notes timing surfaces; fast-repeat regression attribution/fix; full suite completion without host timeouts. No new screenshots: no screen changed this round. Long Tasks are unavailable in WebKit; these are DOM-readiness/frame samples, not GPU presentation measurements. Browser RSS was not measured; server RSS and Chromium JS heap are reported.

Decisions: pin one existing optimized server for both web bundles; use one shared visible-content probe; use the issue's minimum 200 keys per direction while retaining the full fixture. Keep missing navigation as a coverage gap. Do not revert application behavior from an ambiguous isolated probe. No new product design decision.

UX gaps closed: none in product code; measurement false-frame, stopped-pass and pre-dispatch defects fixed. UX gaps left: Settings arrows, Mail held-arrow selection, Notes opening the focused Note, Files fast-repeat behavior.

For the merge round / remaining work: on the perf VM, use the recorded environment in /root/blaze-641-20261002 and run BLAZE_FIRST_PAIR=3 bash bench/blaze-ab.sh bundles results-final files with the staged resume wrapper. It holds the perf lock per run and keeps pairs 1/2. The sampler and bundles remain frozen. Complete the third pair, then attribute/fix the Chromium regression and obtain a three-pair verdict. Run cd apps/web && bun run test --maxWorkers=2 once on the combined branch to obtain a complete suite result; run cd apps/web && CALTERNAL_E2E_ASSET_OVERRIDE=1 bun run test:e2e:files to verify real preview identities and controls. Full Rust, adversarial, release, deployment and Mac checks remain with the merge round.

Report and raw evidence: docs/perf/2026-10-02-blaze-surfaces-641.md; artifacts/blaze/perf-vm-20261002/final/; provenance.json and stopped-round.log; VM /root/blaze-641-20261002/results-final/ and diagnostics-final/. Partial A3 has no JSON and does not enter the matched table. Issue remains open.

Final report (partial). Stopped at the four-hour job limit. The requested three-pair comparison and regression fix are not complete. Head: c97afaec365837ebdb6e30d8cb0f76418256612c. Branch: job/blaze-surfaces. No push, deployment or owner merge. The authorized origin/dev merge was ffaad902f, with A pinned to c4faf184df726a9375ae0c13bdfb6018ac2cf57e. B's production bundle is ffaad902f; later commits change the measurement code and documentation only. Built: fixed-deadline repeats and visible-content readiness, a serialized perf-VM comparison driver, native input ordering/lifecycle regressions, stopped-pass cleanup, a pre-dispatch sample guard, and a pair-resume option. The frozen sampler producer is ff104ac0a. An independent read-only code review found no remaining definite sampler defect. No product interaction was added or changed this round. Files: bench/blaze.mjs, bench/blaze-ab.sh, bench/blaze-native.test.mjs, bench/blaze.md; apps/web/src/lib/navigation/blazeMetrics.ts, blazeMetrics.test.ts, blazeReadiness.test.ts, blazeCadence.test.ts; docs/perf/2026-10-02-blaze-surfaces-641.md. docs/DESIGN.md changed only through the authorized origin/dev merge. Coverage: Settings has six identity runs (19,200 desktop keys), but Account remains selected in every phase; arrow navigation is absent in both revisions, so no navigation latency is available. Phone Settings is a drill-in list. Mail lacks held-arrow message selection and a profile. Notes arrows move tree focus without opening the focused Note and have no valid profile. Those interactions were not fabricated. Two complete Files A/B pairs were measured on root@10.69.69.63 with /root/perf.lock held for every run. Same pinned optimized server, production bundles, sampler and HDD settings: direct I/O, 8 ms read/write delay, 200 IOPS; full 1,200-file fixture, 200 keys each direction. Cold uses a restarted server/context and dropped Linux page caches; warm reuses them. Server SHA-256: 5eb0ba4775542f517948c2b396ba436df44cd9e0a0cb2ad7fd41c3b7e10e9a9d. All earlier Files counters (results, results-fixed, results-ordered) are excluded due to driver or sampler defects. No build-host timing enters this table. Only this job's measurement scope was active during the spot check. Per-run load readings are retained in the raw logs. Provisional table (two pairs, not three): Matched production runs: files-A1, files-A2, files-B1, files-B2 | Browser | Width | Repeat | Phase | A complete/eligible by pair | B complete/eligible by pair | A/B p95 frame median ms | A/B longest task median ms | |---|---:|---:|---|---|---|---:|---:| | chromium | 1440 | 33 | cold | 0/63, 5/78 | 7/72, 13/82 | 266.7 / 233.3 | 294.0 / 98.5 | | chromium | 1440 | 33 | warm | 4/56, 3/88 | 3/80, 11/73 | 291.6 / 249.9 | 602.5 / 611.5 | | chromium | 1440 | 15 | cold | 1/33, 3/38 | 1/32, 4/46 | 174.9 / 283.3 | 151.0 / 81.0 | | chromium | 1440 | 15 | warm | 2/22, 7/39 | 2/30, 4/47 | 241.6 / 283.4 | 1822.0 / 1135.0 | | chromium | 390 | 33 | cold | 2/114, 8/115 | 10/155, 12/166 | 174.9 / 133.3 | 859.0 / 1001.5 | | chromium | 390 | 33 | warm | 17/106, 11/132 | 13/173, 36/190 | 175.0 / 108.2 | 526.5 / 133.5 | | chromium | 390 | 15 | cold | 2/85, 10/100 | 1/50, 2/52 | 99.9 / 200.0 | 79.0 / 241.0 | | chromium | 390 | 15 | warm | 11/88, 11/97 | 2/53, 4/53 | 158.3 / 216.7 | 495.0 / 463.5 | | webkit | 1440 | 33 | cold | 0/68, 0/65 | 1/46, 1/56 | 392.0 / 380.5 | unavailable / unavailable | | webkit | 1440 | 33 | warm | 0/65, 0/66 | 4/56, 3/57 | 337.5 / 367.5 | unavailable / unavailable | | webkit | 1440 | 15 | cold | 0/66, 0/70 | 0/27, 2/27 | 361.5 / 307.5 | unavailable / unavailable | | webkit | 1440 | 15 | warm | 0/61, 0/65 | 0/29, 2/24 | 347.5 / 328.5 | unavailable / unavailable | | webkit | 390 | 33 | cold | 0/98, 0/97 | 24/98, 8/79 | 273.5 / 219.0 | unavailable / unavailable | | webkit | 390 | 33 | warm | 2/100, 0/91 | 48/116, 4/73 | 210.5 / 231.5 | unavailable / unavailable | | webkit | 390 | 15 | cold | 0/97, 0/92 | 0/38, 0/39 | 248.5 / 316.0 | unavailable / unavailable | | webkit | 390 | 15 | warm | 0/82, 0/92 | 2/38, 0/36 | 218.5 / 247.0 | unavailable / unavailable | Complete means selected content ready within 33.33 ms of native event creation. Superseded same-frame targets are excluded from eligible. Frame intervals measure DOM sampling, not GPU presentation. WebKit does not expose Long Tasks. A: complete/eligible=99/2529, keys=12800 serverCpuPercent: median=26.85, max=83.5 serverRssAverageKiB: median=354368.0, max=455000 serverRssPeakKiB: median=392106.0, max=523292 jsHeapPeakBytes: median=17150000.0, max=23100000 B: complete/eligible=224/2193, keys=12800 serverCpuPercent: median=25.200000000000003, max=44.3 serverRssAverageKiB: median=232431.0, max=458020 serverRssPeakKiB: median=346010.0, max=458020 jsHeapPeakBytes: median=15200000.0, max=21700000 Verdict: mixed and incomplete. B completes 224/2193 eligible targets (10.2%), A 99/2529 (3.9%). The 33 ms cases generally improve. Chromium phone/15 ms regresses in both pairs: cold median p95 frame interval 99.9 → 200.0 ms; warm 158.3 → 216.7 ms. Cold complete targets fall 12 → 3, warm 22 → 6. Neither revision meets the zero-incomplete-target target. Attribution probes used the same frozen final sampler and lock. C removes only Quick Look's frame queue (4e58f2779): phone/15 ms cold 2/33 complete, p95 749.9 ms; warm 5/21, 900 ms. D disables image prefetch while retaining that queue: cold 0/63, 183.3 ms; warm 3/67, 166.7 ms. C worsens frame intervals. D improves some intervals but does not restore the reference completion rate. These isolated probes do not establish one offending commit, so no unsupported application revert was made. Attribution and repair of the fast-repeat regression remain unfinished. Final gates, verbatim excerpts: ```text $ node scripts/check-user-storage.mjs && node scripts/check-type-tokens.mjs && node scripts/check-motion-tokens.mjs && svelte-kit sync && svelte-check --tsconfig ./tsconfig.json User browser caches use userStorage; only documented device/public-link exceptions remain. Text sizes and UI shape values use shared role tokens. UI transitions and animation options use shared motion tokens or documented exceptions. Loading svelte-check in workspace: /home/kayg/Developer/calternal-wt/blaze-surfaces/apps/web Getting Svelte diagnostics... svelte-check found 0 errors and 0 warnings ``` `bun run test --maxWorkers=2` exited 1: ```text Test Files 3 failed | 151 passed (154) Tests 3 failed | 970 passed (973) Errors 10 errors Start at 23:25:58 Duration 1591.55s (transform 39%, import 33%, environment 17%, setup 6%, tests 5%) ``` All three source-scan failures report `Error: Test timed out in 5000ms.` All ten worker errors report `[vitest-pool-runner]: Timeout waiting for worker to respond`. No assertion failure is reported. The unchanged shortcut-registry source test passes alone. Existing expectations and timeout limits were not changed. Full log: artifacts/blaze/perf-vm-20261002/test-final.log. Focused metrics and cadence: two files, six tests passed. The initial focused attempt lost a readiness worker at startup; readiness then passed alone with one worker (one file, six tests). Native regression, exit 0: ```text chromium: native sampler ordering and lifecycle passed webkit: native sampler ordering and lifecycle passed ``` `cargo fmt --check`: exit 0, no output. No Rust crate changed, so no Rust Clippy/test run. `cargo clean`: exit 0, `Removed 1 file, 356B total`. Local web build output, baseline tree and copied archives were deleted; ignored artifacts remain. The unfinished A3 run was stopped and excluded. The perf lock was released and the resume wrapper is staged on the VM. Known gaps: pair 3; valid Settings/Mail/Notes timing surfaces; fast-repeat regression attribution/fix; full suite completion without host timeouts. No new screenshots: no screen changed this round. Long Tasks are unavailable in WebKit; these are DOM-readiness/frame samples, not GPU presentation measurements. Browser RSS was not measured; server RSS and Chromium JS heap are reported. Decisions: pin one existing optimized server for both web bundles; use one shared visible-content probe; use the issue's minimum 200 keys per direction while retaining the full fixture. Keep missing navigation as a coverage gap. Do not revert application behavior from an ambiguous isolated probe. No new product design decision. UX gaps closed: none in product code; measurement false-frame, stopped-pass and pre-dispatch defects fixed. UX gaps left: Settings arrows, Mail held-arrow selection, Notes opening the focused Note, Files fast-repeat behavior. For the merge round / remaining work: on the perf VM, use the recorded environment in /root/blaze-641-20261002 and run `BLAZE_FIRST_PAIR=3 bash bench/blaze-ab.sh bundles results-final files` with the staged resume wrapper. It holds the perf lock per run and keeps pairs 1/2. The sampler and bundles remain frozen. Complete the third pair, then attribute/fix the Chromium regression and obtain a three-pair verdict. Run `cd apps/web && bun run test --maxWorkers=2` once on the combined branch to obtain a complete suite result; run `cd apps/web && CALTERNAL_E2E_ASSET_OVERRIDE=1 bun run test:e2e:files` to verify real preview identities and controls. Full Rust, adversarial, release, deployment and Mac checks remain with the merge round. Report and raw evidence: docs/perf/2026-10-02-blaze-surfaces-641.md; artifacts/blaze/perf-vm-20261002/final/; provenance.json and stopped-round.log; VM /root/blaze-641-20261002/results-final/ and diagnostics-final/. Partial A3 has no JSON and does not enter the matched table. Issue remains open.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#641
No description provided.