PERF GUARD: require route-open profiles and validate budget reports (#663) #798

Open
opened 2026-10-02 13:11:01 +00:00 by kayg · 0 comments
Owner

PERF GUARD: require route-open profiles and validate budget reports (#663)

Finding and context

origin/dev apps/web/e2e/route-perf.mjs:16–27 has a fixed route list, some empty
readiness selectors, and :66 defaults to 3 runs. #663 requires at least 5 and a
usable view at 10k items. job/blaze-settings already asserts warm Settings open
<=100 ms at settings-open-642.mjs:193. Reuse it and bench/tab-switch.mjs rather
than create another timing harness. A missing/malformed profile is a guard
failure even when generic route tests pass.

Exact detection and report validation

Every route/surface registry record names executable open and common-action
profiles. Every list profile records populated (10k) and empty states, cold
and warm separately, and a large-data burst. Validate runner discovery and
required measurement fields in CI. Each report has source SHA, production
binary hash, SPA build identity, fixture item/byte count, viewport, theme,
macOS platform, CPU settings, perf lock evidence, load average captured inside
the lock, timing-boundary definitions, all raw samples and run status.
Profile verification rejects a generic empty-container predicate for populated
runs. Usable means the first real item/detail content is present, correctly
identified and accepts its primary action. Counts are not on that boundary.

A dedicated perf-budget command checks >=5 successful samples per measured
cell, exact raw-sample/stat consistency (nearest-rank p95, median and max),
finite nonnegative times and no omitted failed/incomplete attempts. Zero,
missing, skipped or incomplete records produce status 2, never PASS. It compares
cached open <=100 ms, warm Tab switch <=100 ms, input-to-visible accepted action
<=150 ms and first usable 10k view <=1500 ms. Use p95 for the threshold, also
retain median/max; this is a proposed statistic policy because DESIGN names
all three but does not select one for comparison. Accepted and durable action
times are distinct; network durability must not stand in for accepted paint.

Measured profiles run on the perf VM using the existing shared release build,
HDD emulation and flock /root/perf.lock, recording load inside that lock. Local
runs are labelled local and never compared as equivalent VM samples. Reuse
bench/run.sh and baseline.json comparison logic. Report p50/p95/CPU/RSS next
to the matching baseline; missing baseline is explicit. Do not remove slow
samples or rerun until passing. Above-budget results exit 1 and file one
regression issue per tight group after duplicate search. Required CI enforces
profile coverage/schema and deterministic blocked-network readiness tests;
periodic timing results do not block merges under CLAUDE.md.

Tests and acceptance

Validator fixtures cover 4 samples, one failed attempt, missing binary hash,
empty-state evidence in a 10k run, mixed cold/warm samples, mismatched statistics,
NaN, exactly-on-limit pass and over-limit fail. Add test-discovery fixtures for
new/missing routes. Run once on the VM when implementing the runner. Keep
Settings #642, blaze #641 and route switch #549 as owners of their tests.

Scope and exception contract

Specification only, from #663 perf-guards. Implement in a later job. Reuse the
shared guard registry and ledger described in the coverage issue. Until it
exists, use this exact contract: an exception has rule ID, repo-relative file,
symbol or route/surface ID, normalized syntax hash (or exact build metric),
owner issue, reason, replacement test and UTC expiry date. No directory globs,
line-number-only entries or ignore all comments. Fail on expired, changed,
duplicate and unused exceptions. Existing debt needs an explicit entry and a
non-growing limit. Print exceptions in CI. An exception does not waive access,
session clearing or accessibility tests. Generated and test fixtures are
excluded by explicit source class, not substring matching. Tests must prove
both violation detection and valid exceptions. No product fix belongs here.

The static/contract command exits 1 for violations, 2 for invalid configuration
or parse/build errors, and 0 only for complete passing coverage. Output includes
rule ID, file:line, symbol/route, actual value and required value. No User content.
Measured host timing stays periodic under CLAUDE.md. Deterministic guards fail
required CI; the dedicated perf-budget runner reports measured budget failures
with nonzero status but is not a required merge check.

# PERF GUARD: require route-open profiles and validate budget reports (#663) ## Finding and context origin/dev apps/web/e2e/route-perf.mjs:16–27 has a fixed route list, some empty readiness selectors, and :66 defaults to 3 runs. #663 requires at least 5 and a usable view at 10k items. job/blaze-settings already asserts warm Settings open <=100 ms at settings-open-642.mjs:193. Reuse it and bench/tab-switch.mjs rather than create another timing harness. A missing/malformed profile is a guard failure even when generic route tests pass. ## Exact detection and report validation Every route/surface registry record names executable open and common-action profiles. Every list profile records populated (10k) and empty states, cold and warm separately, and a large-data burst. Validate runner discovery and required measurement fields in CI. Each report has source SHA, production binary hash, SPA build identity, fixture item/byte count, viewport, theme, macOS platform, CPU settings, perf lock evidence, load average captured inside the lock, timing-boundary definitions, all raw samples and run status. Profile verification rejects a generic empty-container predicate for populated runs. Usable means the first real item/detail content is present, correctly identified and accepts its primary action. Counts are not on that boundary. A dedicated perf-budget command checks >=5 successful samples per measured cell, exact raw-sample/stat consistency (nearest-rank p95, median and max), finite nonnegative times and no omitted failed/incomplete attempts. Zero, missing, skipped or incomplete records produce status 2, never PASS. It compares cached open <=100 ms, warm Tab switch <=100 ms, input-to-visible accepted action <=150 ms and first usable 10k view <=1500 ms. Use p95 for the threshold, also retain median/max; this is a proposed statistic policy because DESIGN names all three but does not select one for comparison. Accepted and durable action times are distinct; network durability must not stand in for accepted paint. Measured profiles run on the perf VM using the existing shared release build, HDD emulation and flock /root/perf.lock, recording load inside that lock. Local runs are labelled local and never compared as equivalent VM samples. Reuse bench/run.sh and baseline.json comparison logic. Report p50/p95/CPU/RSS next to the matching baseline; missing baseline is explicit. Do not remove slow samples or rerun until passing. Above-budget results exit 1 and file one regression issue per tight group after duplicate search. Required CI enforces profile coverage/schema and deterministic blocked-network readiness tests; periodic timing results do not block merges under CLAUDE.md. ## Tests and acceptance Validator fixtures cover 4 samples, one failed attempt, missing binary hash, empty-state evidence in a 10k run, mixed cold/warm samples, mismatched statistics, NaN, exactly-on-limit pass and over-limit fail. Add test-discovery fixtures for new/missing routes. Run once on the VM when implementing the runner. Keep Settings #642, blaze #641 and route switch #549 as owners of their tests. ## Scope and exception contract Specification only, from #663 perf-guards. Implement in a later job. Reuse the shared guard registry and ledger described in the coverage issue. Until it exists, use this exact contract: an exception has rule ID, repo-relative file, symbol or route/surface ID, normalized syntax hash (or exact build metric), owner issue, reason, replacement test and UTC expiry date. No directory globs, line-number-only entries or `ignore all` comments. Fail on expired, changed, duplicate and unused exceptions. Existing debt needs an explicit entry and a non-growing limit. Print exceptions in CI. An exception does not waive access, session clearing or accessibility tests. Generated and test fixtures are excluded by explicit source class, not substring matching. Tests must prove both violation detection and valid exceptions. No product fix belongs here. The static/contract command exits 1 for violations, 2 for invalid configuration or parse/build errors, and 0 only for complete passing coverage. Output includes rule ID, file:line, symbol/route, actual value and required value. No User content. Measured host timing stays periodic under CLAUDE.md. Deterministic guards fail required CI; the dedicated perf-budget runner reports measured budget failures with nonzero status but is not a required merge check.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#798
No description provided.