PERF GUARD: fail closed on surface coverage and scoped exceptions (#663) #791

Open
opened 2026-10-02 13:10:51 +00:00 by kayg · 12 comments
Owner

PERF GUARD: fail closed on surface coverage and scoped exceptions (#663)

Finding and context

At origin/dev c4a61e8cf, .forgejo/workflows/ci.yml:30–73 has no perf guard
stage. apps/web/package.json:11 has storage/type/motion checks, but no #663
coverage check. Routes and tests use hand-written lists. A new Tab or list can
therefore omit the performance contract without failing a gate. This is a
coverage finding, not a measured speed claim. Existing #669–#704 own adoption.

Exact detection and integration

Create one registry for UI surfaces and API operations. Each UI record has
stable surface ID, route pattern, implementation symbol, list/detail/action
classification, item source, first-usable predicate, empty predicate, contract
test IDs, blaze adapter ID where applicable, bench profile ID and bundle group.
Each API operation has method/path (operation ID when available), handler,
interactive-read/list/mutation/background/stream/raw-download/static class,
row and byte caps, cursor policy, cache variant and contract test IDs.

Enumerate UI routes from the SvelteKit route tree and endpoint operations from
the generated OpenAPI contract. Reuse scripts/action_registry.py and route
classification machinery. Compare router declarations too: an undocumented
route is not an exemption. Every discovered route must map to a record; every
record must map back to a live route and symbol. A route may have many surfaces.
Require arrays rendered from API-backed data and list action-registry entries
to map to a surface, so one route record cannot hide its secondary lists.
Unresolved expressions require an exact ledger entry, not a silent pass.

Fail on an unclassified route/list, absent test export, duplicate identity,
missing enabled browser adapter, skipped required test, missing profile, or
stale record. Check symbols with parsed source and test-runner discovery;
matching a text name alone is insufficient. Add one perf-lint --check entry
in required CI. Run the web part from bun run check; keep shared registry
validation in one command. Rust contract tests run per affected crate.

Tests and acceptance

Fixtures add a route, endpoint, nested list, renamed handler and skipped test;
each must fail until explicitly classified. Fixtures remove a route and leave
a stale record; it fails. Test the full ledger contract below, including a
changed syntax hash and a past expiry. Emit a machine-readable coverage report
for Calendar, Notes/Tasks, Files, Photos, Mail, Money, Settings, Search and Admin.
No live private data enters the registry. This issue owns the common schema
and runner; child guards extend it instead of adding separate exception files.

Scope and exception contract

Specification only, from #663 perf-guards. Implement in a later job. Reuse the
shared guard registry and ledger described in the coverage issue. Until it
exists, use this exact contract: an exception has rule ID, repo-relative file,
symbol or route/surface ID, normalized syntax hash (or exact build metric),
owner issue, reason, replacement test and UTC expiry date. No directory globs,
line-number-only entries or ignore all comments. Fail on expired, changed,
duplicate and unused exceptions. Existing debt needs an explicit entry and a
non-growing limit. Print exceptions in CI. An exception does not waive access,
session clearing or accessibility tests. Generated and test fixtures are
excluded by explicit source class, not substring matching. Tests must prove
both violation detection and valid exceptions. No product fix belongs here.

The static/contract command exits 1 for violations, 2 for invalid configuration
or parse/build errors, and 0 only for complete passing coverage. Output includes
rule ID, file:line, symbol/route, actual value and required value. No User content.
Measured host timing stays periodic under CLAUDE.md. Deterministic guards fail
required CI; the dedicated perf-budget runner reports measured budget failures
with nonzero status but is not a required merge check.

# PERF GUARD: fail closed on surface coverage and scoped exceptions (#663) ## Finding and context At origin/dev c4a61e8cf, `.forgejo/workflows/ci.yml:30–73` has no perf guard stage. `apps/web/package.json:11` has storage/type/motion checks, but no #663 coverage check. Routes and tests use hand-written lists. A new Tab or list can therefore omit the performance contract without failing a gate. This is a coverage finding, not a measured speed claim. Existing #669–#704 own adoption. ## Exact detection and integration Create one registry for UI surfaces and API operations. Each UI record has stable surface ID, route pattern, implementation symbol, list/detail/action classification, item source, first-usable predicate, empty predicate, contract test IDs, blaze adapter ID where applicable, bench profile ID and bundle group. Each API operation has method/path (operation ID when available), handler, interactive-read/list/mutation/background/stream/raw-download/static class, row and byte caps, cursor policy, cache variant and contract test IDs. Enumerate UI routes from the SvelteKit route tree and endpoint operations from the generated OpenAPI contract. Reuse `scripts/action_registry.py` and route classification machinery. Compare router declarations too: an undocumented route is not an exemption. Every discovered route must map to a record; every record must map back to a live route and symbol. A route may have many surfaces. Require arrays rendered from API-backed data and list action-registry entries to map to a surface, so one route record cannot hide its secondary lists. Unresolved expressions require an exact ledger entry, not a silent pass. Fail on an unclassified route/list, absent test export, duplicate identity, missing enabled browser adapter, skipped required test, missing profile, or stale record. Check symbols with parsed source and test-runner discovery; matching a text name alone is insufficient. Add one `perf-lint --check` entry in required CI. Run the web part from bun run check; keep shared registry validation in one command. Rust contract tests run per affected crate. ## Tests and acceptance Fixtures add a route, endpoint, nested list, renamed handler and skipped test; each must fail until explicitly classified. Fixtures remove a route and leave a stale record; it fails. Test the full ledger contract below, including a changed syntax hash and a past expiry. Emit a machine-readable coverage report for Calendar, Notes/Tasks, Files, Photos, Mail, Money, Settings, Search and Admin. No live private data enters the registry. This issue owns the common schema and runner; child guards extend it instead of adding separate exception files. ## Scope and exception contract Specification only, from #663 perf-guards. Implement in a later job. Reuse the shared guard registry and ledger described in the coverage issue. Until it exists, use this exact contract: an exception has rule ID, repo-relative file, symbol or route/surface ID, normalized syntax hash (or exact build metric), owner issue, reason, replacement test and UTC expiry date. No directory globs, line-number-only entries or `ignore all` comments. Fail on expired, changed, duplicate and unused exceptions. Existing debt needs an explicit entry and a non-growing limit. Print exceptions in CI. An exception does not waive access, session clearing or accessibility tests. Generated and test fixtures are excluded by explicit source class, not substring matching. Tests must prove both violation detection and valid exceptions. No product fix belongs here. The static/contract command exits 1 for violations, 2 for invalid configuration or parse/build errors, and 0 only for complete passing coverage. Output includes rule ID, file:line, symbol/route, actual value and required value. No User content. Measured host timing stays periodic under CLAUDE.md. Deterministic guards fail required CI; the dedicated perf-budget runner reports measured budget failures with nonzero status but is not a required merge check.
Author
Owner

Started job/perfguards-impl at base 2f4482ded0. Reading #791–#798 and the audit before implementing the common registry and exception ledger. No push or deploy.

Started job/perfguards-impl at base 2f4482ded066d9c5d9c59130377907f7fd2916c9. Reading #791–#798 and the audit before implementing the common registry and exception ledger. No push or deploy.
Author
Owner

Framework commit 84a0bb835 and child-rule fixture slice are committed. The fixture suite passes 54 tests. Native tree-sitter 0.26.0 crashed while walking the full Rust source; verified 0.25.2 is pinned instead. Existing split router declarations require method-qualified identities. Full registry seeding and per-crate Rust gates are in progress. Unknown adoption remains explicit debt; these source checks do not claim measured speed.

Framework commit 84a0bb835 and child-rule fixture slice are committed. The fixture suite passes 54 tests. Native tree-sitter 0.26.0 crashed while walking the full Rust source; verified 0.25.2 is pinned instead. Existing split router declarations require method-qualified identities. Full registry seeding and per-crate Rust gates are in progress. Unknown adoption remains explicit debt; these source checks do not claim measured speed.
Author
Owner

The guard suite passes 69 tests; the existing Tab-switch harness passes 2 tests under bun test. The initial production shell is 369,604 gzip bytes versus the strict 200,000-byte budget; its exact exception belongs to #497. Guard-tool report validation measured locally: small-input p95 0.194 ms; 10k-cell burst 87.1 ms; peak RSS 74,010,624 bytes. No matching baseline exists. The perf VM lock was occupied. The populated local benchmark and Rust gates remain in progress on a host with load average above 100. Source-capture collisions are fixed before accepting the initial ledger.

The guard suite passes 69 tests; the existing Tab-switch harness passes 2 tests under bun test. The initial production shell is 369,604 gzip bytes versus the strict 200,000-byte budget; its exact exception belongs to #497. Guard-tool report validation measured locally: small-input p95 0.194 ms; 10k-cell burst 87.1 ms; peak RSS 74,010,624 bytes. No matching baseline exists. The perf VM lock was occupied. The populated local benchmark and Rust gates remain in progress on a host with load average above 100. Source-capture collisions are fixed before accepting the initial ledger.
Author
Owner

perfguards-impl / #798 verification finding (2026-10-02)

One local run reused bench/tab-switch.mjs and the shared release binary. The perf VM lock was occupied. The command requested Chromium, 1440 px, light, five warm Calendar → Files samples, macOS platform emulation, and the existing large Home fixture.

Setup failed before any route sample:

Error: Photos indexed 0 of 50000 fixture Items before the timeout
    at waitForPhotoIndex (.../bench/tab-switch.mjs:1158:12)
    at async main (.../bench/tab-switch.mjs:1375:26)

The build host load average was above 100 during the run. This is not a measured route regression and is not a PASS. I did not rerun. The initial legacy harness did not save a report for a setup failure. #798 now fixes that harness gap: setup and later failures save an incomplete report, preserve collected samples, redact credentials, and exit 2. Four harness tests pass, including this case. The populated profile setup remains unverified. Please investigate the shared-build/Home/projection compatibility in the existing #549 profile before accepting timing numbers.

perfguards-impl / #798 verification finding (2026-10-02) One local run reused `bench/tab-switch.mjs` and the shared release binary. The perf VM lock was occupied. The command requested Chromium, 1440 px, light, five warm Calendar → Files samples, macOS platform emulation, and the existing large Home fixture. Setup failed before any route sample: ``` Error: Photos indexed 0 of 50000 fixture Items before the timeout at waitForPhotoIndex (.../bench/tab-switch.mjs:1158:12) at async main (.../bench/tab-switch.mjs:1375:26) ``` The build host load average was above 100 during the run. This is not a measured route regression and is not a PASS. I did not rerun. The initial legacy harness did not save a report for a setup failure. #798 now fixes that harness gap: setup and later failures save an incomplete report, preserve collected samples, redact credentials, and exit 2. Four harness tests pass, including this case. The populated profile setup remains unverified. Please investigate the shared-build/Home/projection compatibility in the existing #549 profile before accepting timing numbers.
Author
Owner

Implementation committed through 6f53b2041c2bdfc10b561df42a91763c10d2614b on job/perfguards-impl.

The registry discovers 56 SvelteKit page modules, 261 nested lists, 309 router declarations and 340 OpenAPI operations. Initial source debt has 10,223 exact exceptions; the bundle has one #497 exception. Each expires 2026-11-16. Parser updates now bind Utoipa annotations across doc comments and decode escaped SQL fragments. Checks do not update limits.

Verbatim gate output:

perf-lint: PASS; 0 violations; 10223 scoped exceptions
perf-lint: PASS; 0 violations; 1 scoped exceptions
Ran 78 tests in 0.196s

OK

The source and bundle checks passed; the 78 fixture tests include the final populated-readiness predicate check. The final web type check is running. Per-crate Cargo clippy continues to compile dependencies on a host with load above 100. Cargo test is queued behind the build lock. The standalone server guard test passed (1 test, 368.29s); this is supplemental evidence, not a completed Cargo gate.

The full web suite returned four 5-second timeouts: 1,069 tests passed. No expectations or timeouts were changed. One local benchmark run produced no samples: Photos indexed 0 of 50,000 fixture Items before setup timed out. Evidence is on #549. The perf VM lock was busy; no unlocked measurement was taken.

Decisions: strict decimal 200,000-byte shell limit; initial route JS/CSS/lazy ceilings are measured value plus one for a strict comparison; p95 uses nearest rank; Python tooling uses existing uv and pinned MIT parser packages. Missing product adoption tests/profiles remain visible, bounded debt. Existing auth/session/access test suites stay mandatory; the per-surface adoption checks do not prove those behaviors.

Implementation committed through `6f53b2041c2bdfc10b561df42a91763c10d2614b` on `job/perfguards-impl`. The registry discovers 56 SvelteKit page modules, 261 nested lists, 309 router declarations and 340 OpenAPI operations. Initial source debt has 10,223 exact exceptions; the bundle has one #497 exception. Each expires 2026-11-16. Parser updates now bind Utoipa annotations across doc comments and decode escaped SQL fragments. Checks do not update limits. Verbatim gate output: ``` perf-lint: PASS; 0 violations; 10223 scoped exceptions perf-lint: PASS; 0 violations; 1 scoped exceptions Ran 78 tests in 0.196s OK ``` The source and bundle checks passed; the 78 fixture tests include the final populated-readiness predicate check. The final web type check is running. Per-crate Cargo clippy continues to compile dependencies on a host with load above 100. Cargo test is queued behind the build lock. The standalone server guard test passed (1 test, 368.29s); this is supplemental evidence, not a completed Cargo gate. The full web suite returned four 5-second timeouts: 1,069 tests passed. No expectations or timeouts were changed. One local benchmark run produced no samples: Photos indexed 0 of 50,000 fixture Items before setup timed out. Evidence is on #549. The perf VM lock was busy; no unlocked measurement was taken. Decisions: strict decimal 200,000-byte shell limit; initial route JS/CSS/lazy ceilings are measured value plus one for a strict comparison; p95 uses nearest rank; Python tooling uses existing uv and pinned MIT parser packages. Missing product adoption tests/profiles remain visible, bounded debt. Existing auth/session/access test suites stay mandatory; the per-surface adoption checks do not prove those behaviors.
Author
Owner

Head 113d2c7bb640089ab6091ec575fa8e3944d68768 is committed and the worktree is clean.

Final source output, verbatim:

perf-lint: PASS; 0 violations; 17956 scoped exceptions

Bundle output, verbatim:

perf-lint: PASS; 0 violations; 1 scoped exceptions

Fixture output, verbatim:

...........................................................................................
----------------------------------------------------------------------
Ran 91 tests in 0.056s

OK

Web output, verbatim:

svelte-check found 0 errors and 0 warnings

Review findings closed: receiver method names and macro expansion no longer count as IO safety evidence; Rust ignore/conditional tests and Python class skips cannot satisfy evidence; array spreads cannot count as fixed enums; blaze reports must cover the full registered matrix; missing/incomplete overall status fails; resident model row/byte bounds and finite identity-keyed cache variants are separate checks. The explicit initial capture added 6,968 unresolved receiver/macro sites and 765 missing model-bound contracts from unchanged product source. There are now 17,957 total exact exceptions, each expiring 2026-11-16.

Required gates are not all green: the web suite returned four 5-second timeouts (1,069 passed); server clippy still compiles dependencies and server tests wait for its build lock. No expectations or timeout settings were changed. The job will stop unfinished Cargo work and clean its target before the five-hour limit. Full final evidence will follow here.

Head `113d2c7bb640089ab6091ec575fa8e3944d68768` is committed and the worktree is clean. Final source output, verbatim: ``` perf-lint: PASS; 0 violations; 17956 scoped exceptions ``` Bundle output, verbatim: ``` perf-lint: PASS; 0 violations; 1 scoped exceptions ``` Fixture output, verbatim: ``` ........................................................................................... ---------------------------------------------------------------------- Ran 91 tests in 0.056s OK ``` Web output, verbatim: ``` svelte-check found 0 errors and 0 warnings ``` Review findings closed: receiver method names and macro expansion no longer count as IO safety evidence; Rust ignore/conditional tests and Python class skips cannot satisfy evidence; array spreads cannot count as fixed enums; blaze reports must cover the full registered matrix; missing/incomplete overall status fails; resident model row/byte bounds and finite identity-keyed cache variants are separate checks. The explicit initial capture added 6,968 unresolved receiver/macro sites and 765 missing model-bound contracts from unchanged product source. There are now 17,957 total exact exceptions, each expiring 2026-11-16. Required gates are not all green: the web suite returned four 5-second timeouts (1,069 passed); server clippy still compiles dependencies and server tests wait for its build lock. No expectations or timeout settings were changed. The job will stop unfinished Cargo work and clean its target before the five-hour limit. Full final evidence will follow here.
Author
Owner

Built the #791–#798 guard rollout; required verification is incomplete on job/perfguards-impl.
Head: 586d0ea3a38641d0520c7ed1eaee4d21917eea7c. Base: 2f4482ded066d9c5d9c59130377907f7fd2916c9.
git fetch origin and git merge origin/dev ran once. Merge output:

Already up to date.

No push or deploy. No product route or UI change.

Built:

  • One parsed coverage registry and exact owner/expiry debt ledger. Checks fail closed on parse errors, new/stale coverage, missing/disabled tests, changed/expired/duplicate/unused debt and growing limits.
  • SQL/keyset/list-bound and interactive IO guards; layout/material guards; finite identity-keyed finite identity-keyed revision-cache and synchronous-selection adoption contracts; nested render and separate resident model row/byte bounds, virtualization and strict blaze report coverage.
  • Production static manifest closures, preload validation and reproducible gzip shell/route/CSS/lazy-group budgets.
  • Optional measured-budget validation with raw attempts, build/fixture/environment identity, CPU/RSS, nearest-rank p95, incomplete-run status 2 and report retention. Timing thresholds do not enter mandatory CI.
  • Mandatory web check/test hooks, a server integration test and CI bundle validation. Pinned MIT syntax parsers use the existing uv tool.

Coverage: 56 SvelteKit page modules, 261 nested lists, 309 router declarations and 340 OpenAPI operations. Final initial-capture corrections include 6,968 unresolved receiver/macro sites and 765 missing resident-model bounds/tests. A final routing capture added 72 unbound method/opaque-service records. Known candidate handler bodies are pinned; this does not claim full dynamic Rust resolution. These were existing sites on the unchanged product source. The initial ledger has 18,028 source exceptions and one shell exception. Each expires 2026-11-16. Current shell: 369,604 gzip bytes, capped at that value with #497 ownership. Initial route JS/CSS/lazy ceilings prohibit growth. Initial rollout capture corrections were explicit; normal checks never change accepted inputs.

Files:
scripts/perf-lint, scripts/perf_guards/*, apps/web/scripts/perf-source.mjs, apps/web/scripts/perf-source.test.mjs, contracts/perf/{registry,exceptions}.json, docs/perf/guards.md, .gitignore, apps/web/package.json, .forgejo/workflows/ci.yml, crates/calternal-server/tests/perf_guards.rs, bench/perf-guards.py, bench/tab-switch.mjs, bench/tab-switch.test.mjs.

Gate output, verbatim (full output remains in worktree artifacts/):
Source (artifacts/perf-check-final.log):

perf-lint: PASS; 0 violations; 18028 scoped exceptions

Bundle (artifacts/perf-bundle-final.log):

DEBT bundle.js apps/web/package.json shell:js limit=369604 owner=https://git.kayg.org/kayg/calternal/issues/497 expires=2026-11-16
perf-lint: PASS; 0 violations; 1 scoped exceptions

Guard fixture tests (artifacts/perf-test-final.log):

............................................................................................
----------------------------------------------------------------------
Ran 92 tests in 0.028s

OK

Web check, exit 0 (artifacts/bun-check-final.log):

svelte-check found 0 errors and 0 warnings

Rust formatting: exit 0, no output (artifacts/cargo-fmt-final.log).
Rust clippy: incomplete, stopped for time-box cleanup; exit 143. It was still compiling dependencies (artifacts/cargo-clippy-early.log). Last output, verbatim:

    Checking ed25519-compact v2.4.2
    Checking tantivy-columnar v0.7.0
    Checking hyper v0.14.32
   Compiling sha2 v0.11.0
    Checking tantivy-query-grammar v0.26.0
    Checking superboring v0.1.14
    Checking markup5ever v0.40.0
    Checking tokio-native-tls v0.3.1

Rust tests: incomplete, stopped for time-box cleanup; exit 143. It never obtained the build lock. Output, verbatim:

    Blocking waiting for file lock on package cache
    Blocking waiting for file lock on package cache
    Blocking waiting for file lock on package cache
    Blocking waiting for file lock on build directory

Full web tests ran once (bun run test --maxWorkers=2, no assertions or timeouts changed):

 Test Files  3 failed | 151 passed (154)
      Tests  4 failed | 1069 passed (1073)
   Start at  16:11:41
   Duration  933.08s (transform 37%, environment 28%, import 18%, tests 13%, setup 4%)

Environment  |component| jsdom was created 49 times · 392.62s total, 38% of tracked time
             create it once per worker with pool: 'vmThreads' (keeps per-file isolation) or isolate: false (shares it across files)
             learn more: https://vitest.dev/guide/improving-performance#test-environments

error: script "test" exited with code 1

All four failures report Error: Test timed out in 5000ms. They affect tray keyboard focus, KeyboardShortcutsCard grouping/Escape, and two ConnectedAccountsSection provider-setup interactions. The host load was above 100. These are timeout evidence, not proven functional regressions. 1,069 tests passed.

Supplemental Rust evidence, not a substitute for Cargo gates (artifacts/rust-guard-standalone.log):

running 1 test
test deterministic_performance_guards has been running for over 60 seconds
test deterministic_performance_guards ... ok

test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 368.29s

The same Rust integration file also passed standalone clippy-driver with -D warnings (no output).
The final source check ran after the last router-binding change; the web type check and production build ran before that Python-only change. Their inputs did not change. The two real Svelte compiler fixtures pass; the four Tab-switch fixtures pass. The former rejects [...rows] as a fixed-enum bound. The latter checks Mac identity, retained target identity, redaction and incomplete-run sample retention. These tests are wired into the guard/web test entry points.

Adversarial round:
One bounded round covered exact/stale/expired/duplicate debt, skipped test discovery, SQL escaped fragments and malformed queries, helper/layout/IO dispatch, bundle preloads and malformed/incomplete reports. Final CLI probes reject omitted attempts and unregistered blaze adapters. Missing/incomplete overall run status is invalid even if retained samples are complete. The full blaze matrix cannot hide missing browser/width/theme/cadence/cold-warm cells. No API changed, so no new API probe was required.

Performance:
The perf VM lock probe returned busy. No unlocked measurement ran. One local existing Tab-switch profile used the shared release binary with 100k Files and 50k Photos. Setup failed before samples: Photos indexed 0 of 50000 fixture Items before the timeout. Evidence was posted on existing #549; artifacts/perf-local-798-incomplete.json retains the incomplete status. There are no route p50/p95/CPU/RSS values or valid baseline comparison from this run.
The added guard-tool profile ran locally once (source 6a1e46bce5152777d8d992596fcdc41654afef30, artifacts/perf-guard-tools.json): p50 0.036199 ms, p95 0.193787 ms; 10k-cell burst 87.121306 ms, CPU 84.606454 ms, RSS 74,010,624 bytes. Matching docs/perf/baseline.json entry: absent. This measures validator tooling, not User interaction or VM route performance.

Known gaps:

  • Full per-crate Rust clippy/test gates did not finish within the job limit. The standalone hook evidence does not replace these gates. Required gates remain incomplete.
  • The full web suite has four SLOW-only timeouts. No rerun-until-pass or changed expectations.
  • Product receipt/cache/virtualizer/profile/blaze adoption is not fabricated. Missing per-surface suites and profiles are explicit pinned debt. The queued blaze harness/adapters are absent from this base; the CLI rejects an unregistered adapter. The canonical report validator is implemented, but there is no completed real blaze/profile run today.
  • Existing auth/session/access suites are mandatory and cannot be waived. Per-surface session-clearing adoption tests remain separately pinned debt; their replacement fixture proves pinning, not session behavior. This interpretation needs owner review against the stricter no-waiver rule.
  • Static Rust/TypeScript analysis is conservative, not complete type inference. Ambiguous dispatch/macros and runtime style construction need exact debt. Initial owner links require adoption-owner review.

Decisions:

  • Use 200,000 decimal compressed JS bytes with strict <; initial route/CSS/lazy ceilings use measured value plus one for a strict comparison.
  • Use nearest-rank p95 for budget comparisons and retain median/max/all attempts.
  • Use 100 rendered items as the proposed virtualizer threshold, with numeric bounds at 390/820/1440.
  • Use existing uv plus pinned Python parsers and installed Svelte/TypeScript compilers. Do not add product dependencies or duplicate queued product primitives.
  • Keep unknown bounds/readiness/tests/profile fields null with 45-day exact debt. Fail closed on unresolved analysis and incomplete reports.
  • Treat conditional Rust tests as unknown/disabled evidence unless they are unconditional cfg(test) tests. Python replacements must be real unittest methods.

UX gaps closed: no product UI change. Tooling now retains incomplete reports, uses Mac identity in sign-in and measurement contexts, and checks populated item/action readiness.
UX gaps left: no new product screen; per-surface primary-action/readiness checks remain adoption debt.
Cleanup: Cargo clean completed, verbatim:

     Removed 6331 files, 2.0GiB total

Removed own web build output and temporary fixture data. Review artifacts remain untracked. The worktree is clean.

Built the #791–#798 guard rollout; required verification is incomplete on `job/perfguards-impl`. Head: `586d0ea3a38641d0520c7ed1eaee4d21917eea7c`. Base: `2f4482ded066d9c5d9c59130377907f7fd2916c9`. `git fetch origin` and `git merge origin/dev` ran once. Merge output: ``` Already up to date. ``` No push or deploy. No product route or UI change. Built: - One parsed coverage registry and exact owner/expiry debt ledger. Checks fail closed on parse errors, new/stale coverage, missing/disabled tests, changed/expired/duplicate/unused debt and growing limits. - SQL/keyset/list-bound and interactive IO guards; layout/material guards; finite identity-keyed finite identity-keyed revision-cache and synchronous-selection adoption contracts; nested render and separate resident model row/byte bounds, virtualization and strict blaze report coverage. - Production static manifest closures, preload validation and reproducible gzip shell/route/CSS/lazy-group budgets. - Optional measured-budget validation with raw attempts, build/fixture/environment identity, CPU/RSS, nearest-rank p95, incomplete-run status 2 and report retention. Timing thresholds do not enter mandatory CI. - Mandatory web check/test hooks, a server integration test and CI bundle validation. Pinned MIT syntax parsers use the existing uv tool. Coverage: 56 SvelteKit page modules, 261 nested lists, 309 router declarations and 340 OpenAPI operations. Final initial-capture corrections include 6,968 unresolved receiver/macro sites and 765 missing resident-model bounds/tests. A final routing capture added 72 unbound method/opaque-service records. Known candidate handler bodies are pinned; this does not claim full dynamic Rust resolution. These were existing sites on the unchanged product source. The initial ledger has 18,028 source exceptions and one shell exception. Each expires 2026-11-16. Current shell: 369,604 gzip bytes, capped at that value with #497 ownership. Initial route JS/CSS/lazy ceilings prohibit growth. Initial rollout capture corrections were explicit; normal checks never change accepted inputs. Files: `scripts/perf-lint`, `scripts/perf_guards/*`, `apps/web/scripts/perf-source.mjs`, `apps/web/scripts/perf-source.test.mjs`, `contracts/perf/{registry,exceptions}.json`, `docs/perf/guards.md`, `.gitignore`, `apps/web/package.json`, `.forgejo/workflows/ci.yml`, `crates/calternal-server/tests/perf_guards.rs`, `bench/perf-guards.py`, `bench/tab-switch.mjs`, `bench/tab-switch.test.mjs`. Gate output, verbatim (full output remains in worktree `artifacts/`): Source (`artifacts/perf-check-final.log`): ``` perf-lint: PASS; 0 violations; 18028 scoped exceptions ``` Bundle (`artifacts/perf-bundle-final.log`): ``` DEBT bundle.js apps/web/package.json shell:js limit=369604 owner=https://git.kayg.org/kayg/calternal/issues/497 expires=2026-11-16 perf-lint: PASS; 0 violations; 1 scoped exceptions ``` Guard fixture tests (`artifacts/perf-test-final.log`): ``` ............................................................................................ ---------------------------------------------------------------------- Ran 92 tests in 0.028s OK ``` Web check, exit 0 (`artifacts/bun-check-final.log`): ``` svelte-check found 0 errors and 0 warnings ``` Rust formatting: exit 0, no output (`artifacts/cargo-fmt-final.log`). Rust clippy: incomplete, stopped for time-box cleanup; exit 143. It was still compiling dependencies (`artifacts/cargo-clippy-early.log`). Last output, verbatim: ``` Checking ed25519-compact v2.4.2 Checking tantivy-columnar v0.7.0 Checking hyper v0.14.32 Compiling sha2 v0.11.0 Checking tantivy-query-grammar v0.26.0 Checking superboring v0.1.14 Checking markup5ever v0.40.0 Checking tokio-native-tls v0.3.1 ``` Rust tests: incomplete, stopped for time-box cleanup; exit 143. It never obtained the build lock. Output, verbatim: ``` Blocking waiting for file lock on package cache Blocking waiting for file lock on package cache Blocking waiting for file lock on package cache Blocking waiting for file lock on build directory ``` Full web tests ran once (`bun run test --maxWorkers=2`, no assertions or timeouts changed): ``` Test Files 3 failed | 151 passed (154) Tests 4 failed | 1069 passed (1073) Start at 16:11:41 Duration 933.08s (transform 37%, environment 28%, import 18%, tests 13%, setup 4%) Environment |component| jsdom was created 49 times · 392.62s total, 38% of tracked time create it once per worker with pool: 'vmThreads' (keeps per-file isolation) or isolate: false (shares it across files) learn more: https://vitest.dev/guide/improving-performance#test-environments error: script "test" exited with code 1 ``` All four failures report `Error: Test timed out in 5000ms.` They affect tray keyboard focus, KeyboardShortcutsCard grouping/Escape, and two ConnectedAccountsSection provider-setup interactions. The host load was above 100. These are timeout evidence, not proven functional regressions. 1,069 tests passed. Supplemental Rust evidence, not a substitute for Cargo gates (`artifacts/rust-guard-standalone.log`): ``` running 1 test test deterministic_performance_guards has been running for over 60 seconds test deterministic_performance_guards ... ok test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 368.29s ``` The same Rust integration file also passed standalone clippy-driver with `-D warnings` (no output). The final source check ran after the last router-binding change; the web type check and production build ran before that Python-only change. Their inputs did not change. The two real Svelte compiler fixtures pass; the four Tab-switch fixtures pass. The former rejects `[...rows]` as a fixed-enum bound. The latter checks Mac identity, retained target identity, redaction and incomplete-run sample retention. These tests are wired into the guard/web test entry points. Adversarial round: One bounded round covered exact/stale/expired/duplicate debt, skipped test discovery, SQL escaped fragments and malformed queries, helper/layout/IO dispatch, bundle preloads and malformed/incomplete reports. Final CLI probes reject omitted attempts and unregistered blaze adapters. Missing/incomplete overall run status is invalid even if retained samples are complete. The full blaze matrix cannot hide missing browser/width/theme/cadence/cold-warm cells. No API changed, so no new API probe was required. Performance: The perf VM lock probe returned busy. No unlocked measurement ran. One local existing Tab-switch profile used the shared release binary with 100k Files and 50k Photos. Setup failed before samples: `Photos indexed 0 of 50000 fixture Items before the timeout`. Evidence was posted on existing #549; `artifacts/perf-local-798-incomplete.json` retains the incomplete status. There are no route p50/p95/CPU/RSS values or valid baseline comparison from this run. The added guard-tool profile ran locally once (source `6a1e46bce5152777d8d992596fcdc41654afef30`, `artifacts/perf-guard-tools.json`): p50 0.036199 ms, p95 0.193787 ms; 10k-cell burst 87.121306 ms, CPU 84.606454 ms, RSS 74,010,624 bytes. Matching `docs/perf/baseline.json` entry: absent. This measures validator tooling, not User interaction or VM route performance. Known gaps: - Full per-crate Rust clippy/test gates did not finish within the job limit. The standalone hook evidence does not replace these gates. Required gates remain incomplete. - The full web suite has four SLOW-only timeouts. No rerun-until-pass or changed expectations. - Product receipt/cache/virtualizer/profile/blaze adoption is not fabricated. Missing per-surface suites and profiles are explicit pinned debt. The queued blaze harness/adapters are absent from this base; the CLI rejects an unregistered adapter. The canonical report validator is implemented, but there is no completed real blaze/profile run today. - Existing auth/session/access suites are mandatory and cannot be waived. Per-surface session-clearing adoption tests remain separately pinned debt; their replacement fixture proves pinning, not session behavior. This interpretation needs owner review against the stricter no-waiver rule. - Static Rust/TypeScript analysis is conservative, not complete type inference. Ambiguous dispatch/macros and runtime style construction need exact debt. Initial owner links require adoption-owner review. Decisions: - Use 200,000 decimal compressed JS bytes with strict `<`; initial route/CSS/lazy ceilings use measured value plus one for a strict comparison. - Use nearest-rank p95 for budget comparisons and retain median/max/all attempts. - Use 100 rendered items as the proposed virtualizer threshold, with numeric bounds at 390/820/1440. - Use existing uv plus pinned Python parsers and installed Svelte/TypeScript compilers. Do not add product dependencies or duplicate queued product primitives. - Keep unknown bounds/readiness/tests/profile fields null with 45-day exact debt. Fail closed on unresolved analysis and incomplete reports. - Treat conditional Rust tests as unknown/disabled evidence unless they are unconditional `cfg(test)` tests. Python replacements must be real unittest methods. UX gaps closed: no product UI change. Tooling now retains incomplete reports, uses Mac identity in sign-in and measurement contexts, and checks populated item/action readiness. UX gaps left: no new product screen; per-surface primary-action/readiness checks remain adoption debt. Cleanup: Cargo clean completed, verbatim: ``` Removed 6331 files, 2.0GiB total ``` Removed own web build output and temporary fixture data. Review artifacts remain untracked. The worktree is clean.
Author
Owner

Independent read-only review: request changes.

Review branch: job/rev2-perfguards; base
c4faf184df726a9375ae0c13bdfb6018ac2cf57e; review head
d9b906b18bafa32d20a89ae8a9cc70b68ec7b98c.
Implementation reviewed: 586d0ea3a38641d0520c7ed1eaee4d21917eea7c.
Files committed: audit-findings.md, review-perfguards.md. No product change.

P1 — F1: Session clearing tests can be waived. scripts/perf_guards/core.py:70
rejects session.*, but line 99 emits contract.session-clear. The ledger
contains 317 such waivers; the first is contracts/perf/exceptions.json:146.
#791 forbids these waivers. Reject the emitted rule, add an exact negative
fixture and supply real session clearing contracts. Do not rename the rule.

P2 — F3: Test names are not test-runner discovery.
apps/web/scripts/perf-source.mjs:25–34 accepts declarations inside unused
helpers and test.runIf(false) as enabled. It does not use runner include
patterns. Use runner collection metadata, or reject unsupported declarations.
Test dead helpers, conditional skips and excluded files.

P3 — F6: scripts/perf_guards/runner.py:56 maps Settings debt to #701, which
owns Notes. Settings shared-contract adoption is #674. Use an explicit table
by surface and rule. Review initial owner links and test this mapping.

Other P2 findings were added to their existing owners:

  • #794, F2: rules.py:176–180 checks awaits only in named arrow-key source;
    the Svelte exporter checks AwaitBlock. Ordinary open/cache-return animation
    waits are unchecked. Add call resolution and blocked-animation fixtures.
  • #795, F4: rules.py:171–172 flags all fetch/apiFetch calls with no registry
    acceptance path for a valid adapter, mutation or download. core.py:180
    validates key strings without binding them to the adapter. Add typed
    bindings and complete positive and negative fixtures.
  • #793, F5: rules.py:86–88,114–115 always flags external calls. Projection
    tests cannot resolve them. The positive SQL fixture at test_rules.py:153
    checks only one absent rule. Add reviewed capability bindings and prove a
    complete projected read passes with no ledger.

Ledger verdict: 18,028 source waivers plus one bundle waiver, all expiring
2026-11-16. Changed, duplicate, expired and unused entries fail. They are not
permanent blanket waivers. There is no staged shrink requirement or frozen
total across ledger edits. All use one generic synthetic pinning test, not
18,029 product contracts. 12,703 entries are unresolved calls; F4/F5 prevent
valid adoption from clearing some debt.

Detection verdict: new EachBlocks and list endpoints receive useful coverage
and cap checks. New queries usually fail through changed debt or unresolved
calls, but no query-count rule proves N+1 behavior. Ordinary animation waits
and unused revision-key declarations remain gaps.

Four saved web failures are all Error: Test timed out in 5000ms. Guard
fixtures run before Vitest and do not enter those components. The integrations
branch supplied Connected Accounts changes. No direct causal evidence ties
these timeouts to guards. Host load remains a possible cause, not a diagnosis.
Saved output, verbatim:

 Test Files  3 failed | 151 passed (154)
      Tests  4 failed | 1069 passed (1073)
   Start at  16:11:41
   Duration  933.08s (transform 37%, environment 28%, import 18%, tests 13%, setup 4%)

No builds or tests ran in this LIGHT review. Prior saved gate output, verbatim:

perf-lint: PASS; 0 violations; 18028 scoped exceptions

Ran git fetch origin && git merge origin/dev once. Output: Already up to date.
Read both review documents again and ran git diff --check with no output.

Known gaps: static findings have not been reproduced by execution. The six
fixes remain open. No performance or visual claim. No new issue or duplicate
owner was created; existing owners are #791, #793, #794 and #795.

Decisions: DESIGN §58 in this tree is Agent discovery (#630), so use the
explicit #663/#791–#795 requirements and the owner loading rule. Use existing
owners. Leave timeout assertions and limits unchanged.

For the merge round, after fixes:
scripts/perf-lint test and
scripts/perf-lint --check --report artifacts/perf-coverage.json must prove
exact no-waiver enforcement, real discovery and valid adapter/projection
acceptance. cargo test -p calternal-server --test perf_guards must prove the
shared hook. Run
cd apps/web && bunx vitest run src/lib/tray.svelte.test.ts src/lib/components/KeyboardShortcutsCard.svelte.test.ts src/routes/settings/connected-accounts/ConnectedAccountsSection.svelte.test.ts --maxWorkers=2
to diagnose the four recorded timeouts without changed expectations.

Independent read-only review: request changes. Review branch: `job/rev2-perfguards`; base `c4faf184df726a9375ae0c13bdfb6018ac2cf57e`; review head `d9b906b18bafa32d20a89ae8a9cc70b68ec7b98c`. Implementation reviewed: `586d0ea3a38641d0520c7ed1eaee4d21917eea7c`. Files committed: `audit-findings.md`, `review-perfguards.md`. No product change. P1 — F1: Session clearing tests can be waived. `scripts/perf_guards/core.py:70` rejects `session.*`, but line 99 emits `contract.session-clear`. The ledger contains 317 such waivers; the first is `contracts/perf/exceptions.json:146`. #791 forbids these waivers. Reject the emitted rule, add an exact negative fixture and supply real session clearing contracts. Do not rename the rule. P2 — F3: Test names are not test-runner discovery. `apps/web/scripts/perf-source.mjs:25–34` accepts declarations inside unused helpers and `test.runIf(false)` as enabled. It does not use runner include patterns. Use runner collection metadata, or reject unsupported declarations. Test dead helpers, conditional skips and excluded files. P3 — F6: `scripts/perf_guards/runner.py:56` maps Settings debt to #701, which owns Notes. Settings shared-contract adoption is #674. Use an explicit table by surface and rule. Review initial owner links and test this mapping. Other P2 findings were added to their existing owners: - #794, F2: `rules.py:176–180` checks awaits only in named arrow-key source; the Svelte exporter checks AwaitBlock. Ordinary open/cache-return animation waits are unchecked. Add call resolution and blocked-animation fixtures. - #795, F4: `rules.py:171–172` flags all fetch/apiFetch calls with no registry acceptance path for a valid adapter, mutation or download. `core.py:180` validates key strings without binding them to the adapter. Add typed bindings and complete positive and negative fixtures. - #793, F5: `rules.py:86–88,114–115` always flags external calls. Projection tests cannot resolve them. The positive SQL fixture at `test_rules.py:153` checks only one absent rule. Add reviewed capability bindings and prove a complete projected read passes with no ledger. Ledger verdict: 18,028 source waivers plus one bundle waiver, all expiring 2026-11-16. Changed, duplicate, expired and unused entries fail. They are not permanent blanket waivers. There is no staged shrink requirement or frozen total across ledger edits. All use one generic synthetic pinning test, not 18,029 product contracts. 12,703 entries are unresolved calls; F4/F5 prevent valid adoption from clearing some debt. Detection verdict: new EachBlocks and list endpoints receive useful coverage and cap checks. New queries usually fail through changed debt or unresolved calls, but no query-count rule proves N+1 behavior. Ordinary animation waits and unused revision-key declarations remain gaps. Four saved web failures are all `Error: Test timed out in 5000ms.` Guard fixtures run before Vitest and do not enter those components. The integrations branch supplied Connected Accounts changes. No direct causal evidence ties these timeouts to guards. Host load remains a possible cause, not a diagnosis. Saved output, verbatim: ``` Test Files 3 failed | 151 passed (154) Tests 4 failed | 1069 passed (1073) Start at 16:11:41 Duration 933.08s (transform 37%, environment 28%, import 18%, tests 13%, setup 4%) ``` No builds or tests ran in this LIGHT review. Prior saved gate output, verbatim: ``` perf-lint: PASS; 0 violations; 18028 scoped exceptions ``` Ran `git fetch origin && git merge origin/dev` once. Output: `Already up to date.` Read both review documents again and ran `git diff --check` with no output. Known gaps: static findings have not been reproduced by execution. The six fixes remain open. No performance or visual claim. No new issue or duplicate owner was created; existing owners are #791, #793, #794 and #795. Decisions: DESIGN §58 in this tree is Agent discovery (#630), so use the explicit #663/#791–#795 requirements and the owner loading rule. Use existing owners. Leave timeout assertions and limits unchanged. For the merge round, after fixes: `scripts/perf-lint test` and `scripts/perf-lint --check --report artifacts/perf-coverage.json` must prove exact no-waiver enforcement, real discovery and valid adapter/projection acceptance. `cargo test -p calternal-server --test perf_guards` must prove the shared hook. Run `cd apps/web && bunx vitest run src/lib/tray.svelte.test.ts src/lib/components/KeyboardShortcutsCard.svelte.test.ts src/routes/settings/connected-accounts/ConnectedAccountsSection.svelte.test.ts --maxWorkers=2` to diagnose the four recorded timeouts without changed expectations.
Author
Owner

Review fixes: perf guards (job perfguards-impl)

Branch job/perfguards-impl, head 42a5c6a64fb8ea9cd561cf3cd4cdebf021d75179. Not pushed.

Fixes for each finding

  • F1 (P1), session clearing waivers. An exception cannot waive a forbidden category under any prefix: access.*, session.*, a11y.*, and any rule whose kind starts with access, session or a11y (for example contract.session-clear). The 317 contract.session-clear entries are removed from the ledger. Seeding skips forbidden rules. A surface now proves session clearing in one of two ways:

    1. Paired surface tests.
    2. retention: "shared-user-storage". This is valid only when the shared userStorage session-end tests are enabled. A surface file that keeps module state also needs its own session-clear test (session_clear.module_tests).

    The guard found two real gaps. PhotoGrid.svelte and SearchPreview.svelte kept User data in module caches. These caches were not cleared on calternal:mode-caches-clear. Both now clear on that event and on auth change. Each has a test that fails without the fix.

  • Ratchet (owner brief). contracts/perf/ratchet.json holds the exception count for each rule. The gate fails in four cases:

    • a count grows,
    • a count shrinks but its ceiling is not lowered,
    • a ceiling or the total is above the ratchet at the merge base with origin/dev (PERF_RATCHET_BASE overrides the base),
    • a forbidden rule has a ceiling above 0.

    Every exception still needs an owner issue and an expiry. The expiry must now be at most 60 days away.

  • F2 (P2), paint barriers. New rule paint.await-motion. It fails when a function assigns a value only after it awaits an animation (.finished, getAnimations, waitFor*), a transition or animation end, or a timer of motion length. A .then continuation or an inline transitionend/animationend listener also counts. An assignment made before the wait passes, so cleanup that resets a value set before the wait also passes.

  • F3 (P2), test discovery. A test counts as enabled only when all of these are true:

    • a Vitest config collects the file (include/exclude patterns are parsed from apps/web/vite.config.ts and packages/editor/vitest.config.ts; packages/ui has no runner),
    • it/test/describe are imported from vitest,
    • the call is a top-level statement or a direct statement of a reachable describe.

    skip, todo, runIf, skipIf, only, each, for and fails disable a test. Declarations in helpers, branches or loops are disabled.

  • F4 (P2), cache adapters. A raw fetch passes when it is inside a registered fetch_boundaries entry with enabled paired tests. The classes are revision-cache-adapter, mutation, raw-download and stream. A mutation must give a literal non-GET method. A raw download must read the raw body. A revision-cache adapter must send If-None-Match and must read every key that its operations declare. cache_variant.adapter must name a registered adapter.

  • F5 (P2), unresolved calls. call_bindings resolve an external call by its exact path and the chain of methods called on its result (for example sqlx::query(..).bind(..).fetch_one(..)), or by its macro name. A binding counts only when its tests are enabled. A method on an unknown receiver, an unlisted method or an unknown macro stays unresolved. A binding to an IO capability is reported as a finding.

  • F6 (P3), owners. Owners now come from an explicit table of area and rule (adoption issue, rule 1, rule 2, rule 8). Areas match whole path words. All ledger owners were reassigned. Settings snapshot debt now goes to #674.

  • N+1. New rule sql.query-in-loop. It fails when a query, or a helper that runs a query, is called in a loop or in an iterator closure on a list or interactive-read path.

Fixtures (new violations fail)

  • New N+1 query: direct, and through a helper in .map.
  • New unbounded list: coverage.surfaces, render.bound and sql.limit.
  • Render that awaits an animation: animation.finished, .then, listener and motion timer. The valid order (assign first) passes.
  • Cache without a revision key: cache.variant-binding. A complete adapter passes with no debt.

Also tested: a complete projection read passes with no ledger; the exact emitted contract.session-clear finding cannot be waived; the ratchet catches growth, a shrink without a lower ceiling, growth above the base and debt moved between rules; Settings debt goes to #674; dead-helper, runIf(false) and outside-include tests are disabled.

Ledger changes

  • 18,029 → 17,803 entries.
  • Removed: the 317 forbidden waivers.
  • Re-pinned: 61 entries in PhotoGrid and SearchPreview. These have the same rule and scope, but a new hash because the files changed. The counts did not change.
  • Added: 91 sql.query-in-loop entries. The new rule found these in existing reads, for example per-item index::identity_at_path and known_hash in files/shares.rs, and per-item INSERT/UPDATE in notes/store.rs and tasks_store.rs. They are pinned as debt that expires on 2026-11-16 and start the ratchet. No real code hit paint.await-motion. No real call_bindings or fetch_boundaries are registered yet. Adoption work (#793, #795) adds them with real tests, and then the ledger can shrink.

Four saved web test timeouts

On origin/dev c4faf184d, tray.svelte.test.ts and KeyboardShortcutsCard.svelte.test.ts pass. ConnectedAccountsSection.svelte.test.ts does not exist there: it comes from the merged integrations work. On this branch, all three files pass when run alone (load average about 38). The saved timeouts were from host load during the full suite, not from the guards.

Gates (verbatim)

  • scripts/perf-lint test: Ran 128 tests in 0.052s / OK
  • bun test scripts/perf-source.test.mjs: 7 pass / 0 fail
  • bun run check: perf-lint: PASS; 0 violations; 17802 scoped exceptions … COMPLETED 1995 FILES 0 ERRORS 0 WARNINGS 0 FILES_WITH_PROBLEMS
  • Focused Vitest (tray, KeyboardShortcutsCard, ConnectedAccountsSection, PhotoGrid, SearchPreview): Test Files 5 passed (5) / Tests 38 passed (38)
  • End-to-end with the real runner:
    • one exception over a ceiling: perf-lint: INVALID: exception ratchet: sql.offset: 4 exceptions exceed the ratchet ceiling 3; fix the new violation instead
    • a session-clear entry added back: contract.session-clear: forbidden rule has 1 exceptions and ceiling 0; it must be 0

Not run and known gaps

  • perf-lint bundle was not run because it needs a production build. Its ledger entry only got a new owner.
  • The full Vitest suite and the cargo perf_guards hook were not run. The hook runs the same perf-lint --check.
  • In CI without origin/dev, the base ratchet comparison is skipped with a printed note. The equality check against the committed ratchet still runs.
  • Module caches in plain .ts stores are outside surface files. They are not checked for session clearing beyond the shared path.
  • DESIGN §59 (instant interactions) is on the merge-round branch and not yet in this branch's DESIGN.md. Comments cite §59.
## Review fixes: perf guards (job perfguards-impl) Branch `job/perfguards-impl`, head `42a5c6a64fb8ea9cd561cf3cd4cdebf021d75179`. Not pushed. ### Fixes for each finding - **F1 (P1), session clearing waivers.** An exception cannot waive a forbidden category under any prefix: `access.*`, `session.*`, `a11y.*`, and any rule whose kind starts with access, session or a11y (for example `contract.session-clear`). The 317 `contract.session-clear` entries are removed from the ledger. Seeding skips forbidden rules. A surface now proves session clearing in one of two ways: 1. Paired surface tests. 2. `retention: "shared-user-storage"`. This is valid only when the shared userStorage session-end tests are enabled. A surface file that keeps module state also needs its own `session-clear` test (`session_clear.module_tests`). The guard found two real gaps. `PhotoGrid.svelte` and `SearchPreview.svelte` kept User data in module caches. These caches were not cleared on `calternal:mode-caches-clear`. Both now clear on that event and on auth change. Each has a test that fails without the fix. - **Ratchet (owner brief).** `contracts/perf/ratchet.json` holds the exception count for each rule. The gate fails in four cases: - a count grows, - a count shrinks but its ceiling is not lowered, - a ceiling or the total is above the ratchet at the merge base with `origin/dev` (`PERF_RATCHET_BASE` overrides the base), - a forbidden rule has a ceiling above 0. Every exception still needs an owner issue and an expiry. The expiry must now be at most 60 days away. - **F2 (P2), paint barriers.** New rule `paint.await-motion`. It fails when a function assigns a value only after it awaits an animation (`.finished`, `getAnimations`, `waitFor*`), a transition or animation end, or a timer of motion length. A `.then` continuation or an inline `transitionend`/`animationend` listener also counts. An assignment made before the wait passes, so cleanup that resets a value set before the wait also passes. - **F3 (P2), test discovery.** A test counts as enabled only when all of these are true: - a Vitest config collects the file (include/exclude patterns are parsed from `apps/web/vite.config.ts` and `packages/editor/vitest.config.ts`; `packages/ui` has no runner), - `it`/`test`/`describe` are imported from `vitest`, - the call is a top-level statement or a direct statement of a reachable `describe`. `skip`, `todo`, `runIf`, `skipIf`, `only`, `each`, `for` and `fails` disable a test. Declarations in helpers, branches or loops are disabled. - **F4 (P2), cache adapters.** A raw fetch passes when it is inside a registered `fetch_boundaries` entry with enabled paired tests. The classes are revision-cache-adapter, mutation, raw-download and stream. A mutation must give a literal non-GET method. A raw download must read the raw body. A revision-cache adapter must send `If-None-Match` and must read every key that its operations declare. `cache_variant.adapter` must name a registered adapter. - **F5 (P2), unresolved calls.** `call_bindings` resolve an external call by its exact path and the chain of methods called on its result (for example `sqlx::query(..).bind(..).fetch_one(..)`), or by its macro name. A binding counts only when its tests are enabled. A method on an unknown receiver, an unlisted method or an unknown macro stays unresolved. A binding to an IO capability is reported as a finding. - **F6 (P3), owners.** Owners now come from an explicit table of area and rule (adoption issue, rule 1, rule 2, rule 8). Areas match whole path words. All ledger owners were reassigned. Settings snapshot debt now goes to #674. - **N+1.** New rule `sql.query-in-loop`. It fails when a query, or a helper that runs a query, is called in a loop or in an iterator closure on a list or interactive-read path. ### Fixtures (new violations fail) - New N+1 query: direct, and through a helper in `.map`. - New unbounded list: `coverage.surfaces`, `render.bound` and `sql.limit`. - Render that awaits an animation: `animation.finished`, `.then`, listener and motion timer. The valid order (assign first) passes. - Cache without a revision key: `cache.variant-binding`. A complete adapter passes with no debt. Also tested: a complete projection read passes with no ledger; the exact emitted `contract.session-clear` finding cannot be waived; the ratchet catches growth, a shrink without a lower ceiling, growth above the base and debt moved between rules; Settings debt goes to #674; dead-helper, `runIf(false)` and outside-include tests are disabled. ### Ledger changes - 18,029 → 17,803 entries. - Removed: the 317 forbidden waivers. - Re-pinned: 61 entries in PhotoGrid and SearchPreview. These have the same rule and scope, but a new hash because the files changed. The counts did not change. - Added: 91 `sql.query-in-loop` entries. The new rule found these in existing reads, for example per-item `index::identity_at_path` and `known_hash` in `files/shares.rs`, and per-item INSERT/UPDATE in `notes/store.rs` and `tasks_store.rs`. They are pinned as debt that expires on 2026-11-16 and start the ratchet. No real code hit `paint.await-motion`. No real `call_bindings` or `fetch_boundaries` are registered yet. Adoption work (#793, #795) adds them with real tests, and then the ledger can shrink. ### Four saved web test timeouts On `origin/dev` c4faf184d, `tray.svelte.test.ts` and `KeyboardShortcutsCard.svelte.test.ts` pass. `ConnectedAccountsSection.svelte.test.ts` does not exist there: it comes from the merged integrations work. On this branch, all three files pass when run alone (load average about 38). The saved timeouts were from host load during the full suite, not from the guards. ### Gates (verbatim) - `scripts/perf-lint test`: `Ran 128 tests in 0.052s` / `OK` - `bun test scripts/perf-source.test.mjs`: `7 pass` / `0 fail` - `bun run check`: `perf-lint: PASS; 0 violations; 17802 scoped exceptions` … `COMPLETED 1995 FILES 0 ERRORS 0 WARNINGS 0 FILES_WITH_PROBLEMS` - Focused Vitest (tray, KeyboardShortcutsCard, ConnectedAccountsSection, PhotoGrid, SearchPreview): `Test Files 5 passed (5)` / `Tests 38 passed (38)` - End-to-end with the real runner: - one exception over a ceiling: `perf-lint: INVALID: exception ratchet: sql.offset: 4 exceptions exceed the ratchet ceiling 3; fix the new violation instead` - a session-clear entry added back: `contract.session-clear: forbidden rule has 1 exceptions and ceiling 0; it must be 0` ### Not run and known gaps - `perf-lint bundle` was not run because it needs a production build. Its ledger entry only got a new owner. - The full Vitest suite and the cargo `perf_guards` hook were not run. The hook runs the same `perf-lint --check`. - In CI without `origin/dev`, the base ratchet comparison is skipped with a printed note. The equality check against the committed ratchet still runs. - Module caches in plain `.ts` stores are outside surface files. They are not checked for session clearing beyond the shared path. - DESIGN §59 (instant interactions) is on the merge-round branch and not yet in this branch's DESIGN.md. Comments cite §59.
Author
Owner

Review follow-up for #791 on job/perfguards-impl. Starting head: 42a5c6a64fb8ea9cd561cf3cd4cdebf021d75179. Read the independent review and its issue comment. Existing review fixes are present; verify them and add regressions for any remaining gaps. Fetch and merge of origin/dev found one DESIGN.md append conflict; keep both Agent discovery and Canvas/history sections. No new dependency or product feature is planned.

Review follow-up for #791 on `job/perfguards-impl`. Starting head: `42a5c6a64fb8ea9cd561cf3cd4cdebf021d75179`. Read the independent review and its issue comment. Existing review fixes are present; verify them and add regressions for any remaining gaps. Fetch and merge of origin/dev found one DESIGN.md append conflict; keep both Agent discovery and Canvas/history sections. No new dependency or product feature is planned.
Author
Owner

Follow-up findings verified with failing regression fixtures:

  • F4: unused key declarations, string key names, keys split across separate cache lookups, and an If-None-Match header with a constant value all passed the earlier fix. Four new fixtures failed before the change. Commit 157f394b3 binds all identity references to one parsed cache key and the revision to the fetch header value.
  • F3: an expect as test import and an async describe callback counted as runner evidence. Two new fixtures failed before the change. Commit 352fcc803 uses imported runner identity and rejects async groups and unknown modifiers.
  • F2: local helpers returning animation.finished hid delayed retained-content assignment. Two new fixtures failed before the change. Commit d4a2ebc5b follows local helper chains with reverse call edges. A valid assign-before-wait fixture passes.

The combined fixture suite passes 135 tests. Bun parser fixtures pass 9 tests. No exception or ceiling was added or raised in this follow-up. The web check and final source gate are in progress.

Follow-up findings verified with failing regression fixtures: - F4: unused key declarations, string key names, keys split across separate cache lookups, and an If-None-Match header with a constant value all passed the earlier fix. Four new fixtures failed before the change. Commit `157f394b3` binds all identity references to one parsed cache key and the revision to the fetch header value. - F3: an `expect as test` import and an async describe callback counted as runner evidence. Two new fixtures failed before the change. Commit `352fcc803` uses imported runner identity and rejects async groups and unknown modifiers. - F2: local helpers returning animation.finished hid delayed retained-content assignment. Two new fixtures failed before the change. Commit `d4a2ebc5b` follows local helper chains with reverse call edges. A valid assign-before-wait fixture passes. The combined fixture suite passes 135 tests. Bun parser fixtures pass 9 tests. No exception or ceiling was added or raised in this follow-up. The web check and final source gate are in progress.
Author
Owner

Review follow-up complete on job/perfguards-impl.
Head: d4a2ebc5b9255a96673958c17bc5076eb6ad0f37.
Starting head: 42a5c6a64fb8ea9cd561cf3cd4cdebf021d75179.
Fetched and merged origin/dev first. Kept both sides of the DESIGN.md append conflict. No push or deploy.

Built and verified:

  • F1: verified the existing exact contract.session-clear rejection fixture, removed 317 waivers, shared retention proof and debt ratchet. The ledger contains no forbidden waiver. Owner issue and expiry remain required. This follow-up added no exception and raised no ceiling.
  • F2: added local call-chain resolution for helpers that return or await motion. Two new negative fixtures failed before the change. The assign-before-motion case passes.
  • F3: runner imports now use their imported identity, including aliases. Async describe groups and unknown modifiers cannot prove collection. Two new negative fixtures failed before the change. Existing dead-helper, conditional and outside-include fixtures also pass.
  • F4: all four identity fields must be read in one cache key expression. The fetch header value must read the revision. Four new negative fixtures failed before the change: unused declarations, string key names, fields split across keys and a constant validator. The complete adapter and mutation positive fixtures still pass without debt.
  • F5: verified the existing complete projection positive fixture with no debt, and the disabled-test, unknown receiver, provider and macro negative fixtures.
  • F6: verified the existing Settings owner fixture (#674).

Atomic follow-up commits: 157f394b3, 352fcc803, d4a2ebc5b.
Files changed in those commits: scripts/perf_guards/{rules,test_rules}.py, apps/web/scripts/perf-source{,.test}.mjs, docs/perf/guards.md. The merge also carries origin/dev changes to DESIGN.md, deploy/media-sandbox and tests/adversarial/prepare-media-runtime.sh.

Gate output, verbatim excerpts:

perf-lint: PASS; 0 violations; 17802 scoped exceptions
Ran 135 tests in 0.101s

OK
 9 pass
 0 fail
 17 expect() calls
Ran 9 tests across 1 file. [972.00ms]

bun run check, exit 0:

svelte-check found 0 errors and 0 warnings

Focused PhotoGrid/SearchPreview session-clearing Vitest, exit 0:

 Test Files  2 passed (2)
      Tests  2 passed (2)
   Start at  09:09:22
   Duration  56.62s (transform 79%, environment 12%, import 5%, setup 3%, tests 2%)

cargo fmt --check: exit 0, no output. No Rust code changed in this follow-up, so no crate clippy/test build ran. git diff --check: exit 0, no output. Logs and the coverage report are in worktree artifacts/review-*. Doc comments re-read. Existing test expectations unchanged. Worktree clean.

Cleanup:

     Removed 1 file, 356B total

Web production build output removed.

Known gaps: 17,803 total ledger entries remain (17,802 source, one bundle). Adoption has not yet registered real call_bindings or fetch_boundaries. This is source/contract evidence, not a runtime timing claim. Cache key/header analysis is local expression proof, not TypeScript type inference. Opaque builders need an explicit resolution path. Motion helper resolution covers local parsed identifiers, not arbitrary external dispatch. No product rendering changed in this follow-up; no screenshot or perf measurement was required.

Decisions: use conservative direct cache expressions and synchronous runner collection as supported proof. Resolve motion helper chains with reverse call edges to avoid repeated whole-function scans. Preserve Agent discovery and Canvas/history design sections from both merge sides.

UX gaps closed / left: no user-facing UI change in this follow-up. Existing cache-clearing fixes are covered by the two focused tests.

For the merge round:

  • cargo test -p calternal-server --test perf_guards: prove the server hook runs the corrected common source gate.
  • cd apps/web && bun run test: full suite once on the combined branch.
  • cd apps/web && bunx vitest run src/lib/tray.svelte.test.ts src/lib/components/KeyboardShortcutsCard.svelte.test.ts src/routes/settings/connected-accounts/ConnectedAccountsSection.svelte.test.ts --maxWorkers=2: verify the four previously saved timeouts with unchanged expectations.
  • cd apps/web && bun run build, then scripts/perf-lint bundle --report artifacts/perf-bundles.json: verify production bundle debt and byte limits on the combined build.
Review follow-up complete on `job/perfguards-impl`. Head: `d4a2ebc5b9255a96673958c17bc5076eb6ad0f37`. Starting head: `42a5c6a64fb8ea9cd561cf3cd4cdebf021d75179`. Fetched and merged origin/dev first. Kept both sides of the DESIGN.md append conflict. No push or deploy. Built and verified: - F1: verified the existing exact `contract.session-clear` rejection fixture, removed 317 waivers, shared retention proof and debt ratchet. The ledger contains no forbidden waiver. Owner issue and expiry remain required. This follow-up added no exception and raised no ceiling. - F2: added local call-chain resolution for helpers that return or await motion. Two new negative fixtures failed before the change. The assign-before-motion case passes. - F3: runner imports now use their imported identity, including aliases. Async describe groups and unknown modifiers cannot prove collection. Two new negative fixtures failed before the change. Existing dead-helper, conditional and outside-include fixtures also pass. - F4: all four identity fields must be read in one cache key expression. The fetch header value must read the revision. Four new negative fixtures failed before the change: unused declarations, string key names, fields split across keys and a constant validator. The complete adapter and mutation positive fixtures still pass without debt. - F5: verified the existing complete projection positive fixture with no debt, and the disabled-test, unknown receiver, provider and macro negative fixtures. - F6: verified the existing Settings owner fixture (#674). Atomic follow-up commits: `157f394b3`, `352fcc803`, `d4a2ebc5b`. Files changed in those commits: `scripts/perf_guards/{rules,test_rules}.py`, `apps/web/scripts/perf-source{,.test}.mjs`, `docs/perf/guards.md`. The merge also carries origin/dev changes to DESIGN.md, deploy/media-sandbox and tests/adversarial/prepare-media-runtime.sh. Gate output, verbatim excerpts: ``` perf-lint: PASS; 0 violations; 17802 scoped exceptions ``` ``` Ran 135 tests in 0.101s OK ``` ``` 9 pass 0 fail 17 expect() calls Ran 9 tests across 1 file. [972.00ms] ``` `bun run check`, exit 0: ``` svelte-check found 0 errors and 0 warnings ``` Focused PhotoGrid/SearchPreview session-clearing Vitest, exit 0: ``` Test Files 2 passed (2) Tests 2 passed (2) Start at 09:09:22 Duration 56.62s (transform 79%, environment 12%, import 5%, setup 3%, tests 2%) ``` `cargo fmt --check`: exit 0, no output. No Rust code changed in this follow-up, so no crate clippy/test build ran. `git diff --check`: exit 0, no output. Logs and the coverage report are in worktree `artifacts/review-*`. Doc comments re-read. Existing test expectations unchanged. Worktree clean. Cleanup: ``` Removed 1 file, 356B total ``` Web production build output removed. Known gaps: 17,803 total ledger entries remain (17,802 source, one bundle). Adoption has not yet registered real call_bindings or fetch_boundaries. This is source/contract evidence, not a runtime timing claim. Cache key/header analysis is local expression proof, not TypeScript type inference. Opaque builders need an explicit resolution path. Motion helper resolution covers local parsed identifiers, not arbitrary external dispatch. No product rendering changed in this follow-up; no screenshot or perf measurement was required. Decisions: use conservative direct cache expressions and synchronous runner collection as supported proof. Resolve motion helper chains with reverse call edges to avoid repeated whole-function scans. Preserve Agent discovery and Canvas/history design sections from both merge sides. UX gaps closed / left: no user-facing UI change in this follow-up. Existing cache-clearing fixes are covered by the two focused tests. For the merge round: - `cargo test -p calternal-server --test perf_guards`: prove the server hook runs the corrected common source gate. - `cd apps/web && bun run test`: full suite once on the combined branch. - `cd apps/web && bunx vitest run src/lib/tray.svelte.test.ts src/lib/components/KeyboardShortcutsCard.svelte.test.ts src/routes/settings/connected-accounts/ConnectedAccountsSection.svelte.test.ts --maxWorkers=2`: verify the four previously saved timeouts with unchanged expectations. - `cd apps/web && bun run build`, then `scripts/perf-lint bundle --report artifacts/perf-bundles.json`: verify production bundle debt and byte limits on the combined build.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#791
No description provided.