P1: Bind UI intents and successful fixture evidence to the parity gate #834

Open
opened 2026-10-02 13:23:02 +00:00 by kayg · 1 comment
Owner

Research follow-up under #484. Source snapshot: c4a61e8cf090170f35b1bed3350d9de20c83ecd5.

Current route coverage is broad: 333 operations, 315 generated tools, zero non-exempt adapter gaps. That does not prove complete workflows.

Evidence:

  • scripts/parity_matrix.py prints 145 static menu rows with unresolved route mappings and does not fail on those rows.
  • The gate scans literal apiFetch calls and static patterns. The Notes client uses a separate request wrapper; dynamic provider commands, editor transactions and template-literal calls require semantic review.
  • Historical documentation records 170/288 successful tools per surface. Today's eligible set is 315; no current full successful-operation artifact was produced by this research job.
  • apps/web/e2e/webmcp.mjs --require-full-smoke already rejects missing successful fixtures. Preserve it.

Acceptance:

  1. Continue #581: give each data-changing UI action one registry ID and an explicit reason for genuine presentation/sign-in exceptions.
  2. Cover dynamic menus, commands, Settings behavior and editor mutations. Reject an unbound non-exempt intent.
  3. Require real successful calls for all eligible tools with real identities/revisions; use a throwaway Home for writes and controlled provider fixtures.
  4. Extend existing cross-User, admin, read-only, Home-prefix, revoke and disabled-surface checks; keep denial evidence separate from successful smoke coverage.
  5. Record full denominator and source SHA. A discovered tool or rejected call is not a successful fixture.

Matrix covers 58 capability groups, 333 operations, 145 menus and 126 shortcuts. This issue is the semantic/proof slice under #484, not a replacement for the current route gate.

Full evidence and decisions: docs/research/agent-surfaces.md on branch job/research-surfaces. No runtime change was made by the research job.

Research follow-up under #484. Source snapshot: `c4a61e8cf090170f35b1bed3350d9de20c83ecd5`. Current route coverage is broad: 333 operations, 315 generated tools, zero non-exempt adapter gaps. That does not prove complete workflows. Evidence: - `scripts/parity_matrix.py` prints 145 static menu rows with unresolved route mappings and does not fail on those rows. - The gate scans literal `apiFetch` calls and static patterns. The Notes client uses a separate request wrapper; dynamic provider commands, editor transactions and template-literal calls require semantic review. - Historical documentation records 170/288 successful tools per surface. Today's eligible set is 315; no current full successful-operation artifact was produced by this research job. - `apps/web/e2e/webmcp.mjs --require-full-smoke` already rejects missing successful fixtures. Preserve it. Acceptance: 1. Continue #581: give each data-changing UI action one registry ID and an explicit reason for genuine presentation/sign-in exceptions. 2. Cover dynamic menus, commands, Settings behavior and editor mutations. Reject an unbound non-exempt intent. 3. Require real successful calls for all eligible tools with real identities/revisions; use a throwaway Home for writes and controlled provider fixtures. 4. Extend existing cross-User, admin, read-only, Home-prefix, revoke and disabled-surface checks; keep denial evidence separate from successful smoke coverage. 5. Record full denominator and source SHA. A discovered tool or rejected call is not a successful fixture. Matrix covers 58 capability groups, 333 operations, 145 menus and 126 shortcuts. This issue is the semantic/proof slice under #484, not a replacement for the current route gate. Full evidence and decisions: `docs/research/agent-surfaces.md` on branch `job/research-surfaces`. No runtime change was made by the research job.
Author
Owner

P2: Parity evidence does not check the source at completion or at the gate

  • Evidence: apps/web/e2e/webmcp.mjs:23 records HEAD, cleanliness and registry
    digest once before the fixtures. writeEvidence at line 32 reuses that
    object, including for the final complete: true output at line 548.
  • Gate: scripts/parity_matrix.py:192 reads current HEAD and registry digest,
    but does not check current source cleanliness. Line 151 trusts the stored
    source_dirty: false field.
  • Effect: a tracked source edit after the initial snapshot can leave the same
    HEAD and registry digest. A completed artifact can then describe clean source
    although the fixture run or gate used changed source. An untracked source file
    added after the initial snapshot has the same problem.
  • Rule: #834 acceptance 5 and the documented complete clean-source parity gate.
  • Fix: compare HEAD, registry digest and source cleanliness at completion with
    the initial snapshot. Refuse complete evidence if they differ. Check current
    source cleanliness when the release gate consumes evidence. Reuse one helper
    for the identity checks and retain the current untracked-file check.
  • Test idea: edit a tracked source file or add an untracked source file after
    the initial snapshot. Both completion and release validation must reject it.
  • Existing issue: #834 owns fixture evidence. No duplicate issue is required.
## P2: Parity evidence does not check the source at completion or at the gate - Evidence: `apps/web/e2e/webmcp.mjs:23` records HEAD, cleanliness and registry digest once before the fixtures. `writeEvidence` at line 32 reuses that object, including for the final `complete: true` output at line 548. - Gate: `scripts/parity_matrix.py:192` reads current HEAD and registry digest, but does not check current source cleanliness. Line 151 trusts the stored `source_dirty: false` field. - Effect: a tracked source edit after the initial snapshot can leave the same HEAD and registry digest. A completed artifact can then describe clean source although the fixture run or gate used changed source. An untracked source file added after the initial snapshot has the same problem. - Rule: #834 acceptance 5 and the documented complete clean-source parity gate. - Fix: compare HEAD, registry digest and source cleanliness at completion with the initial snapshot. Refuse complete evidence if they differ. Check current source cleanliness when the release gate consumes evidence. Reuse one helper for the identity checks and retain the current untracked-file check. - Test idea: edit a tracked source file or add an untracked source file after the initial snapshot. Both completion and release validation must reject it. - Existing issue: #834 owns fixture evidence. No duplicate issue is required.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#834
No description provided.