VOICE: revive and merge local Parakeet transcription (job/voice, #373) + backfill-transcribe existing voice memos #619

Open
opened 2026-10-01 10:13:05 +00:00 by kayg · 56 comments
Owner

Owner question (2026-10-01): "When is Parakeet coming? And auto-transcribing existing voice notes?"

State: local transcription (#373, Parakeet v2 → S1-mini, background model download, inline mic morph) was built on job/voice but never merged (40 commits ahead, last touched ~37 h ago). It also implements the owner's #304 decision: a recording made on a Log entry becomes a linked note (a child bullet whose title is the link text) that holds the title + recording + transcription; on a task → the task note; on a note → that note.
Job:

  1. Merge origin/dev into job/voice and resolve conflicts (the composer has changed a lot since: #617 Voice Pill and #618 Voice Memos are queued; coordinate: keep the recorder hooks generic so the Voice Pill can wrap them).
  2. Verify end to end on a production build: record on a Log entry → linked note created → transcription arrives automatically (English, local only, on by default) → the transcript is searchable.
  3. Backfill: a background job transcribes existing voice memos (everything under Documents/Voice Memos and Documents/Voice memos plus audio attachments referenced by entries), newest first, resumable, low priority, with progress in Settings → Maintenance (Jobs). It never re-transcribes a file whose transcript exists (hash-keyed).
  4. Performance: report the real-time factor on the build host and on the perf VM for a 1 min and a 30 min recording (chunk long audio, as decided), plus peak RSS.
    Gates per crate + web. Screenshots of the record → transcript flow. Time limit 5 h.
## Owner question (2026-10-01): "When is Parakeet coming? And auto-transcribing existing voice notes?" **State:** local transcription (#373, Parakeet v2 → S1-mini, background model download, inline mic morph) was **built on `job/voice` but never merged** (40 commits ahead, last touched ~37 h ago). It also implements the owner's #304 decision: a recording made on a Log entry becomes a **linked note** (a child bullet whose title is the link text) that holds the title + recording + transcription; on a task → the task note; on a note → that note. **Job:** 1. Merge origin/dev into `job/voice` and resolve conflicts (the composer has changed a lot since: #617 Voice Pill and #618 Voice Memos are queued; coordinate: keep the recorder hooks generic so the Voice Pill can wrap them). 2. Verify end to end on a production build: record on a Log entry → linked note created → transcription arrives automatically (English, local only, on by default) → the transcript is searchable. 3. **Backfill:** a background job transcribes **existing** voice memos (everything under `Documents/Voice Memos` and `Documents/Voice memos` plus audio attachments referenced by entries), newest first, resumable, low priority, with progress in Settings → Maintenance (Jobs). It never re-transcribes a file whose transcript exists (hash-keyed). 4. Performance: report the real-time factor on the build host and on the perf VM for a 1 min and a 30 min recording (chunk long audio, as decided), plus peak RSS. Gates per crate + web. Screenshots of the record → transcript flow. Time limit 5 h.
Author
Owner

Started voice-619 on branch job/voice-619 at base SHA 9a6c38c3b7c8bbbdad9919cd2419f773d9c34b6e (origin/dev was 687ff703136e71e89f8dfba139e93cd0788b25c1, seven commits ahead). I am reviewing the binding voice and Jobs decisions, then merging the existing job/voice work into this branch and adapting it to the current Composer. I will keep recorder hooks generic for Voice Pill integration and add resumable, hash-keyed backfill for existing voice memos and Journal/task audio attachments.

Started voice-619 on branch `job/voice-619` at base SHA `9a6c38c3b7c8bbbdad9919cd2419f773d9c34b6e` (origin/dev was `687ff703136e71e89f8dfba139e93cd0788b25c1`, seven commits ahead). I am reviewing the binding voice and Jobs decisions, then merging the existing `job/voice` work into this branch and adapting it to the current Composer. I will keep recorder hooks generic for Voice Pill integration and add resumable, hash-keyed backfill for existing voice memos and Journal/task audio attachments.
Author
Owner

Finding: bun run check stopped before Svelte checking because apps/web/node_modules does not contain typescript (ERR_MODULE_NOT_FOUND). I am installing from the committed lockfile, then I will rerun the check.

Finding: `bun run check` stopped before Svelte checking because `apps/web/node_modules` does not contain `typescript` (`ERR_MODULE_NOT_FOUND`). I am installing from the committed lockfile, then I will rerun the check.
Author
Owner

Findings during integration:

  • The first Web check could not start because typescript was absent from apps/web/node_modules; bun install --frozen-lockfile restored the lockfile dependencies.
  • The merged VoiceRecordingPill used raw CSS motion/shape values. check-type-tokens and check-motion-tokens identified them; I replaced them with the shared tokens.
  • The old branch and current dev had duplicate appendEntryAttachments implementations and incompatible ItemPreview edit callbacks. The app type check exposed these, so I am retaining the current Journal API and routing Search edits through the shared Composer.
Findings during integration: - The first Web check could not start because `typescript` was absent from `apps/web/node_modules`; `bun install --frozen-lockfile` restored the lockfile dependencies. - The merged VoiceRecordingPill used raw CSS motion/shape values. `check-type-tokens` and `check-motion-tokens` identified them; I replaced them with the shared tokens. - The old branch and current `dev` had duplicate `appendEntryAttachments` implementations and incompatible `ItemPreview` edit callbacks. The app type check exposed these, so I am retaining the current Journal API and routing Search edits through the shared Composer.
Author
Owner

Finding: the first Rust build stopped in llama-cpp-sys-2's bindgen step because this build host has no discoverable libclang.so (Unable to find libclang ... set LIBCLANG_PATH). I am checking the installed toolchain and will continue as soon as the available library path is identified.

Finding: the first Rust build stopped in `llama-cpp-sys-2`'s bindgen step because this build host has no discoverable `libclang.so` (`Unable to find libclang ... set LIBCLANG_PATH`). I am checking the installed toolchain and will continue as soon as the available library path is identified.
Author
Owner

The libclang finding is resolved on this build host: libclang-19-dev was absent; the distro package is installed and I am restarting the same build with its library path. This changed no repository files.

The libclang finding is resolved on this build host: `libclang-19-dev` was absent; the distro package is installed and I am restarting the same build with its library path. This changed no repository files.
Author
Owner

The next compiler attempt reached native model builds and found the same image is also missing cmake; this is an environment dependency, not a Rust diagnostic. I am installing the distro build tool and checking the native build's worker count before retrying so it stays within the shared-host limit.

The next compiler attempt reached native model builds and found the same image is also missing `cmake`; this is an environment dependency, not a Rust diagnostic. I am installing the distro build tool and checking the native build's worker count before retrying so it stays within the shared-host limit.
Author
Owner

Decisions for points not set in DESIGN: files found only in Documents/Voice Memos or Documents/Voice memos become one standalone Note in Documents/Notes, titled from the audio filename and linked to the original audio. A Journal attachment creates or reuses the entry's linked child Note; a Note or Task attachment writes the transcript to that existing destination. The transcript cache is private per User and keyed by the audio BLAKE3 hash. The performance sample will be deterministic English speech synthesized locally so both 1-minute and 30-minute timings can be repeated without shipping fixture audio or sending speech to a service.

Decisions for points not set in DESIGN: files found only in `Documents/Voice Memos` or `Documents/Voice memos` become one standalone Note in `Documents/Notes`, titled from the audio filename and linked to the original audio. A Journal attachment creates or reuses the entry's linked child Note; a Note or Task attachment writes the transcript to that existing destination. The transcript cache is private per User and keyed by the audio BLAKE3 hash. The performance sample will be deterministic English speech synthesized locally so both 1-minute and 30-minute timings can be repeated without shipping fixture audio or sending speech to a service.
Author
Owner

The Notes crate compile surfaced two integration issues from the merge: both controlled and ordinary backfill handlers returned String errors where the shared job queue requires JobError; they now map that error explicitly. The newer User Settings API also adds InsufficientStorage; the Voice settings writer now returns HTTP 507 for quota failures. I am rerunning the crate check before continuing.

The Notes crate compile surfaced two integration issues from the merge: both controlled and ordinary backfill handlers returned `String` errors where the shared job queue requires `JobError`; they now map that error explicitly. The newer User Settings API also adds `InsufficientStorage`; the Voice settings writer now returns HTTP 507 for quota failures. I am rerunning the crate check before continuing.
Author
Owner

Finding: both historical folders can contain byte-identical files. A Note ID derived from the audio hash would collapse those two paths into one Note and leave the second source file without its own link. Standalone backfill Notes now use a stable identity per Home-relative source path; the transcript cache stays keyed by User and audio hash, so the duplicate bytes share recognition while both files keep a Note link. The E2E fixture seeds the same speech file in both folder spellings to cover this case.

Finding: both historical folders can contain byte-identical files. A Note ID derived from the audio hash would collapse those two paths into one Note and leave the second source file without its own link. Standalone backfill Notes now use a stable identity per Home-relative source path; the transcript cache stays keyed by User and audio hash, so the duplicate bytes share recognition while both files keep a Note link. The E2E fixture seeds the same speech file in both folder spellings to cover this case.
Author
Owner

Integration finding: bun run check found the Voice routes were absent from contracts/openapi.json, and the Search preview was passing an attachment callback shape that no longer matched ItemPreview. I am generating the contract from the server and restoring the Search preview's file picker through its current callback while keeping the recorder pill as a generic snippet hook.

Integration finding: `bun run check` found the Voice routes were absent from `contracts/openapi.json`, and the Search preview was passing an attachment callback shape that no longer matched `ItemPreview`. I am generating the contract from the server and restoring the Search preview's file picker through its current callback while keeping the recorder pill as a generic snippet hook.
Author
Owner

Finding and fix: the first server OpenAPI build reported that crates/plugins/files/src/media.rs still contained an unused MediaProcessGroup; the shared process-group ownership now lives in calternal-plugin::media_sandbox. I removed the obsolete duplicate guard without changing the timeout regression test. OpenAPI generation also requires apps/web/build because the server embeds the production SPA, so I am building that first.

Finding and fix: the first server OpenAPI build reported that `crates/plugins/files/src/media.rs` still contained an unused `MediaProcessGroup`; the shared process-group ownership now lives in `calternal-plugin::media_sandbox`. I removed the obsolete duplicate guard without changing the timeout regression test. OpenAPI generation also requires `apps/web/build` because the server embeds the production SPA, so I am building that first.
Author
Owner

Correction to the earlier design note: standalone backfill Note IDs use the Home-relative source path, not the audio hash. This keeps two byte-identical memos in the two historical folder spellings individually reachable. The transcription cache and destination marker remain keyed by User, BLAKE3 audio hash, and destination.

Correction to the earlier design note: standalone backfill Note IDs use the Home-relative source path, not the audio hash. This keeps two byte-identical memos in the two historical folder spellings individually reachable. The transcription cache and destination marker remain keyed by User, BLAKE3 audio hash, and destination.
Author
Owner

Finding and fix: the web suite reported that the merged shortcut registry changed calendar.attach from the contract's A key to Cmd+Shift+A, while the issue branch's existing test expects A. I restored the registered A binding and wired it to the focused Journal attachment action in ItemPreview. I left both existing expectations unchanged; I will rerun the focused and full web tests.

Finding and fix: the web suite reported that the merged shortcut registry changed `calendar.attach` from the contract's `A` key to Cmd+Shift+A, while the issue branch's existing test expects `A`. I restored the registered `A` binding and wired it to the focused Journal attachment action in ItemPreview. I left both existing expectations unchanged; I will rerun the focused and full web tests.
Author
Owner

Benchmark fixture check: FFprobe measured the original 30-minute WebM at 1800.008 seconds, 8 ms above the API's hard 30-minute decode cap. I truncated only the ignored test input to 1799.021 seconds; the benchmark accepts its duration within 1 second and the server cap remains unchanged.

Benchmark fixture check: FFprobe measured the original 30-minute WebM at 1800.008 seconds, 8 ms above the API's hard 30-minute decode cap. I truncated only the ignored test input to 1799.021 seconds; the benchmark accepts its duration within 1 second and the server cap remains unchanged.
Author
Owner

Rust finding: cargo clippy -p calternal-plugin-files --all-targets -- -D warnings reached the Notes dependency and rejected ensure_journal_voice_note for eight arguments. I am grouping the audio path/hash and Journal block identity/title into one documented value type; no lint suppression or behavior change.

Rust finding: `cargo clippy -p calternal-plugin-files --all-targets -- -D warnings` reached the Notes dependency and rejected `ensure_journal_voice_note` for eight arguments. I am grouping the audio path/hash and Journal block identity/title into one documented value type; no lint suppression or behavior change.
Author
Owner

Owner override applied: inherited #527 code still set all motion tokens to 0 ms for data-input="keyboard", and Settings added an instantOpen path. The job instructions withdraw that behavior. I am removing the keyboard-motion suppression, including the Settings/sidebar special case, while preserving reduced-motion handling and the keyboard flag used by focus rings.

Owner override applied: inherited #527 code still set all motion tokens to 0 ms for `data-input="keyboard"`, and Settings added an `instantOpen` path. The job instructions withdraw that behavior. I am removing the keyboard-motion suppression, including the Settings/sidebar special case, while preserving reduced-motion handling and the keyboard flag used by focus rings.
Author
Owner

Found that origin/dev already uses Notes migration 0024 for Journal day read snapshots in crates/plugins/notes/migrations/0024_journal_day_snapshots.sql. Renumbered the Voice transcript migration and registration to 0025 before final gates, preserving both migrations.

Found that origin/dev already uses Notes migration 0024 for Journal day read snapshots in crates/plugins/notes/migrations/0024_journal_day_snapshots.sql. Renumbered the Voice transcript migration and registration to 0025 before final gates, preserving both migrations.
Author
Owner

During the benchmark audit, measure() called request_audio() with four arguments while the helper required five; the extra path parameter was unused. Removed that parameter so the requested timing run can reach transcription.

During the benchmark audit, `measure()` called `request_audio()` with four arguments while the helper required five; the extra `path` parameter was unused. Removed that parameter so the requested timing run can reach transcription.
Author
Owner

The production E2E found two backfill candidates but produced no Notes or transcript cache entries. I reproduced the decoder call directly: local voice-mode media sandbox exits 127 (prlimit: failed to execute ffmpeg: No such file or directory). This host resolves FFmpeg inside the Nix store, outside the sandbox mounts. I am correcting the local runtime setup and will rerun the real-server flow.

The production E2E found two backfill candidates but produced no Notes or transcript cache entries. I reproduced the decoder call directly: local voice-mode media sandbox exits 127 (`prlimit: failed to execute ffmpeg: No such file or directory`). This host resolves FFmpeg inside the Nix store, outside the sandbox mounts. I am correcting the local runtime setup and will rerun the real-server flow.
Author
Owner

Fixed the local media runtime staging and the production runtime image to link FFmpeg’s real BLAS and LAPACK libraries into the read-only decoder tool tree; /etc/alternatives remains outside the sandbox. The direct voice-mode sandbox decode now exits 0 and emits 512,000 PCM bytes from the spoken fixture.

Fixed the local media runtime staging and the production runtime image to link FFmpeg’s real BLAS and LAPACK libraries into the read-only decoder tool tree; `/etc/alternatives` remains outside the sandbox. The direct voice-mode sandbox decode now exits 0 and emits 512,000 PCM bytes from the spoken fixture.
Author
Owner

After fixing the sandbox, backfill created both Notes and stored one transcript for the identical audio hash. The next E2E assertion failed because /api/v1/search returns the indexed Note file by its Home-relative path (plugin_id: files, id: Notes/...md), while the probe expected the Note API UUID. I am changing the probe to search a transcript-only word and assert that indexed path.

After fixing the sandbox, backfill created both Notes and stored one transcript for the identical audio hash. The next E2E assertion failed because `/api/v1/search` returns the indexed Note file by its Home-relative path (`plugin_id: files`, `id: Notes/...md`), while the probe expected the Note API UUID. I am changing the probe to search a transcript-only word and assert that indexed path.
Author
Owner

The revised Search probe passed both backfill cases and the flow reached the Journal Log preview. It then timed out because the merged Calendar now names that accessible dialog Journal; the E2E still looked for Log entry. I am aligning the locator with the current Calendar label.

The revised Search probe passed both backfill cases and the flow reached the Journal Log preview. It then timed out because the merged Calendar now names that accessible dialog `Journal`; the E2E still looked for `Log entry`. I am aligning the locator with the current Calendar label.
Author
Owner

Correction after checking the merged Calendar keyboard contract: the Log row’s Enter action opens its Composer for editing; Space selects the Log and opens its preview. The Journal label was correct. I am updating the E2E activation to Space so it records the intended preview action.

Correction after checking the merged Calendar keyboard contract: the Log row’s Enter action opens its Composer for editing; Space selects the Log and opens its preview. The `Journal` label was correct. I am updating the E2E activation to Space so it records the intended preview action.
Author
Owner

The cached-model E2E reached the recorder and captured calendar-recorder-390-light.png. Its next action failed because the broad Discard locator also matched a hidden Composer Discard draft button. I am making that action exact so the screenshot loop can continue.

The cached-model E2E reached the recorder and captured `calendar-recorder-390-light.png`. Its next action failed because the broad `Discard` locator also matched a hidden Composer `Discard draft` button. I am making that action exact so the screenshot loop can continue.
Author
Owner

The recorder viewport checks now complete, and the live recording reaches Keep and Save changes. After the Composer closes, the Calendar preview does not expose its expected Open note link within 30 seconds. I will preserve the throwaway server data on the next run to inspect whether the linked Note or the preview refresh is missing.

The recorder viewport checks now complete, and the live recording reaches Keep and Save changes. After the Composer closes, the Calendar preview does not expose its expected `Open note` link within 30 seconds. I will preserve the throwaway server data on the next run to inspect whether the linked Note or the preview refresh is missing.
Author
Owner

Preserved E2E data shows the product did create the linked child Note: the Journal attachment projection contains both the WebM and Notes/20261002-voice-linked-memo-proof-ce78bd1f.md, and the Note file exists with the recording link and pending transcript marker. The probe held a locator for the pre-edit preview, so I will reopen the saved Log and assert the fresh preview link.

Preserved E2E data shows the product did create the linked child Note: the Journal attachment projection contains both the WebM and `Notes/20261002-voice-linked-memo-proof-ce78bd1f.md`, and the Note file exists with the recording link and pending transcript marker. The probe held a locator for the pre-edit preview, so I will reopen the saved Log and assert the fresh preview link.
Author
Owner

The re-opened Journal preview now exposes the linked Note. Its transcription Job completed, and the transcript-only Search query returned the indexed Note path. The Note shows both cleaned and raw transcript text, so the screen-capture wait matched two elements; I am selecting the first visible transcript block.

The re-opened Journal preview now exposes the linked Note. Its transcription Job completed, and the transcript-only Search query returned the indexed Note path. The Note shows both cleaned and raw transcript text, so the screen-capture wait matched two elements; I am selecting the first visible transcript block.
Author
Owner

Production E2E finding (2026-10-02): the 9-second recording is present in the Calendar Log attachments and in its linked Note, but the local notes.voice.transcribe job is absent from .system/index.sqlite after the edit save. The transcript wait therefore cannot pass. The route-level model download and old-file backfill jobs completed successfully. I am adding a production-browser assertion on the transcription POST response so the next run reports the queue failure at the save boundary.

Production E2E finding (2026-10-02): the 9-second recording is present in the Calendar Log attachments and in its linked Note, but the local `notes.voice.transcribe` job is absent from `.system/index.sqlite` after the edit save. The transcript wait therefore cannot pass. The route-level model download and old-file backfill jobs completed successfully. I am adding a production-browser assertion on the transcription POST response so the next run reports the queue failure at the save boundary.
Author
Owner

Performance finding (2026-10-02): the production UI matrix passes, and the live Journal transcript job completes. The local benchmark accepted its 60-second audio, but that queued job remained leased for more than 20 minutes without an error. The build host load average was 25.43/26.28/24.00 on an 8-core host. This is a slow-path result under heavy shared load; the 30-minute local run has not started yet.

Performance finding (2026-10-02): the production UI matrix passes, and the live Journal transcript job completes. The local benchmark accepted its 60-second audio, but that queued job remained leased for more than 20 minutes without an error. The build host load average was 25.43/26.28/24.00 on an 8-core host. This is a slow-path result under heavy shared load; the 30-minute local run has not started yet.
Author
Owner

#619 implementation report

Head: 4011a4fad on job/voice-619.

Built

  • Local English transcription with Parakeet v2 and S1-mini, automatic downloads enabled by default, and hash-keyed result reuse.
  • Journal recordings create a linked Note with the audio and transcript. Existing audio is backfilled from both Documents/Voice Memos spellings and audio attachments; the newest items run first in a low-priority resumable job shown in Maintenance.
  • Production E2E verified the linked Note, completed live transcript, searchable transcript text, both legacy-folder backfills, and the six width/theme combinations. Thirty screenshots are attached below.

Gate output

  • cargo fmt --check: no output; exit 0.
  • cargo clippy -p calternal-fs --all-targets -- -D warnings: Finished dev profile [unoptimized + debuginfo] target(s) in 29.48s.
  • cargo test -p calternal-fs: test result: ok. 53 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; storage: 42 passed; 0 failed.
  • cargo clippy -p calternal-plugin --all-targets -- -D warnings: Finished dev profile [unoptimized + debuginfo] target(s) in 1m 17s.
  • cargo test -p calternal-plugin: 25 passed; 0 failed; doc tests: 0 passed; 0 failed.
  • cargo clippy -p calternal-plugin-notes --all-targets -- -D warnings: Finished dev profile [unoptimized + debuginfo] target(s) in 55.74s.
  • Earlier post-merge gates: Notes tests 179 passed; 0 failed; 1 ignored; Files clippy passed; Files tests 146 passed; 0 failed; 1 ignored.
  • cargo clippy -p calternal-server --all-targets -- -D warnings: Finished dev profile [unoptimized + debuginfo] target(s) in 5m 07s.
  • cargo test -p calternal-server was still compiling at the five-hour job limit.
  • bun run check: svelte-check found 0 errors and 0 warnings.
  • bun run test: Test Files 153 passed (153); Tests 1056 passed (1056).
  • Production bun run build passed earlier: ✓ built in 37.06s; Wrote site to "build"; ✔ done.

Performance and gaps

The build host load average was 25.43, 26.28, 24.00 on the shared 8-core host. The benchmark accepted a 60-second recording, but its job remained leased for over 20 minutes without an error; I stopped the run at the time limit. I do not have a completed local RTF, the 30-minute local result, or perf VM RTF/RSS numbers. The existing release server on the perf VM did not contain the Voice route. The adversarial round and cargo test -p calternal-server remain incomplete. I did not run cargo clean before the time limit.

Decisions

  • A Journal Log recording belongs in its linked child Note, following owner decision #304.
  • Long audio uses five-minute recognition chunks with a half-second overlap. Backfill de-duplicates by audio content hash per User and keeps both historical folder spellings.
  • Keyboard-triggered UI motion follows the 2026-10-01 owner override and reduced-motion preferences.

Production screenshots

#619 implementation report **Head:** `4011a4fad` on `job/voice-619`. ## Built - Local English transcription with Parakeet v2 and S1-mini, automatic downloads enabled by default, and hash-keyed result reuse. - Journal recordings create a linked Note with the audio and transcript. Existing audio is backfilled from both `Documents/Voice Memos` spellings and audio attachments; the newest items run first in a low-priority resumable job shown in Maintenance. - Production E2E verified the linked Note, completed live transcript, searchable transcript text, both legacy-folder backfills, and the six width/theme combinations. Thirty screenshots are attached below. ## Gate output - `cargo fmt --check`: no output; exit 0. - `cargo clippy -p calternal-fs --all-targets -- -D warnings`: `Finished `dev` profile [unoptimized + debuginfo] target(s) in 29.48s`. - `cargo test -p calternal-fs`: `test result: ok. 53 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out`; storage: `42 passed; 0 failed`. - `cargo clippy -p calternal-plugin --all-targets -- -D warnings`: `Finished `dev` profile [unoptimized + debuginfo] target(s) in 1m 17s`. - `cargo test -p calternal-plugin`: `25 passed; 0 failed`; doc tests: `0 passed; 0 failed`. - `cargo clippy -p calternal-plugin-notes --all-targets -- -D warnings`: `Finished `dev` profile [unoptimized + debuginfo] target(s) in 55.74s`. - Earlier post-merge gates: Notes tests `179 passed; 0 failed; 1 ignored`; Files clippy passed; Files tests `146 passed; 0 failed; 1 ignored`. - `cargo clippy -p calternal-server --all-targets -- -D warnings`: `Finished `dev` profile [unoptimized + debuginfo] target(s) in 5m 07s`. - `cargo test -p calternal-server` was still compiling at the five-hour job limit. - `bun run check`: `svelte-check found 0 errors and 0 warnings`. - `bun run test`: `Test Files 153 passed (153)`; `Tests 1056 passed (1056)`. - Production `bun run build` passed earlier: `✓ built in 37.06s`; `Wrote site to "build"`; `✔ done`. ## Performance and gaps The build host load average was `25.43, 26.28, 24.00` on the shared 8-core host. The benchmark accepted a 60-second recording, but its job remained leased for over 20 minutes without an error; I stopped the run at the time limit. I do not have a completed local RTF, the 30-minute local result, or perf VM RTF/RSS numbers. The existing release server on the perf VM did not contain the Voice route. The adversarial round and `cargo test -p calternal-server` remain incomplete. I did not run `cargo clean` before the time limit. ## Decisions - A Journal Log recording belongs in its linked child Note, following owner decision #304. - Long audio uses five-minute recognition chunks with a half-second overlap. Backfill de-duplicates by audio content hash per User and keeps both historical folder spellings. - Keyboard-triggered UI motion follows the 2026-10-01 owner override and reduced-motion preferences. ## Production screenshots - [calendar-recorder-1440-dark.png](https://git.kayg.org/attachments/05d3781d-0290-4f3c-a815-f84b2d2d770b) - [calendar-recorder-1440-light.png](https://git.kayg.org/attachments/d29e3fd9-2537-4d74-a938-9e80986eb6f4) - [calendar-recorder-390-dark.png](https://git.kayg.org/attachments/39136a56-5524-4a9e-94aa-fff98ff1461f) - [calendar-recorder-390-light.png](https://git.kayg.org/attachments/a6daa0b1-58e0-485c-bb3e-e85b02111524) - [calendar-recorder-820-dark.png](https://git.kayg.org/attachments/dc611889-6788-4f6e-a323-43c6fde03c1f) - [calendar-recorder-820-light.png](https://git.kayg.org/attachments/d1f1790f-e6aa-4276-b1f6-36dca2a86cc0) - [maintenance-jobs-1440-dark.png](https://git.kayg.org/attachments/b2b99b34-7fba-49d2-9715-63d0999e9e57) - [maintenance-jobs-1440-light.png](https://git.kayg.org/attachments/045cc09b-5c0f-4618-9d7a-174b88be69ff) - [maintenance-jobs-390-dark.png](https://git.kayg.org/attachments/d2047524-fef3-4411-8277-04f0b28a46a6) - [maintenance-jobs-390-light.png](https://git.kayg.org/attachments/c36e7857-fbf4-46a2-842d-aa8261c21b05) - [maintenance-jobs-820-dark.png](https://git.kayg.org/attachments/85e893c5-83bb-46df-929f-dcd27ff40645) - [maintenance-jobs-820-light.png](https://git.kayg.org/attachments/fe2c6950-0413-4844-9b10-6a1626e59d97) - [search-voice-1440-dark.png](https://git.kayg.org/attachments/b86cb5f2-51d8-4d8e-b4d8-581919ab2872) - [search-voice-1440-light.png](https://git.kayg.org/attachments/24ba9f1c-0223-476a-bf2a-87479c289b3b) - [search-voice-390-dark.png](https://git.kayg.org/attachments/2a31d93e-ce25-4ce0-969a-6c71c5dfc193) - [search-voice-390-light.png](https://git.kayg.org/attachments/38f63f78-91c9-4936-addf-cccf50f43a51) - [search-voice-820-dark.png](https://git.kayg.org/attachments/53766cd5-644e-418a-8a12-021f2fbef2d7) - [search-voice-820-light.png](https://git.kayg.org/attachments/8635ef71-b5bf-47b8-8285-0d777f1832f1) - [transcript-note-1440-dark.png](https://git.kayg.org/attachments/f3de72d7-52ae-4d55-922c-d94503e88245) - [transcript-note-1440-light.png](https://git.kayg.org/attachments/814be8bd-4cd2-4606-a693-8dcf135566d3) - [transcript-note-390-dark.png](https://git.kayg.org/attachments/b13475e1-1dde-4e1a-a1ae-37697220035b) - [transcript-note-390-light.png](https://git.kayg.org/attachments/fbf4e58f-4761-4685-9483-f5f1d84aefa6) - [transcript-note-820-dark.png](https://git.kayg.org/attachments/2652742f-d40a-4178-932a-26b9d59634b0) - [transcript-note-820-light.png](https://git.kayg.org/attachments/b2b865e6-a308-4b4b-99fd-2b6711959949) - [voice-settings-1440-dark.png](https://git.kayg.org/attachments/5da763b6-b4a8-41bf-9d14-eb279ffc35b3) - [voice-settings-1440-light.png](https://git.kayg.org/attachments/19727d51-7818-4b21-a686-adb9ff789760) - [voice-settings-390-dark.png](https://git.kayg.org/attachments/1e62dc24-aafe-44be-a89d-6437f90b4d3d) - [voice-settings-390-light.png](https://git.kayg.org/attachments/dc5521be-7f2d-45ef-be44-3188481cb699) - [voice-settings-820-dark.png](https://git.kayg.org/attachments/22735569-ed12-47ce-91d7-c327a56bd42e) - [voice-settings-820-light.png](https://git.kayg.org/attachments/abb7a74a-caca-4c51-8a11-7b01b6c8ab3c)
Author
Owner

Report correction: cleanup completed immediately after the report was posted. cargo clean output was Removed 22920 files, 18.0GiB total; apps/web/build and apps/web/.svelte-kit were removed, and git status --short is clean at 4011a4fad.

Report correction: cleanup completed immediately after the report was posted. `cargo clean` output was `Removed 22920 files, 18.0GiB total`; `apps/web/build` and `apps/web/.svelte-kit` were removed, and `git status --short` is clean at `4011a4fad`.
Author
Owner

Round 2 started on branch job/voice-619, based on origin/dev at c4a61e8cf090170f35b1bed3350d9de20c83ecd5. Current head: 4011a4fadd83d82123d2e7f1de37f5168e2917a7. I am tracing the >20-minute benchmark lease, backfill throttling and cancellation, and model/runtime licences before running the perf and adversarial checks.

Round 2 started on branch `job/voice-619`, based on `origin/dev` at `c4a61e8cf090170f35b1bed3350d9de20c83ecd5`. Current head: `4011a4fadd83d82123d2e7f1de37f5168e2917a7`. I am tracing the >20-minute benchmark lease, backfill throttling and cancellation, and model/runtime licences before running the perf and adversarial checks.
Author
Owner

Finding: the durable worker renews a Voice Job lease every lease_for / 3 while its handler future is alive. Voice inference runs on spawn_blocking, so native model calls have no bounded execution time or progress updates; even a wedged model thread keeps the Job leased. The previous 60-second probe was on a host at load average 25.43/26.28/24.00, while the locked perf VM is at 0.04/0.31/0.74. I am reproducing on the VM before selecting the watchdog threshold.

Finding: the durable worker renews a Voice Job lease every `lease_for / 3` while its handler future is alive. Voice inference runs on `spawn_blocking`, so native model calls have no bounded execution time or progress updates; even a wedged model thread keeps the Job leased. The previous 60-second probe was on a host at load average 25.43/26.28/24.00, while the locked perf VM is at 0.04/0.31/0.74. I am reproducing on the VM before selecting the watchdog threshold.
Author
Owner

Round 2 finding: the perf VM was idle inside /root/perf.lock (load average 0.04 / 0.31 / 0.74; 7.0 GiB available), but its installed server binary is from Sep 29 and does not expose the Voice routes. I am building the branch's release server locally before reproducing the 60-second and 30-minute jobs. The earlier >20-minute lease cannot yet be attributed to inference cost until this current build runs on the VM.

Round 2 finding: the perf VM was idle inside `/root/perf.lock` (load average 0.04 / 0.31 / 0.74; 7.0 GiB available), but its installed server binary is from Sep 29 and does not expose the Voice routes. I am building the branch's release server locally before reproducing the 60-second and 30-minute jobs. The earlier >20-minute lease cannot yet be attributed to inference cost until this current build runs on the VM.
Author
Owner

Round 2 finding (2026-10-02): bun run test completed with 1,039 passed and 19 failed across 16 files. The failure report lists each as Test timed out in 5000ms; most are unrelated existing UI/unit tests, consistent with this host's load. I will keep the Voice-specific web checks separate and include the exact gate summary in the final report.

License check: pinned Parakeet TDT v2 INT8 weights are CC BY 4.0; commercial use is allowed with attribution and the INT8 conversion must be identified. S1-mini by Superwhisper uses Apache-2.0 plus its exact-name term; commercial use is allowed while keeping that name condition. Runtime sherpa-onnx 1.13.8 is Apache-2.0 and llama-cpp-2 0.1.157 is MIT OR Apache-2.0. I added these attribution and runtime terms to NOTICE; no separate About/licences page exists in the current app.

Round 2 finding (2026-10-02): `bun run test` completed with 1,039 passed and 19 failed across 16 files. The failure report lists each as `Test timed out in 5000ms`; most are unrelated existing UI/unit tests, consistent with this host's load. I will keep the Voice-specific web checks separate and include the exact gate summary in the final report. License check: pinned Parakeet TDT v2 INT8 weights are CC BY 4.0; commercial use is allowed with attribution and the INT8 conversion must be identified. S1-mini by Superwhisper uses Apache-2.0 plus its exact-name term; commercial use is allowed while keeping that name condition. Runtime sherpa-onnx 1.13.8 is Apache-2.0 and llama-cpp-2 0.1.157 is MIT OR Apache-2.0. I added these attribution and runtime terms to `NOTICE`; no separate About/licences page exists in the current app.
Author
Owner

Benchmark finding (2026-10-02): the concurrent-backfill branch originally reused the baseline 60-second byte stream. Voice transcript caching keys on audio bytes, so the interactive request could have returned from cache and understated latency. Commit 980802151 now requires unique audio bytes for baseline, interactive, and backfill inputs; it requires the backfill stress input to be a bounded 30-minute WebM and reports sample count, mean, p50, and p95. Each case has one sample in this run, so p50/p95 will equal that sample and will be labelled as such.

Benchmark finding (2026-10-02): the concurrent-backfill branch originally reused the baseline 60-second byte stream. Voice transcript caching keys on audio bytes, so the interactive request could have returned from cache and understated latency. Commit `980802151` now requires unique audio bytes for baseline, interactive, and backfill inputs; it requires the backfill stress input to be a bounded 30-minute WebM and reports sample count, mean, p50, and p95. Each case has one sample in this run, so p50/p95 will equal that sample and will be labelled as such.
Author
Owner

Implementation finding (2026-10-02): a direct transcription could hold a lease while waiting on the shared model downloader, but its User Job progress stayed at the prior decode stage. Commit f89b262a9 now reports downloaded/total model bytes on that transcription Job every five seconds and fails visibly if the direct wait exceeds 30 minutes. The shared Voice settings meter remains the source of byte counts.

Implementation finding (2026-10-02): a direct transcription could hold a lease while waiting on the shared model downloader, but its User Job progress stayed at the prior decode stage. Commit `f89b262a9` now reports downloaded/total model bytes on that transcription Job every five seconds and fails visibly if the direct wait exceeds 30 minutes. The shared Voice settings meter remains the source of byte counts.
Author
Owner

Round 2 progress (2026-10-02): merged origin/dev at 8e7d20f1322337ed7d700c2cc725d1c3199ff99e (merge commit on job/voice-619). I started the release build and notes clippy before final gates. They have not reached the application crate after about 1h55 and 48m; /proc/loadavg was 138.54 133.60 129.77, with over 6,700 runnable/total tasks and many concurrent Cargo jobs. The perf VM's existing server binary lacks the Voice route, so no Voice latency sample is valid yet. I will continue using the current build and report measured values only if the Voice server is available.

Round 2 progress (2026-10-02): merged `origin/dev` at `8e7d20f1322337ed7d700c2cc725d1c3199ff99e` (merge commit on `job/voice-619`). I started the release build and notes clippy before final gates. They have not reached the application crate after about 1h55 and 48m; `/proc/loadavg` was `138.54 133.60 129.77`, with over 6,700 runnable/total tasks and many concurrent Cargo jobs. The perf VM's existing server binary lacks the Voice route, so no Voice latency sample is valid yet. I will continue using the current build and report measured values only if the Voice server is available.
Author
Owner

Round 2 follow-up: ffprobe confirms separate benchmark inputs are exactly 60.008s and 1800.008s; the interactive and backfill inputs use distinct bytes to avoid transcript-cache hits. The saved production screenshots cover 30 cases (five touched surfaces × 390/820/1440 × light/dark), but they predate the updated paused-Job screenshot capture, so I will regenerate that set if a server binary becomes available. The local release build is still at generic dependencies and has no executable yet; sccache reported its server had shut down and Cargo fell back to local compilation under the same high host load.

Round 2 follow-up: `ffprobe` confirms separate benchmark inputs are exactly 60.008s and 1800.008s; the interactive and backfill inputs use distinct bytes to avoid transcript-cache hits. The saved production screenshots cover 30 cases (five touched surfaces × 390/820/1440 × light/dark), but they predate the updated paused-Job screenshot capture, so I will regenerate that set if a server binary becomes available. The local release build is still at generic dependencies and has no executable yet; sccache reported its server had shut down and Cargo fell back to local compilation under the same high host load.
Author
Owner

Round 2 build constraint: after 2h27, cargo build --release -p calternal-server had compiled generic dependencies only and /mnt/hdd/targets/jobs/voice-619/release/calternal-server did not exist. The shared host load average was 139.01 149.78 149.32; the server test was waiting on the debug build-directory lock. I stopped only my release-build process (exit 130) to let the required test and adversarial probe proceed. The perf VM still has no Voice-capable binary from this branch, so no perf measurement has been claimed.

Round 2 build constraint: after 2h27, `cargo build --release -p calternal-server` had compiled generic dependencies only and `/mnt/hdd/targets/jobs/voice-619/release/calternal-server` did not exist. The shared host load average was `139.01 149.78 149.32`; the server test was waiting on the debug build-directory lock. I stopped only my release-build process (exit 130) to let the required test and adversarial probe proceed. The perf VM still has no Voice-capable binary from this branch, so no perf measurement has been claimed.
Author
Owner

The requested cargo test -p calternal-server is still blocked after 22 minutes. Its captured output contains only three Blocking waiting for file lock on package cache lines and one Blocking waiting for file lock on build directory; the Cargo process is in uninterruptible filesystem I/O on the shared build volume, and no test has started or failed. The build directory lock became available when I stopped notes clippy, but shared Cargo/cache I/O remains saturated.

The requested `cargo test -p calternal-server` is still blocked after 22 minutes. Its captured output contains only three `Blocking waiting for file lock on package cache` lines and one `Blocking waiting for file lock on build directory`; the Cargo process is in uninterruptible filesystem I/O on the shared build volume, and no test has started or failed. The build directory lock became available when I stopped notes clippy, but shared Cargo/cache I/O remains saturated.
Author
Owner

Forgejo #619 — round 2 report

Branch: job/voice-619
Head: 5465d06eb6d9d478a244bd09e540d4c85e28e6b

Built

  • Added bounded Voice worker waits, retryable/final visible Job errors, durable transcription progress, and model-download byte progress. Added a regression test that leases a real queue Job, hangs a worker, and checks that the Job fails visibly instead of staying leased forever.
  • Limited engine concurrency to one direct transcription and one low-priority backfill; backfill uses nice/idle I/O, bounded paged scans and bounded audio memory, and can pause/resume from Settings → Jobs.
  • Added the isolated ADVERSARIAL_VOICE_ONLY=1 runner path and the hostile Voice probes. Extended the Voice benchmark to reject duplicate recording bytes and measure interactive latency while a real backfill runs.
  • Added Job progress, details, pause/resume, stop, and Copy link actions in Maintenance Jobs; linked Composer and Search notifications to durable Job details.
  • Verified the pinned model and runtime license terms and added the required attribution/license texts to NOTICE and licenses/voice/.

Key files: crates/plugins/notes/src/voice.rs, voice_engine.rs, voice_models.rs, crates/plugins/notes/migrations/0025_voice_transcripts.sql, apps/web/src/routes/settings/jobs/MyJobsGroup.svelte, JobDetails.svelte, bench/voice-transcription.py, tests/adversarial/voice_transcription.py, tests/adversarial/run.sh, contracts/openapi.json, packages/api-client/src/generated.ts, and NOTICE/licenses/voice/. The branch diff against origin/dev is 70 files.

Gate output

cargo fmt --check:

cargo fmt --check exit: 0

bun run check:

svelte-check found 0 errors and 0 warnings

bun run test:

Test Files 16 failed | 137 passed (153)
Tests 19 failed | 1039 passed (1058)
Duration 727.71s (transform 71%, import 10%, environment 10%, tests 6%, setup 3%)

The 19 failures were 5-second test timeouts while the shared build host was heavily loaded.

cargo test -p calternal-server was stopped at 16:20 UTC (exit 130) before compilation reached the server or any test. Captured output ended while compiling dependencies; no test result is available. Notes clippy was also stopped before reaching the Notes crate. Server clippy and the Notes crate test were not run.

bash -n tests/adversarial/run.sh, the benchmark fixture/dedup checks, Python syntax checks, and node --check apps/web/e2e/voice-transcription-619.mjs completed successfully.

Performance and adversarial status

No RTF, peak RSS, CPU, or interactive latency measurement is claimed. The perf VM binary checked for this job lacked the Voice route. The local release build ran for 2h27 without producing a server executable under a host load average of 139.01 149.78 149.32; I stopped it to prioritize the requested test. I did not measure on the perf VM or hold its measurement lock, and I did not run the adversarial probe against a server without the Voice route.

Licenses

The pinned NVIDIA Parakeet INT8 export is CC BY 4.0: commercial use is allowed with attribution and change marking (pinned model, CC BY 4.0 terms). S1-mini permits commercial use under Apache-2.0 plus the model-name condition (model card). Runtime versions checked: sherpa-onnx 1.13.8 (Apache-2.0), llama-cpp-2 0.1.157 (MIT OR Apache-2.0), and linked llama-cpp-sys-2 0.1.158 (MIT) (sherpa-onnx, llama-cpp-2). Attribution and full notices are in NOTICE and licenses/voice/.

UX gaps closed / left

Closed: pause/resume and other standard Job actions are available from the Jobs menu and details view; progress and terminal states are visible; Composer and Search link to durable Job details. Existing screenshot files cover 30 desktop/tablet/phone and light/dark cases.

Left: the saved screenshots predate the macOS user-agent and paused-Job screenshot changes, so there is no fresh macOS production screenshot set. The screenshots were not attached to this issue: the Git credential helper had no non-interactive Forgejo credential. The production E2E script was not rerun because no Voice-capable server executable was available.

Decisions not covered by DESIGN

  • Keep two native engine slots: one for an interactive transcription and one for a background backfill. Give the backfill thread nice 19 and idle I/O priority.
  • Use a 30-minute model-download deadline with byte progress every five seconds. Use a transcription inference deadline of 300 seconds plus five seconds per validated audio second (10 minutes for 60 seconds; 155 minutes for 30 minutes).
  • Record attribution in root NOTICE because the app has no About/licences page.

No push, deploy, merge, or issue close was performed. cargo clean exited 0 (Removed 4914 files, 1.4GiB total); apps/web/build, apps/web/.svelte-kit, and the generated benchmark audio were removed.

# Forgejo #619 — round 2 report Branch: `job/voice-619` Head: `5465d06eb6d9d478a244bd09e540d4c85e28e6b` ## Built - Added bounded Voice worker waits, retryable/final visible Job errors, durable transcription progress, and model-download byte progress. Added a regression test that leases a real queue Job, hangs a worker, and checks that the Job fails visibly instead of staying leased forever. - Limited engine concurrency to one direct transcription and one low-priority backfill; backfill uses nice/idle I/O, bounded paged scans and bounded audio memory, and can pause/resume from Settings → Jobs. - Added the isolated `ADVERSARIAL_VOICE_ONLY=1` runner path and the hostile Voice probes. Extended the Voice benchmark to reject duplicate recording bytes and measure interactive latency while a real backfill runs. - Added Job progress, details, pause/resume, stop, and Copy link actions in Maintenance Jobs; linked Composer and Search notifications to durable Job details. - Verified the pinned model and runtime license terms and added the required attribution/license texts to `NOTICE` and `licenses/voice/`. Key files: `crates/plugins/notes/src/voice.rs`, `voice_engine.rs`, `voice_models.rs`, `crates/plugins/notes/migrations/0025_voice_transcripts.sql`, `apps/web/src/routes/settings/jobs/MyJobsGroup.svelte`, `JobDetails.svelte`, `bench/voice-transcription.py`, `tests/adversarial/voice_transcription.py`, `tests/adversarial/run.sh`, `contracts/openapi.json`, `packages/api-client/src/generated.ts`, and `NOTICE`/`licenses/voice/`. The branch diff against `origin/dev` is 70 files. ## Gate output `cargo fmt --check`: ```text cargo fmt --check exit: 0 ``` `bun run check`: ```text svelte-check found 0 errors and 0 warnings ``` `bun run test`: ```text Test Files 16 failed | 137 passed (153) Tests 19 failed | 1039 passed (1058) Duration 727.71s (transform 71%, import 10%, environment 10%, tests 6%, setup 3%) ``` The 19 failures were 5-second test timeouts while the shared build host was heavily loaded. `cargo test -p calternal-server` was stopped at 16:20 UTC (exit 130) before compilation reached the server or any test. Captured output ended while compiling dependencies; no test result is available. Notes clippy was also stopped before reaching the Notes crate. Server clippy and the Notes crate test were not run. `bash -n tests/adversarial/run.sh`, the benchmark fixture/dedup checks, Python syntax checks, and `node --check apps/web/e2e/voice-transcription-619.mjs` completed successfully. ## Performance and adversarial status No RTF, peak RSS, CPU, or interactive latency measurement is claimed. The perf VM binary checked for this job lacked the Voice route. The local release build ran for 2h27 without producing a server executable under a host load average of `139.01 149.78 149.32`; I stopped it to prioritize the requested test. I did not measure on the perf VM or hold its measurement lock, and I did not run the adversarial probe against a server without the Voice route. ## Licenses The pinned NVIDIA Parakeet INT8 export is CC BY 4.0: commercial use is allowed with attribution and change marking ([pinned model](https://huggingface.co/csukuangfj/sherpa-onnx-nemo-parakeet-tdt-0.6b-v2-int8/tree/b9bd32696df49cbb1e7e499238a22804c82f9ef7), [CC BY 4.0 terms](https://creativecommons.org/licenses/by/4.0/)). S1-mini permits commercial use under Apache-2.0 plus the model-name condition ([model card](https://huggingface.co/superwhisper/s1-mini-GGUF/blob/main/README.md)). Runtime versions checked: sherpa-onnx 1.13.8 (Apache-2.0), llama-cpp-2 0.1.157 (MIT OR Apache-2.0), and linked llama-cpp-sys-2 0.1.158 (MIT) ([sherpa-onnx](https://github.com/k2-fsa/sherpa-onnx/tree/v1.13.8), [llama-cpp-2](https://crates.io/crates/llama-cpp-2/0.1.157)). Attribution and full notices are in `NOTICE` and `licenses/voice/`. ## UX gaps closed / left Closed: pause/resume and other standard Job actions are available from the Jobs menu and details view; progress and terminal states are visible; Composer and Search link to durable Job details. Existing screenshot files cover 30 desktop/tablet/phone and light/dark cases. Left: the saved screenshots predate the macOS user-agent and paused-Job screenshot changes, so there is no fresh macOS production screenshot set. The screenshots were not attached to this issue: the Git credential helper had no non-interactive Forgejo credential. The production E2E script was not rerun because no Voice-capable server executable was available. ## Decisions not covered by DESIGN - Keep two native engine slots: one for an interactive transcription and one for a background backfill. Give the backfill thread nice 19 and idle I/O priority. - Use a 30-minute model-download deadline with byte progress every five seconds. Use a transcription inference deadline of 300 seconds plus five seconds per validated audio second (10 minutes for 60 seconds; 155 minutes for 30 minutes). - Record attribution in root `NOTICE` because the app has no About/licences page. No push, deploy, merge, or issue close was performed. `cargo clean` exited 0 (`Removed 4914 files, 1.4GiB total`); `apps/web/build`, `apps/web/.svelte-kit`, and the generated benchmark audio were removed.
Author
Owner

Correction to my final report: the exact current head is 5465d06eb6d9d478160a4c2449db95046607b09e. Cleanup completed: cargo clean exited 0 and removed 4,914 files; apps/web/build, apps/web/.svelte-kit, and the generated benchmark audio were removed. The final report file now contains the corrected full SHA.

Correction to my final report: the exact current head is `5465d06eb6d9d478160a4c2449db95046607b09e`. Cleanup completed: `cargo clean` exited 0 and removed 4,914 files; `apps/web/build`, `apps/web/.svelte-kit`, and the generated benchmark audio were removed. The final report file now contains the corrected full SHA.
Author
Owner

Starting independent read-only review on branch job/rev2-voice-619, reviewing job/voice-619 at 5465d06eb6d9d478160a4c2449db95046607b09e against origin/dev at 440e19dce23040ac8ebaae88f0469b6535b1afcb. I will inspect the Voice transcription worker and backfill paths, existing helper reuse, route isolation, tests, and commit the review report in this worktree. No build or tests will run.

Starting independent read-only review on branch `job/rev2-voice-619`, reviewing `job/voice-619` at `5465d06eb6d9d478160a4c2449db95046607b09e` against `origin/dev` at `440e19dce23040ac8ebaae88f0469b6535b1afcb`. I will inspect the Voice transcription worker and backfill paths, existing helper reuse, route isolation, tests, and commit the review report in this worktree. No build or tests will run.
Author
Owner

Independent read-only review committed as ff5167cdc2fde7e45523d87d499b7902167ac48c (review-voice-619.md). Findings are ranked by severity:

  1. P1 — Unbounded direct transcription queue and staged storage — crates/plugins/notes/src/voice.rs:418-427, 477-499, 667-674. Request concurrency is capped only while reading each body; queued Jobs and staged audio have no per-User count or byte bound, and crash orphans have no sweeper. The Composer progress link also points to owner-filtered Jobs, but the Voice Job has no owner (apps/web/src/lib/composer/Composer.svelte:941-949, crates/calternal-server/src/wire.rs:3642-3645, 3671-3678). Reserve per-User Job/byte capacity before staging, set owner_user_id, release capacity on terminal paths, and sweep unreferenced staged files.
  2. P1 — Shared transcript cache violates DESIGN §48 and survives User deletion — crates/plugins/notes/migrations/0025_voice_transcripts.sql:3-9, crates/plugins/notes/src/voice.rs:2058-2064. Move the cache into each User's private derived store and purge it during User deletion; add a deletion test.
  3. P1 — Timeout detaches native inference — crates/plugins/notes/src/voice.rs:1855-1894, 1923-1932. The native thread can retain memory and its semaphore slot after Tokio times out. Use a killable, reaped process with resource limits or a cancellable native API; test worker termination and slot recovery.
  4. P1 — Backfill reports success after destination errors and suppresses retry — crates/plugins/notes/src/voice.rs:626-645, 1053-1060, 1195-1232. Persist partial failures as retryable and mark complete only after all destinations update; a disabled run should be skipped/cancelled, not completed. Test a write failure and off-then-on setting change.
  5. P2 — Shared model download has no total 30-minute deadline — crates/plugins/notes/src/voice.rs:649-661, 1252-1281, crates/plugins/notes/src/voice_models.rs:141-145. Bound the shared download Job and the backfill wait; test a downloader that never resolves.
  6. P2 — Backfill scan and hashing block normal-priority async work without checkpoints — crates/plugins/notes/src/voice.rs:782-900, 913-953, 1039-1067, 1111-1153, 1284-1305, 1994-2006. Move bounded scan/hash pages to a low-priority blocking worker and checkpoint progress/cancellation after each page and file.
  7. P2 — Journal backfill can attach one recording's transcript to another recording's Note — crates/plugins/notes/src/voice.rs:1413-1430. Match the Note to the audio identity or add every source audio link when reusing a per-Log Note; test two recordings on one Log.

No weakened existing test expectations were found. This review did not build or test, start a server, use a browser, or run performance measurements, as required for the read-only review. docs/DESIGN.md ends at §57; §58 was not present.

Independent read-only review committed as `ff5167cdc2fde7e45523d87d499b7902167ac48c` (`review-voice-619.md`). Findings are ranked by severity: 1. **P1 — Unbounded direct transcription queue and staged storage** — `crates/plugins/notes/src/voice.rs:418-427, 477-499, 667-674`. Request concurrency is capped only while reading each body; queued Jobs and staged audio have no per-User count or byte bound, and crash orphans have no sweeper. The Composer progress link also points to owner-filtered Jobs, but the Voice Job has no owner (`apps/web/src/lib/composer/Composer.svelte:941-949`, `crates/calternal-server/src/wire.rs:3642-3645, 3671-3678`). Reserve per-User Job/byte capacity before staging, set `owner_user_id`, release capacity on terminal paths, and sweep unreferenced staged files. 2. **P1 — Shared transcript cache violates DESIGN §48 and survives User deletion** — `crates/plugins/notes/migrations/0025_voice_transcripts.sql:3-9`, `crates/plugins/notes/src/voice.rs:2058-2064`. Move the cache into each User's private derived store and purge it during User deletion; add a deletion test. 3. **P1 — Timeout detaches native inference** — `crates/plugins/notes/src/voice.rs:1855-1894, 1923-1932`. The native thread can retain memory and its semaphore slot after Tokio times out. Use a killable, reaped process with resource limits or a cancellable native API; test worker termination and slot recovery. 4. **P1 — Backfill reports success after destination errors and suppresses retry** — `crates/plugins/notes/src/voice.rs:626-645, 1053-1060, 1195-1232`. Persist partial failures as retryable and mark complete only after all destinations update; a disabled run should be skipped/cancelled, not completed. Test a write failure and off-then-on setting change. 5. **P2 — Shared model download has no total 30-minute deadline** — `crates/plugins/notes/src/voice.rs:649-661, 1252-1281`, `crates/plugins/notes/src/voice_models.rs:141-145`. Bound the shared download Job and the backfill wait; test a downloader that never resolves. 6. **P2 — Backfill scan and hashing block normal-priority async work without checkpoints** — `crates/plugins/notes/src/voice.rs:782-900, 913-953, 1039-1067, 1111-1153, 1284-1305, 1994-2006`. Move bounded scan/hash pages to a low-priority blocking worker and checkpoint progress/cancellation after each page and file. 7. **P2 — Journal backfill can attach one recording's transcript to another recording's Note** — `crates/plugins/notes/src/voice.rs:1413-1430`. Match the Note to the audio identity or add every source audio link when reusing a per-Log Note; test two recordings on one Log. No weakened existing test expectations were found. This review did not build or test, start a server, use a browser, or run performance measurements, as required for the read-only review. `docs/DESIGN.md` ends at §57; §58 was not present.
Author
Owner

Round 3 started on branch job/voice-619 at HEAD 5465d06eb6; merge base with origin/dev is 440e19dce2 (origin/dev currently c4faf184df). I have read the independent review and will address all P1/P2 findings with regressions before verification. The host is currently busy (load average 21.67 / 26.22 / 27.58); I will avoid starting resource-heavy gates while I work through the review.

Round 3 started on branch job/voice-619 at HEAD 5465d06eb6d9d478160a4c2449db95046607b09e; merge base with origin/dev is 440e19dce23040ac8ebaae88f0469b6535b1afcb (origin/dev currently c4faf184df726a9375ae0c13bdfb6018ac2cf57e). I have read the independent review and will address all P1/P2 findings with regressions before verification. The host is currently busy (load average 21.67 / 26.22 / 27.58); I will avoid starting resource-heavy gates while I work through the review.
Author
Owner

The first Notes crate test run after the review fixes completed with 191 passed, 7 failed and 1 ignored (885.47 s). The failures were: imap::tests::captured_append_store_expunge_executes_on_real_home (timeout), tests::daily_and_composer_preserve_unrelated_bytes (404 vs 200), tests::fs_change_batch_does_not_reindex_unchanged_notes (projection not updated), tests::journal_patch_storm_finishes_with_one_winner_and_stale_preconditions (timeout), tests::journal_read_does_not_wait_for_reconciliation (100 ms timeout), tests::apple_replay_area_delete_reparents_two_task_links_and_syncs (NotFound), and the new Voice page-cancellation regression (observed Leased before the worker persisted its terminal state). Host load was 60.80 / 62.57 / 67.45 when checked. The six non-Voice failures appear unrelated to this issue; I have not changed their expectations. I fixed the Voice regression's completion synchronization and am rerunning that test alone.

The first Notes crate test run after the review fixes completed with 191 passed, 7 failed and 1 ignored (885.47 s). The failures were: `imap::tests::captured_append_store_expunge_executes_on_real_home` (timeout), `tests::daily_and_composer_preserve_unrelated_bytes` (404 vs 200), `tests::fs_change_batch_does_not_reindex_unchanged_notes` (projection not updated), `tests::journal_patch_storm_finishes_with_one_winner_and_stale_preconditions` (timeout), `tests::journal_read_does_not_wait_for_reconciliation` (100 ms timeout), `tests::apple_replay_area_delete_reparents_two_task_links_and_syncs` (NotFound), and the new Voice page-cancellation regression (observed `Leased` before the worker persisted its terminal state). Host load was 60.80 / 62.57 / 67.45 when checked. The six non-Voice failures appear unrelated to this issue; I have not changed their expectations. I fixed the Voice regression's completion synchronization and am rerunning that test alone.
Author
Owner

The first Voice adversarial runner invocation stopped at its offline cross-User classification gate before starting a server. The new transcription operation exposed unclassified target_id; after adding that rule, the same gate exposed the new detail route's job_id. I added explicit classifications for both, mapped target_id to an owned Note fixture, and extended the Voice probe to post a valid WAV from the second User against the first User's Note and verify that no Voice Job is created. The classification gates now pass: 340 OpenAPI operations / 945 generated entry points, and 39 admin operations. I am rerunning the Voice-only server probe.

The first Voice adversarial runner invocation stopped at its offline cross-User classification gate before starting a server. The new transcription operation exposed unclassified `target_id`; after adding that rule, the same gate exposed the new detail route's `job_id`. I added explicit classifications for both, mapped `target_id` to an owned Note fixture, and extended the Voice probe to post a valid WAV from the second User against the first User's Note and verify that no Voice Job is created. The classification gates now pass: 340 OpenAPI operations / 945 generated entry points, and 39 admin operations. I am rerunning the Voice-only server probe.
Author
Owner

The production E2E run captured the real paused/running/completed backfill states, timeout-failed Job, and Composer Job details, then timed out waiting for the Search preview's Record voice note action. The probe selected the first result matching the Log title without requiring data-kind="log" or activating that row; Search can return both a Journal Log and its linked Note with the same title, while only the Log preview owns the recorder. I updated the probe to hover the matching Log result, assert it is selected, and wait for the action inside that preview. I also reduced only the screenshot backfill fixture to five minutes; the separate local benchmark still uses its byte-distinct 30-minute backfill.

The production E2E run captured the real paused/running/completed backfill states, timeout-failed Job, and Composer Job details, then timed out waiting for the Search preview's `Record voice note` action. The probe selected the first result matching the Log title without requiring `data-kind="log"` or activating that row; Search can return both a Journal Log and its linked Note with the same title, while only the Log preview owns the recorder. I updated the probe to hover the matching Log result, assert it is selected, and wait for the action inside that preview. I also reduced only the screenshot backfill fixture to five minutes; the separate local benchmark still uses its byte-distinct 30-minute backfill.
Author
Owner

The first completed UI run reached the real benchmark and failed with HTTP 403 at the Files TUS backfill fixture. The benchmark sent Origin on its JSON and direct Voice requests, but its TUS create and patch requests omitted it. I added the same-origin Origin header to both TUS requests and am rerunning the browser capture and benchmark against this branch's server binary.

The first completed UI run reached the real benchmark and failed with HTTP 403 at the Files TUS backfill fixture. The benchmark sent Origin on its JSON and direct Voice requests, but its TUS create and patch requests omitted it. I added the same-origin Origin header to both TUS requests and am rerunning the browser capture and benchmark against this branch's server binary.
Author
Owner

The UI run completed its backfill before the contention benchmark. The benchmark then waited for a second scan, but enqueue_voice_backfill correctly skips scheduling while an owner has a completed scan. I confirmed the guard from the Voice code and observed that no pending or leased backfill existed. The E2E fixture now marks only its completed test Job as cancelled, then lets the benchmark call the normal backfill endpoint; the live run entered leased / Transcribing voice memo and started the interactive burst.

The UI run completed its backfill before the contention benchmark. The benchmark then waited for a second scan, but `enqueue_voice_backfill` correctly skips scheduling while an owner has a completed scan. I confirmed the guard from the Voice code and observed that no pending or leased backfill existed. The E2E fixture now marks only its completed test Job as cancelled, then lets the benchmark call the normal backfill endpoint; the live run entered `leased / Transcribing voice memo` and started the interactive burst.
Author
Owner

Voice #619 verification report

Head: 5005b5ac652a2dc2b1b596d1dfe0334acc13896f (Document Voice E2E review flow). The branch is committed and clean. origin/dev was merged once before final gates (bd2db6301); no push or deploy was performed.

Built

  • Stored recordings, transcripts, and cache data in per-User private derived storage, with opaque names and cleanup on User deletion. Added migration 0026_voice_user_store.sql and removed the old shared transcript table.
  • Added owner checks, per-User active-job and upload limits, orphan cleanup, and hostile cross-User destination coverage.
  • Moved inference into a killable server worker with bounded protocol, memory, CPU, and deadline controls. Made model download and backfill waits bounded; made retries and cancellation durable; paged low-priority scans.
  • Added Voice Job labels, stable Job links, and Composer/Search navigation to Job details. Added Settings → Jobs coverage for running, paused, done, and timeout-failed Jobs.
  • Added benchmark profile voice_transcription_619 to docs/perf/baseline.json, a Voice adversarial probe, and a Voice E2E flow.

UX gaps closed

  • Composer recording/transcription and Note attachment paths use the real API. A member targeting another User’s Note receives 404 and does not create a Job.
  • Composer and Search expose working links to Job details. Jobs expose running, paused, done, failed, and detail views. The E2E flow exercised these routes and the Voice backfill/search path.
  • Captured current Settings → Jobs and Composer/Search details at 390, 820, and 1440 px, in light and dark themes, with macOS platform emulation.

UX gaps left: none found in the Voice flow covered here. The full web suite and broad E2E suite remain for the merge round under the verification policy.

Decisions where DESIGN.md is silent

DESIGN.md has no Voice-specific worker protocol, limits, or storage layout. I used its §48 derived-data isolation direction and chose a private per-User store, a killable child process with a bounded framed protocol, a 30-minute / 32 MiB recording cap, and at most four active or pending Voice jobs per User. These are implementation choices for owner confirmation.

Verification

Passed:

cargo fmt --check
(exit 0; no output)

cargo clippy -p calternal-fs --all-targets -- -D warnings
Checking calternal-fs v0.1.0 ...
Finished `dev` profile [unoptimized + debuginfo] target(s) in 1m 01s

cargo clippy -p calternal-plugin-notes --all-targets -- -D warnings
Checking calternal-plugin-notes v0.0.1 ...
Finished `dev` profile [unoptimized + debuginfo] target(s) in 3m 36s

cargo clippy -p calternal-server --all-targets -- -D warnings
Finished `dev` profile [unoptimized + debuginfo] target(s) in 1m 13s

cargo test -p calternal-plugin-notes --lib voice::tests::voice_backfill_scan_is_paged_low_priority_and_cancellable
Finished `test` profile [unoptimized + debuginfo] target(s) in 7m 51s
Running unittests src/lib.rs (...)
running 1 test
test voice::tests::voice_backfill_scan_is_paged_low_priority_and_cancellable ... ok

test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 198 filtered out; finished in 106.49s

cargo test -p calternal-server
test result: ok. 107 passed; 0 failed; 3 ignored; 0 measured; 0 filtered out; finished in 27.23s

bun run check
User browser caches use userStorage; only documented device/public-link exceptions remain.
Text sizes and UI shape values use shared role tokens.
UI transitions and animation options use shared motion tokens or documented exceptions.
Loading svelte-check in workspace: /home/kayg/Developer/calternal-wt/voice-619/apps/web
Getting Svelte diagnostics...

svelte-check found 0 errors and 0 warnings

Voice adversarial probe
Voice adversarial probe passed: auth, malformed and oversized uploads, hostile target, concurrency, User isolation, and bounded one-hour decode.

The full Notes test gate had a shared-host failure. One cancellation race introduced by this work was fixed and its focused regression gate above passed. The remaining failures were unrelated timeout/state failures under heavy host load; no existing expectations were changed.

test result: FAILED. 191 passed; 7 failed; 1 ignored; 0 measured; 0 filtered out; finished in 885.47s

error: test failed, to rerun pass `-p calternal-plugin-notes --lib`

Other checks passed: cargo test -p calternal-fs (53 unit tests and 42 storage tests), cargo test -p calternal-server (107 passed), node --check apps/web/e2e/voice-transcription-619.mjs, python3 -m py_compile bench/voice-transcription.py, JSON validation, git diff --check, and the cross-User classification gates (340 operations; 945 generated entry points; 39 admin operations).

Performance

Local only; the perf VM was not used. docs/perf/baseline.json records no prior Voice baseline (baseline: null). The local host was loaded (load average 10.68 before and 17.75 after). The final benchmark reported:

  • One-minute audio: 24.265 s processing, RTF 0.4044; p50/p95 24.265 s; sampled peak server process-tree RSS 479,064 KiB.
  • Four one-minute requests during a leased backfill: queue/upload p95 2.587 s; processing p50 157.940 s / p95 159.276 s; sampled peak RSS 444,256 KiB. All four transcripts became searchable.
  • Thirty-minute audio: 768.211 s processing, RTF 0.4269; sampled peak RSS 565,520 KiB; transcript searchable.

Screenshots

74 PNGs are under artifacts/voice-619 (ignored by Git). Required groups include maintenance-jobs-{running,paused,done,failed-timeout}-{390,820,1440}-{light,dark}.png, failed-job detail captures, and Composer/Search Job detail captures at all three widths and both themes. Screenshots were not committed.

For the merge round

  • cd apps/web && bun run test --maxWorkers=2 — run the full web suite on the combined branch.
  • cargo test -p calternal-plugin-notes — rerun the Notes suite on the combined branch and determine whether the remaining seven host-load failures reproduce.
  • tests/adversarial/run.sh — run the full adversarial matrices required by the merge round.
  • The Voice-specific E2E already passed here: cd apps/web && bun run test:e2e:voice-619.

Cargo build output was cleaned with cargo clean after verification. No known Voice-specific functional gaps remain; the Notes full-suite failures and local-only performance measurement are the outstanding verification limits.

## Voice #619 verification report Head: `5005b5ac652a2dc2b1b596d1dfe0334acc13896f` (`Document Voice E2E review flow`). The branch is committed and clean. `origin/dev` was merged once before final gates (`bd2db6301`); no push or deploy was performed. ### Built - Stored recordings, transcripts, and cache data in per-User private derived storage, with opaque names and cleanup on User deletion. Added migration `0026_voice_user_store.sql` and removed the old shared transcript table. - Added owner checks, per-User active-job and upload limits, orphan cleanup, and hostile cross-User destination coverage. - Moved inference into a killable server worker with bounded protocol, memory, CPU, and deadline controls. Made model download and backfill waits bounded; made retries and cancellation durable; paged low-priority scans. - Added Voice Job labels, stable Job links, and Composer/Search navigation to Job details. Added Settings → Jobs coverage for running, paused, done, and timeout-failed Jobs. - Added benchmark profile `voice_transcription_619` to `docs/perf/baseline.json`, a Voice adversarial probe, and a Voice E2E flow. ### UX gaps closed - Composer recording/transcription and Note attachment paths use the real API. A member targeting another User’s Note receives 404 and does not create a Job. - Composer and Search expose working links to Job details. Jobs expose running, paused, done, failed, and detail views. The E2E flow exercised these routes and the Voice backfill/search path. - Captured current Settings → Jobs and Composer/Search details at 390, 820, and 1440 px, in light and dark themes, with macOS platform emulation. UX gaps left: none found in the Voice flow covered here. The full web suite and broad E2E suite remain for the merge round under the verification policy. ### Decisions where DESIGN.md is silent DESIGN.md has no Voice-specific worker protocol, limits, or storage layout. I used its §48 derived-data isolation direction and chose a private per-User store, a killable child process with a bounded framed protocol, a 30-minute / 32 MiB recording cap, and at most four active or pending Voice jobs per User. These are implementation choices for owner confirmation. ### Verification Passed: ```text cargo fmt --check (exit 0; no output) cargo clippy -p calternal-fs --all-targets -- -D warnings Checking calternal-fs v0.1.0 ... Finished `dev` profile [unoptimized + debuginfo] target(s) in 1m 01s cargo clippy -p calternal-plugin-notes --all-targets -- -D warnings Checking calternal-plugin-notes v0.0.1 ... Finished `dev` profile [unoptimized + debuginfo] target(s) in 3m 36s cargo clippy -p calternal-server --all-targets -- -D warnings Finished `dev` profile [unoptimized + debuginfo] target(s) in 1m 13s cargo test -p calternal-plugin-notes --lib voice::tests::voice_backfill_scan_is_paged_low_priority_and_cancellable Finished `test` profile [unoptimized + debuginfo] target(s) in 7m 51s Running unittests src/lib.rs (...) running 1 test test voice::tests::voice_backfill_scan_is_paged_low_priority_and_cancellable ... ok test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 198 filtered out; finished in 106.49s cargo test -p calternal-server test result: ok. 107 passed; 0 failed; 3 ignored; 0 measured; 0 filtered out; finished in 27.23s bun run check User browser caches use userStorage; only documented device/public-link exceptions remain. Text sizes and UI shape values use shared role tokens. UI transitions and animation options use shared motion tokens or documented exceptions. Loading svelte-check in workspace: /home/kayg/Developer/calternal-wt/voice-619/apps/web Getting Svelte diagnostics... svelte-check found 0 errors and 0 warnings Voice adversarial probe Voice adversarial probe passed: auth, malformed and oversized uploads, hostile target, concurrency, User isolation, and bounded one-hour decode. ``` The full Notes test gate had a shared-host failure. One cancellation race introduced by this work was fixed and its focused regression gate above passed. The remaining failures were unrelated timeout/state failures under heavy host load; no existing expectations were changed. ```text test result: FAILED. 191 passed; 7 failed; 1 ignored; 0 measured; 0 filtered out; finished in 885.47s error: test failed, to rerun pass `-p calternal-plugin-notes --lib` ``` Other checks passed: `cargo test -p calternal-fs` (53 unit tests and 42 storage tests), `cargo test -p calternal-server` (107 passed), `node --check apps/web/e2e/voice-transcription-619.mjs`, `python3 -m py_compile bench/voice-transcription.py`, JSON validation, `git diff --check`, and the cross-User classification gates (340 operations; 945 generated entry points; 39 admin operations). ### Performance Local only; the perf VM was not used. `docs/perf/baseline.json` records no prior Voice baseline (`baseline: null`). The local host was loaded (load average 10.68 before and 17.75 after). The final benchmark reported: - One-minute audio: 24.265 s processing, RTF 0.4044; p50/p95 24.265 s; sampled peak server process-tree RSS 479,064 KiB. - Four one-minute requests during a leased backfill: queue/upload p95 2.587 s; processing p50 157.940 s / p95 159.276 s; sampled peak RSS 444,256 KiB. All four transcripts became searchable. - Thirty-minute audio: 768.211 s processing, RTF 0.4269; sampled peak RSS 565,520 KiB; transcript searchable. ### Screenshots 74 PNGs are under `artifacts/voice-619` (ignored by Git). Required groups include `maintenance-jobs-{running,paused,done,failed-timeout}-{390,820,1440}-{light,dark}.png`, failed-job detail captures, and Composer/Search Job detail captures at all three widths and both themes. Screenshots were not committed. ### For the merge round - `cd apps/web && bun run test --maxWorkers=2` — run the full web suite on the combined branch. - `cargo test -p calternal-plugin-notes` — rerun the Notes suite on the combined branch and determine whether the remaining seven host-load failures reproduce. - `tests/adversarial/run.sh` — run the full adversarial matrices required by the merge round. - The Voice-specific E2E already passed here: `cd apps/web && bun run test:e2e:voice-619`. Cargo build output was cleaned with `cargo clean` after verification. No known Voice-specific functional gaps remain; the Notes full-suite failures and local-only performance measurement are the outstanding verification limits.
Author
Owner

Per-crate gate addendum

The final crate audit found calternal-plugin and calternal-plugin-files were also touched. Their missing per-crate gates are now complete:

cargo clippy -p calternal-plugin --all-targets -- -D warnings
Checking calternal-plugin v0.0.1 (/home/kayg/Developer/calternal-wt/voice-619/crates/calternal-plugin)
Finished `dev` profile [unoptimized + debuginfo] target(s) in 59.94s

cargo test -p calternal-plugin
test result: ok. 25 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 3.03s

cargo clippy -p calternal-plugin-files --all-targets -- -D warnings
Checking calternal-plugin-files v0.0.1 (/home/kayg/Developer/calternal-wt/voice-619/crates/plugins/files)
Finished `dev` profile [unoptimized + debuginfo] target(s) in 3m 46s

cargo test -p calternal-plugin-files
test result: ok. 146 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 83.24s

cargo clean
     Removed 12360 files, 7.5GiB total

The Files test run printed two tests as running for over 60 seconds; both completed successfully. No code changes were needed. The working tree remains clean at 5005b5ac652a2dc2b1b596d1dfe0334acc13896f.

### Per-crate gate addendum The final crate audit found `calternal-plugin` and `calternal-plugin-files` were also touched. Their missing per-crate gates are now complete: ```text cargo clippy -p calternal-plugin --all-targets -- -D warnings Checking calternal-plugin v0.0.1 (/home/kayg/Developer/calternal-wt/voice-619/crates/calternal-plugin) Finished `dev` profile [unoptimized + debuginfo] target(s) in 59.94s cargo test -p calternal-plugin test result: ok. 25 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 3.03s cargo clippy -p calternal-plugin-files --all-targets -- -D warnings Checking calternal-plugin-files v0.0.1 (/home/kayg/Developer/calternal-wt/voice-619/crates/plugins/files) Finished `dev` profile [unoptimized + debuginfo] target(s) in 3m 46s cargo test -p calternal-plugin-files test result: ok. 146 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 83.24s cargo clean Removed 12360 files, 7.5GiB total ``` The Files test run printed two tests as running for over 60 seconds; both completed successfully. No code changes were needed. The working tree remains clean at `5005b5ac652a2dc2b1b596d1dfe0334acc13896f`.
Author
Owner

Correction to the Notes gate summary: the captured full cargo test -p calternal-plugin-notes result (191 passed; 7 failed; 1 ignored) is from before the backfill cancellation-race fix. Those seven failures comprised six shared-host timeout/state failures and the cancellation race. The race was fixed, and the focused cancellation regression test passed afterward. I did not rerun the full Notes suite after that fix, so the six other failures have not been rechecked. The prior wording that called them “remaining failures” was imprecise; please use this correction.

Correction to the Notes gate summary: the captured full `cargo test -p calternal-plugin-notes` result (`191 passed; 7 failed; 1 ignored`) is from before the backfill cancellation-race fix. Those seven failures comprised six shared-host timeout/state failures and the cancellation race. The race was fixed, and the focused cancellation regression test passed afterward. I did not rerun the full Notes suite after that fix, so the six other failures have not been rechecked. The prior wording that called them “remaining failures” was imprecise; please use this correction.
Author
Owner

Re-verification on a quiet host (load avg about 1), head b5d9c59941c1d2771c336085c189d58627f99578.

First run (before fix), cargo test -p calternal-plugin-notes -- --test-threads=4:

test tests::daily_and_composer_preserve_unrelated_bytes ... FAILED
thread 'tests::daily_and_composer_preserve_unrelated_bytes' panicked at crates/plugins/notes/src/lib.rs:9473:9:
assertion `left == right` failed
  left: 404
 right: 200
test result: FAILED. 197 passed; 1 failed; 1 ignored; 0 measured; 0 filtered out; finished in 6.25s

The other six failures from the last job run (busy host) did not occur again. They were host-load timeouts.

Classification: the remaining failure is the #954 fixture collision. Run alone, it passes 3 out of 3 times (test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 198 filtered out). In parallel, another test holds the process-global USER_LOCKS entry for the shared fixture User ID, so the Journal ID repair skips and the read returns a false 404. The branch did not have the fix. Fix: cherry-picked ff8e857c2 from job/datafix2 (each fixture Home gets its own User ID). No assertion changed. The voice.rs, bookmarks.rs and tasks_dav.rs tests that use the same literal ID only take the blocking lock_user, which serializes them and cannot cause the skip, so they did not need a change.

After the fix, two full runs with --no-fail-fast -- --test-threads=4:

run2: test result: ok. 198 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 7.40s
      apple_replay: test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.07s
run3: test result: ok. 198 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 6.24s
      apple_replay: test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.05s

cargo fmt --check -p calternal-plugin-notes: exit 0, no output.
cargo clippy -p calternal-plugin-notes --all-targets -- -D warnings: exit 0, Finished \dev` profile [unoptimized + debuginfo] target(s) in 1m 20s`.

Ready for merge from the Notes test side: yes. Not pushed.

Re-verification on a quiet host (load avg about 1), head `b5d9c59941c1d2771c336085c189d58627f99578`. **First run (before fix)**, `cargo test -p calternal-plugin-notes -- --test-threads=4`: ``` test tests::daily_and_composer_preserve_unrelated_bytes ... FAILED thread 'tests::daily_and_composer_preserve_unrelated_bytes' panicked at crates/plugins/notes/src/lib.rs:9473:9: assertion `left == right` failed left: 404 right: 200 test result: FAILED. 197 passed; 1 failed; 1 ignored; 0 measured; 0 filtered out; finished in 6.25s ``` The other six failures from the last job run (busy host) did not occur again. They were host-load timeouts. **Classification:** the remaining failure is the #954 fixture collision. Run alone, it passes 3 out of 3 times (`test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 198 filtered out`). In parallel, another test holds the process-global `USER_LOCKS` entry for the shared fixture User ID, so the Journal ID repair skips and the read returns a false 404. The branch did not have the fix. Fix: cherry-picked ff8e857c2 from `job/datafix2` (each fixture Home gets its own User ID). No assertion changed. The voice.rs, bookmarks.rs and tasks_dav.rs tests that use the same literal ID only take the blocking `lock_user`, which serializes them and cannot cause the skip, so they did not need a change. **After the fix**, two full runs with `--no-fail-fast -- --test-threads=4`: ``` run2: test result: ok. 198 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 7.40s apple_replay: test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.07s run3: test result: ok. 198 passed; 0 failed; 1 ignored; 0 measured; 0 filtered out; finished in 6.24s apple_replay: test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.05s ``` `cargo fmt --check -p calternal-plugin-notes`: exit 0, no output. `cargo clippy -p calternal-plugin-notes --all-targets -- -D warnings`: exit 0, `Finished \`dev\` profile [unoptimized + debuginfo] target(s) in 1m 20s`. Ready for merge from the Notes test side: yes. Not pushed.
Author
Owner

Round 7b source review for #867: artifacts/from-7b2/round3-adversarial.log:177 records a short invalid WebM transcription still in processing after the probe's 30-second deadline. No server stage or lifecycle log was supplied. crates/plugins/notes/src/voice.rs sets a 5-minute engine-slot wait (line 65), a 15-minute media decode deadline (line 2211), and a 30-minute model-download wait (line 66). The decoder runs before model verification, but it first acquires the model semaphore. Thus the existing bounds do not establish the probe's 30-second requirement. Production 6074f71d1 has no Voice route.

Unresolved: identify whether the supplied fixture waited on capacity, decoding, or model work. Add a defensive regression that invalid recordings reach a terminal failure before model work, while valid large recordings keep their documented duration and resource limits. Do not clear this finding based on the source deadline alone. No live hostile-input reproduction ran in 7bfix-data.

Round 7b source review for #867: `artifacts/from-7b2/round3-adversarial.log:177` records a short invalid WebM transcription still in `processing` after the probe's 30-second deadline. No server stage or lifecycle log was supplied. `crates/plugins/notes/src/voice.rs` sets a 5-minute engine-slot wait (line 65), a 15-minute media decode deadline (line 2211), and a 30-minute model-download wait (line 66). The decoder runs before model verification, but it first acquires the model semaphore. Thus the existing bounds do not establish the probe's 30-second requirement. Production `6074f71d1` has no Voice route. Unresolved: identify whether the supplied fixture waited on capacity, decoding, or model work. Add a defensive regression that invalid recordings reach a terminal failure before model work, while valid large recordings keep their documented duration and resource limits. Do not clear this finding based on the source deadline alone. No live hostile-input reproduction ran in 7bfix-data.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#619
No description provided.