VOICE: assistant, composer notes, dictation + global shortcuts, capture anywhere, local background transcription (design) #488

Open
opened 2026-09-30 06:44:15 +00:00 by kayg · 2 comments
Owner

Voice everywhere (owner, 2026-09-30); extends #304

Decisions:

  • V1 Assistant (mic next to +): a live voice conversation. One utterance can hold several intents, split into actions (a Money transaction, Tasks, a Reminder, ...). ONE confirmation card before anything is applied. It uses the #484 parity registry as its tool set. The voice and model provider is open; see the research issue on Sign in with ChatGPT for the Realtime API.
  • V2 Composer voice note: transcribe, post-process, then attach a note holding the recording and the transcript (#304 decisions stand).
  • V3 Dictation in every input box, plus global shortcuts with two modes: (a) plain dictation (text at the cursor) and (b) dictation plus post-processing (cleanup, formatting, a custom prompt template).
  • V4 Capture from anywhere: Watch, Drafts, Shortcuts and the Action Button POST audio to a calternal endpoint (App Password scoped) and use the same pipeline.
  • V5 Background transcription of existing media: audio and video in Files, Photos and Mail attachments are transcribed by a low-priority background job so search can find spoken words. All local.
  • V6 Meeting notes: out of scope for now (needs native apps).
  • V7 Voice commands for navigation: no. Instead the web app must work with the OS accessibility tools (macOS and iOS Voice Control, VoiceOver): every control has an accessible name that matches its visible label, so "Click Settings" works. Audit this on the Mac VM.
  • Local only for transcription (#304). The server can get a bigger VM (owner). Whether the assistant's language model must also be local is open.
  • ASR model: measure Phonon-2 (FermionResearch, 164 MB, CC-BY-4.0, English, CPU runtime for x86 Linux) against Parakeet TDT 0.6B v2 (the current pick in docs/research/voice-transcription.md) on the perf VM before choosing.
## Voice everywhere (owner, 2026-09-30); extends #304 Decisions: - **V1 Assistant (mic next to +):** a live voice conversation. One utterance can hold several intents, split into actions (a Money transaction, Tasks, a Reminder, ...). ONE confirmation card before anything is applied. It uses the #484 parity registry as its tool set. The voice and model provider is open; see the research issue on Sign in with ChatGPT for the Realtime API. - **V2 Composer voice note:** transcribe, post-process, then attach a note holding the recording and the transcript (#304 decisions stand). - **V3 Dictation in every input box**, plus **global shortcuts** with two modes: (a) plain dictation (text at the cursor) and (b) dictation plus post-processing (cleanup, formatting, a custom prompt template). - **V4 Capture from anywhere:** Watch, Drafts, Shortcuts and the Action Button POST audio to a calternal endpoint (App Password scoped) and use the same pipeline. - **V5 Background transcription of existing media:** audio and video in Files, Photos and Mail attachments are transcribed by a low-priority background job so search can find spoken words. **All local.** - **V6 Meeting notes:** out of scope for now (needs native apps). - **V7 Voice commands for navigation:** no. Instead the web app must work with the OS accessibility tools (macOS and iOS Voice Control, VoiceOver): every control has an accessible name that matches its visible label, so "Click Settings" works. Audit this on the Mac VM. - **Local only for transcription** (#304). The server can get a bigger VM (owner). Whether the assistant's language model must also be local is open. - **ASR model:** measure Phonon-2 (FermionResearch, 164 MB, CC-BY-4.0, English, CPU runtime for x86 Linux) against Parakeet TDT 0.6B v2 (the current pick in docs/research/voice-transcription.md) on the perf VM before choosing.
Author
Owner

Owner decision (2026-09-30): the voice assistant (V1) uses GPT Live (OpenAI Realtime). Wait for the research job (#490) on Sign in with ChatGPT eligibility and client IDs before building.

**Owner decision (2026-09-30):** the voice assistant (V1) uses **GPT Live** (OpenAI Realtime). Wait for the research job (#490) on Sign in with ChatGPT eligibility and client IDs before building.
Author
Owner

ASR decision from #489 (report merged): keep Parakeet TDT 0.6B v2 — Phonon-2 missed the test-other WER limit and was worse on the voice set, RTF and RSS. BUT Parakeet aborted on a 60-minute file (ONNX allocation error): background transcription (V5) must chunk long audio (e.g. 30–60 s windows with overlap and stitching) and bound memory. Track that in the V5 build.

ASR decision from #489 (report merged): keep **Parakeet TDT 0.6B v2** — Phonon-2 missed the test-other WER limit and was worse on the voice set, RTF and RSS. BUT Parakeet aborted on a 60-minute file (ONNX allocation error): background transcription (V5) must chunk long audio (e.g. 30–60 s windows with overlap and stitching) and bound memory. Track that in the V5 build.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#488
No description provided.