RESEARCH: local voice transcription and rewrite (Superwhisper-style) — on-device/on-server, no cloud (NO implementation) #304

Closed
opened 2026-09-28 07:05:51 +00:00 by kayg · 10 comments
Owner

Owner (2026-09-28): 'file another issue for local transcription / modification of voice notes into text notes (like superwhisper) but we need to run it locally! have a codex research the best way to do this whenever capacity is available.'

Context: calternal records voice notes in the composer and (soon) in the Calendar item preview (the attach/record issue). Recordings are attachments in Home. The owner wants them turned into text notes: a transcription, plus an optional 'modification' step (clean up filler words, punctuate, restructure into a note, apply a style), like Superwhisper's modes. It must run locally: on the calternal server (a Rust server in a rootless podman container; the production VM has 4 vCPU and 7.7 GiB RAM with no GPU) and/or on the user's device (browser WebGPU/WASM; a future native macOS app). Never a cloud API by default (DESIGN business model: local features are always included; model-backed features use the customer's own API key). Licence rule: AGPL-3.0-only project, so the dependencies must be AGPL-compatible and the model weights must allow commercial use (no non-commercial weights).

Research (web search, benchmarks, licences):

  • Speech-to-text: whisper.cpp (and whisper-rs), faster-whisper/CTranslate2, Distil-Whisper, Whisper large-v3-turbo, NVIDIA Parakeet/Canary (check the licences), Moonshine, Kyutai STT, Vosk, sherpa-onnx, transformers.js/WebGPU in the browser, Apple SpeechAnalyzer (macOS 26) for the future native app. For each: accuracy (WER on common benchmarks, English plus the owner's likely languages: English, Hindi and other Indian languages; code-switching), speed on a 4-vCPU CPU (real-time factor), RAM, model size, streaming support, timestamps, diarization, licence of code and weights.
  • 'Modification' step: small local LLMs (Qwen, Llama, Gemma, Phi, Mistral: check the licences and commercial use) through llama.cpp/candle/ort; the CPU cost on the VM; the quality of cleanup; how Superwhisper, MacWhisper, Wispr Flow, Aqua, VoiceInk and Apple's features structure their 'modes'. Also the option to use the customer's own API key when they choose.
  • Where it runs: a server-side job queue (background, bounded concurrency, per-user quotas; must not hurt search latency, #258) vs in-browser vs the native app; a hybrid.
  • Product: auto-transcribe on save vs on demand; where the text goes (the log entry's linked note, a new note, inline under the recording); search indexing of transcripts (#258 targets); editing the transcript; privacy.
    Output: docs/research/voice-transcription.md (ASD-STE100) with a comparison table, a recommendation, a sizing estimate for the production VM, and numbered grill questions (grilling format). Do NOT implement. Commit on the branch, push, and post a summary on this issue.
Owner (2026-09-28): 'file another issue for local transcription / modification of voice notes into text notes (like superwhisper) but we need to run it locally! have a codex research the best way to do this whenever capacity is available.' Context: calternal records voice notes in the composer and (soon) in the Calendar item preview (the attach/record issue). Recordings are attachments in Home. The owner wants them turned into text notes: a transcription, plus an optional 'modification' step (clean up filler words, punctuate, restructure into a note, apply a style), like Superwhisper's modes. It must run **locally**: on the calternal server (a Rust server in a rootless podman container; the production VM has 4 vCPU and 7.7 GiB RAM with no GPU) and/or on the user's device (browser WebGPU/WASM; a future native macOS app). Never a cloud API by default (DESIGN business model: local features are always included; model-backed features use the customer's own API key). Licence rule: AGPL-3.0-only project, so the dependencies must be AGPL-compatible and the model weights must allow commercial use (no non-commercial weights). Research (web search, benchmarks, licences): - Speech-to-text: whisper.cpp (and whisper-rs), faster-whisper/CTranslate2, Distil-Whisper, Whisper large-v3-turbo, NVIDIA Parakeet/Canary (check the licences), Moonshine, Kyutai STT, Vosk, sherpa-onnx, transformers.js/WebGPU in the browser, Apple SpeechAnalyzer (macOS 26) for the future native app. For each: accuracy (WER on common benchmarks, English plus the owner's likely languages: English, Hindi and other Indian languages; code-switching), speed on a 4-vCPU CPU (real-time factor), RAM, model size, streaming support, timestamps, diarization, licence of code and weights. - 'Modification' step: small local LLMs (Qwen, Llama, Gemma, Phi, Mistral: check the licences and commercial use) through llama.cpp/candle/ort; the CPU cost on the VM; the quality of cleanup; how Superwhisper, MacWhisper, Wispr Flow, Aqua, VoiceInk and Apple's features structure their 'modes'. Also the option to use the customer's own API key when they choose. - Where it runs: a server-side job queue (background, bounded concurrency, per-user quotas; must not hurt search latency, #258) vs in-browser vs the native app; a hybrid. - Product: auto-transcribe on save vs on demand; where the text goes (the log entry's linked note, a new note, inline under the recording); search indexing of transcripts (#258 targets); editing the transcript; privacy. Output: docs/research/voice-transcription.md (ASD-STE100) with a comparison table, a recommendation, a sizing estimate for the production VM, and numbered grill questions (grilling format). Do NOT implement. Commit on the branch, push, and post a summary on this issue.
Author
Owner

Starting research on job/voice-research from dev at c9ec6aff84. Scope is research only; I will not run Cargo or Bun builds. I will review the voice attachment flow, search latency constraints, queue behavior, model and code licenses, current benchmark evidence, and local runtime options before writing the recommendation.

Starting research on job/voice-research from dev at c9ec6aff84a9ec810124a2cfbdd07c189eea1b9c. Scope is research only; I will not run Cargo or Bun builds. I will review the voice attachment flow, search latency constraints, queue behavior, model and code licenses, current benchmark evidence, and local runtime options before writing the recommendation.
Author
Owner

Finding: the composer already records voice memos in the browser and saves them under Documents/Voice memos/. The recorder accepts WebM/Opus, MP4/AAC, or Ogg/Opus and caps a recording at 30 minutes (apps/web/src/lib/composer/recorder.svelte.ts; crates/calternal-notes-core/src/attachments.rs). I found no ASR route or audio-transcript indexing path. The existing SQLite Worker supports a separate per-kind concurrency cap and retry/dedup behavior, but the queue has no per-User CPU quota. Search is a hard constraint: the 2026-09-28 100k profile on the shared 8-vCPU dev host reports warm hybrid p95 2,666.51 ms and marks the 150 ms target not met; it did not measure the 4-vCPU production VM (docs/perf/2026-09-28.md). I will recommend an explicit, bounded background job and require a production-VM profile before setting any throughput promise.

Finding: the composer already records voice memos in the browser and saves them under Documents/Voice memos/. The recorder accepts WebM/Opus, MP4/AAC, or Ogg/Opus and caps a recording at 30 minutes (apps/web/src/lib/composer/recorder.svelte.ts; crates/calternal-notes-core/src/attachments.rs). I found no ASR route or audio-transcript indexing path. The existing SQLite Worker supports a separate per-kind concurrency cap and retry/dedup behavior, but the queue has no per-User CPU quota. Search is a hard constraint: the 2026-09-28 100k profile on the shared 8-vCPU dev host reports warm hybrid p95 2,666.51 ms and marks the 150 ms target not met; it did not measure the 4-vCPU production VM (docs/perf/2026-09-28.md). I will recommend an explicit, bounded background job and require a production-VM profile before setting any throughput promise.
Author
Owner

Research is complete. The report is in docs/research/voice-transcription.md.

Recommendation: start with on-demand, local server transcription as a bounded background job. Pilot Qwen3-ASR-0.6B INT8 through sherpa-onnx; compare it with Vosk Hindi large and Whisper large-v3-turbo on the same consented samples. Save the raw result separately from cleanup. Let the User explicitly create a Note linked to the voice memo by stable identity.

Evidence and limits: the current recorder caps voice memos at 30 minutes. The Worker concurrency cap is per process and does not give per-User CPU quotas. The current Search profile misses its hybrid p95 target on the shared 8-vCPU host; it did not measure the 4-vCPU production VM. No controlled RTF or peak-RSS result was found for candidate models on that VM. The report includes language coverage, license and model caveats, product patterns, sizing estimates, a profiling plan, and 11 owner questions.

Decisions recommended for owner confirmation: required languages and Hindi script; on-demand versus opt-in automatic transcription; transcript persistence as Derived data or an explicit Note; live text and diarization scope; cleanup modes and cloud opt-in; retention and per-User resource limits. No product design was treated as settled where the design doc is silent.

Branch: job/voice-research. Base SHA: c9ec6aff84a9ec810124a2cfbdd07c189eea1b9c. Head SHA: 6924fba9c879a3d421b45dd41a54b1c659eb8216. The branch was pushed to origin.

Gate output:

  • git diff --cached --check: no output; exit 0.
  • Cargo/Bun build and test gates: Not run (research-only issue scope; builds prohibited; no code changed).
  • cargo clean: Removed 4 files, 3.3KiB total.

Known gap: no model or production-VM benchmark was run. This issue expressly prohibits builds, and the host is loaded.

Research is complete. The report is in `docs/research/voice-transcription.md`. Recommendation: start with on-demand, local server transcription as a bounded background job. Pilot Qwen3-ASR-0.6B INT8 through `sherpa-onnx`; compare it with Vosk Hindi large and Whisper large-v3-turbo on the same consented samples. Save the raw result separately from cleanup. Let the User explicitly create a Note linked to the voice memo by stable identity. Evidence and limits: the current recorder caps voice memos at 30 minutes. The Worker concurrency cap is per process and does not give per-User CPU quotas. The current Search profile misses its hybrid p95 target on the shared 8-vCPU host; it did not measure the 4-vCPU production VM. No controlled RTF or peak-RSS result was found for candidate models on that VM. The report includes language coverage, license and model caveats, product patterns, sizing estimates, a profiling plan, and 11 owner questions. Decisions recommended for owner confirmation: required languages and Hindi script; on-demand versus opt-in automatic transcription; transcript persistence as Derived data or an explicit Note; live text and diarization scope; cleanup modes and cloud opt-in; retention and per-User resource limits. No product design was treated as settled where the design doc is silent. Branch: `job/voice-research`. Base SHA: `c9ec6aff84a9ec810124a2cfbdd07c189eea1b9c`. Head SHA: `6924fba9c879a3d421b45dd41a54b1c659eb8216`. The branch was pushed to `origin`. Gate output: - `git diff --cached --check`: no output; exit 0. - Cargo/Bun build and test gates: `Not run (research-only issue scope; builds prohibited; no code changed).` - `cargo clean`: `Removed 4 files, 3.3KiB total`. Known gap: no model or production-VM benchmark was run. This issue expressly prohibits builds, and the host is loaded.
Author
Owner

Owner decisions (2026-09-28) on docs/research/voice-transcription.md:

  • V1 English only at launch. The owner asks: 'are you sure Qwen is the correct approach there? not Parakeet or S1-mini by Superwhisper?' The earlier pick (Qwen3-ASR-0.6B) was driven by Hindi, which is no longer needed. Follow-up research required (below).
  • V2 not needed (English only).
  • V3 B: automatic transcription after every voice note, on by default, as a low-priority background job. Plus a new Settings → Admin → Jobs page (like Immich) that shows what is running and lets an admin pause, resume or stop queues and jobs (filed separately).
  • V4 live text is not needed for now. The owner can increase the VM core count, so research how speed scales with cores (4 / 8 / 16 vCPU) and recommend a size.
  • V5 where the transcript goes depends on where the recording is added:
    • to an event / log entry → a linked note (the sub-bullet rule: a nested child bullet under the entry, with the note title as the link text) that holds Title + the voice recording + the transcription;
    • to a task → the task note;
    • to a note → that note itself (the recording and its transcription go into it).
      All of them appear on the Calendar.
  • V6 local only by default. Owner: 'voice is one area which we can go local.' (A cloud key stays out of scope for voice.)

Follow-up research (English only): compare NVIDIA Parakeet TDT 0.6B v2 (English) and v3, Whisper large-v3-turbo (whisper.cpp), Distil-Whisper, Moonshine, Kyutai STT, Qwen3-ASR-0.6B, and Superwhisper's own models (the owner named 'S1-mini'; verify that it exists, what it is, and whether it can be self-hosted or licensed at all; do not guess). For each: English WER (Open ASR Leaderboard and other sources), real-time factor on CPU at 4, 8 and 16 vCPU (published numbers first; say clearly which numbers are measured and which are estimated), RAM, a Rust path (sherpa-onnx, ort, whisper-rs, candle), punctuation and casing out of the box, long-form (30 min) handling, and the licence of the weights for commercial use (CC-BY-4.0 attribution rules for Parakeet, etc.). Recommend one model plus a VM size. Also propose a small measurement protocol (fixed English clips, interleaved runs) that a later measurement job can run on a quiet host.

Owner decisions (2026-09-28) on docs/research/voice-transcription.md: - V1 **English only** at launch. The owner asks: 'are you sure Qwen is the correct approach there? not Parakeet or S1-mini by Superwhisper?' The earlier pick (Qwen3-ASR-0.6B) was driven by Hindi, which is no longer needed. **Follow-up research required** (below). - V2 not needed (English only). - V3 **B: automatic transcription after every voice note, on by default**, as a low-priority background job. Plus a new **Settings → Admin → Jobs** page (like Immich) that shows what is running and lets an admin pause, resume or stop queues and jobs (filed separately). - V4 live text is not needed for now. The owner **can increase the VM core count**, so research how speed scales with cores (4 / 8 / 16 vCPU) and recommend a size. - V5 **where the transcript goes depends on where the recording is added**: - to an **event / log entry** → a **linked note** (the sub-bullet rule: a nested child bullet under the entry, with the note title as the link text) that holds **Title + the voice recording + the transcription**; - to a **task** → the **task note**; - to a **note** → that **note itself** (the recording and its transcription go into it). All of them appear on the Calendar. - V6 **local only by default.** Owner: 'voice is one area which we can go local.' (A cloud key stays out of scope for voice.) **Follow-up research (English only):** compare NVIDIA **Parakeet TDT 0.6B v2** (English) and **v3**, **Whisper large-v3-turbo** (whisper.cpp), **Distil-Whisper**, **Moonshine**, **Kyutai STT**, Qwen3-ASR-0.6B, and Superwhisper's own models (the owner named '**S1-mini**'; verify that it exists, what it is, and whether it can be self-hosted or licensed at all; do not guess). For each: English WER (Open ASR Leaderboard and other sources), real-time factor on CPU at 4, 8 and 16 vCPU (published numbers first; say clearly which numbers are measured and which are estimated), RAM, a Rust path (sherpa-onnx, ort, whisper-rs, candle), punctuation and casing out of the box, long-form (30 min) handling, and the **licence of the weights for commercial use** (CC-BY-4.0 attribution rules for Parakeet, etc.). Recommend one model plus a VM size. Also propose a small measurement protocol (fixed English clips, interleaved runs) that a later measurement job can run on a quiet host.
Author
Owner

Starting the English-only follow-up research on job/voice-models, based on dev at c99fc31dd9. I will verify the Parakeet, Superwhisper S1-mini, and Qwen comparison from primary sources, update docs/research/voice-transcription.md, and run no builds.

Starting the English-only follow-up research on job/voice-models, based on dev at c99fc31dd9e0f542a02956a787010bbf1250384c. I will verify the Parakeet, Superwhisper S1-mini, and Qwen comparison from primary sources, update docs/research/voice-transcription.md, and run no builds.
Author
Owner

Owner link: https://huggingface.co/superwhisper/s1-mini. Orchestrator check of the model card (2026-09-28): S1-mini is not an ASR model. It is Superwhisper's post-ASR text normalizer: a 596M-parameter causal LM fine-tuned from Qwen3-0.6B, English only (v1). It removes fillers, resolves self-corrections, applies punctuation and casing, and converts spoken forms (numbers, dates, currency, emails) to written form, with style control lines (casual, semi-casual, semi-formal, formal). Formats: BF16 safetensors plus GGUF quantizations (Q4_K_M = 462 MB) for llama.cpp, Ollama and vLLM; it runs on CPU. The card reports 94.8% token accuracy on 7,519 held-out English cases (Q4_K_M, greedy). Licence: 'Apache 2.0 + naming clause'; the name 'S1-mini' by 'Superwhisper' must be kept with exact capitalization wherever it is used.
Implication: the pipeline is ASR (Parakeet/Whisper, to be chosen) → S1-mini cleanup, both local. S1-mini is the candidate for the 'Clean up' step that the first research left unpromised. The follow-up must: evaluate S1-mini as the cleanup stage (CPU cost per minute of audio at 4/8/16 vCPU via llama.cpp and a Rust binding; memory next to the ASR model; the naming-clause obligations in the UI and the NOTICE file); decide whether the raw transcript is kept next to the cleaned one (proposal: keep both, show the cleaned text, 'Show raw' toggle); and map the style control lines to a Settings → Voice choice.

Owner link: https://huggingface.co/superwhisper/s1-mini. Orchestrator check of the model card (2026-09-28): **S1-mini is not an ASR model.** It is Superwhisper's **post-ASR text normalizer**: a 596M-parameter causal LM fine-tuned from Qwen3-0.6B, English only (v1). It removes fillers, resolves self-corrections, applies punctuation and casing, and converts spoken forms (numbers, dates, currency, emails) to written form, with style control lines (casual, semi-casual, semi-formal, formal). Formats: BF16 safetensors plus GGUF quantizations (Q4_K_M = 462 MB) for llama.cpp, Ollama and vLLM; it runs on CPU. The card reports 94.8% token accuracy on 7,519 held-out English cases (Q4_K_M, greedy). Licence: **'Apache 2.0 + naming clause'**; the name 'S1-mini' by 'Superwhisper' must be kept with exact capitalization wherever it is used. Implication: the pipeline is **ASR (Parakeet/Whisper, to be chosen) → S1-mini cleanup**, both local. S1-mini is the candidate for the 'Clean up' step that the first research left unpromised. The follow-up must: evaluate S1-mini as the cleanup stage (CPU cost per minute of audio at 4/8/16 vCPU via llama.cpp and a Rust binding; memory next to the ASR model; the naming-clause obligations in the UI and the NOTICE file); decide whether the raw transcript is kept next to the cleaned one (proposal: keep both, show the cleaned text, 'Show raw' toggle); and map the style control lines to a Settings → Voice choice.
Author
Owner

Follow-up research complete

Updated docs/research/voice-transcription.md with the English-only pipeline, model comparison, 4/8/16-vCPU guidance, measurement protocol, and owner decisions. No code changes or builds were made.

Recommendation

  • Use NVIDIA Parakeet TDT 0.6B v2 as the first ASR candidate. Its model card reports 6.05 mean WER on the Open ASR Leaderboard, built-in punctuation/capitalization/word times, and an official INT8 ONNX path with a Rust example through sherpa-onnx.
  • Run S1-mini by Superwhisper as a separate local text-cleanup stage after ASR. Its 94.8% token accuracy is from 7,519 held-out cleanup cases and is not ASR WER or calternal quality evidence. Keep raw and cleaned text. The research recommends cleanup on by default with a Settings → Voice opt-out, cleaned text shown first, and a Show raw control.
  • Use 4 vCPU / 7.7 GiB for a one-job pilot. Start production at 8 vCPU / 16 GiB with one low-priority active job. Use 16 vCPU / 32 GiB only if an 8-vCPU measurement shows a queue or latency need.

There is no published controlled 4/8/16-vCPU result for Parakeet v2 or S1-mini. The available CPU data is a third-party Parakeet v3 INT8 profile using another runtime on an 8-vCPU host: 20.3× RTF at 4 threads and 26.3× at 8 threads on a 30-second clip, with about 1.5 GB peak RSS. It is not a calternal VM result and does not establish 30-minute or linear scaling performance. The note records this limitation and a same-host measurement protocol.

Owner decisions recorded

English only; automatic transcription after a memo is saved; no live transcription; local-only processing; transcript destination follows the Event/Log entry, Task, or Note; 4/8/16-vCPU comparison; Admin Jobs UI is a separate issue.

The research recommendations still needing owner confirmation are 8 vCPU / 16 GiB as the initial production size, S1-mini cleanup by default, its style defaults, five-minute ASR chunks, sequential model loading, and a license review before redistributing S1-mini. Its naming and NOTICE requirements are called out in the note.

Commit and gates

Branch: job/voice-models
Head: 48a96f3e410ac5f567d8bbadec23d0ca21fd7afa (pushed)

Gate output:

$ git diff --check

(exit code 0; no stdout or stderr)

Cargo and web build/test gates were not run because this issue is research-only and prohibits builds. cargo clean output:

Removed 2 files, 636B total

Local dev was already at the branch base; git merge dev output was Already up to date.

## Follow-up research complete Updated [docs/research/voice-transcription.md](https://git.kayg.org/kayg/calternal/src/branch/job/voice-models/docs/research/voice-transcription.md) with the English-only pipeline, model comparison, 4/8/16-vCPU guidance, measurement protocol, and owner decisions. No code changes or builds were made. ### Recommendation - Use NVIDIA Parakeet TDT 0.6B v2 as the first ASR candidate. Its model card reports 6.05 mean WER on the Open ASR Leaderboard, built-in punctuation/capitalization/word times, and an official INT8 ONNX path with a Rust example through `sherpa-onnx`. - Run S1-mini by Superwhisper as a separate local text-cleanup stage after ASR. Its 94.8% token accuracy is from 7,519 held-out cleanup cases and is not ASR WER or calternal quality evidence. Keep raw and cleaned text. The research recommends cleanup on by default with a Settings → Voice opt-out, cleaned text shown first, and a Show raw control. - Use 4 vCPU / 7.7 GiB for a one-job pilot. Start production at 8 vCPU / 16 GiB with one low-priority active job. Use 16 vCPU / 32 GiB only if an 8-vCPU measurement shows a queue or latency need. There is no published controlled 4/8/16-vCPU result for Parakeet v2 or S1-mini. The available CPU data is a third-party Parakeet v3 INT8 profile using another runtime on an 8-vCPU host: 20.3× RTF at 4 threads and 26.3× at 8 threads on a 30-second clip, with about 1.5 GB peak RSS. It is not a calternal VM result and does not establish 30-minute or linear scaling performance. The note records this limitation and a same-host measurement protocol. ### Owner decisions recorded English only; automatic transcription after a memo is saved; no live transcription; local-only processing; transcript destination follows the Event/Log entry, Task, or Note; 4/8/16-vCPU comparison; Admin Jobs UI is a separate issue. The research recommendations still needing owner confirmation are 8 vCPU / 16 GiB as the initial production size, S1-mini cleanup by default, its style defaults, five-minute ASR chunks, sequential model loading, and a license review before redistributing S1-mini. Its naming and NOTICE requirements are called out in the note. ### Commit and gates Branch: `job/voice-models` Head: `48a96f3e410ac5f567d8bbadec23d0ca21fd7afa` (pushed) Gate output: ```text $ git diff --check (exit code 0; no stdout or stderr) ``` Cargo and web build/test gates were not run because this issue is research-only and prohibits builds. `cargo clean` output: ```text Removed 2 files, 636B total ``` Local `dev` was already at the branch base; `git merge dev` output was `Already up to date.`
Author
Owner

Owner decisions (2026-09-28) on the follow-up research (merged 82843b08):

  • Pipeline accepted: Parakeet TDT 0.6B v2 (sherpa-onnx INT8, CC-BY-4.0 attribution) → S1-mini by Superwhisper cleanup on the text (cleaned text by default, a 'Show raw' toggle; Settings → Voice: style / structure / context, defaults semi-formal / prose / general).
  • V7: accept S1-mini's licence (Apache-2.0 + naming clause). Name it exactly 'S1-mini by Superwhisper' in Settings → Voice and in the credits/NOTICE. Do not bundle the weights in the AGPL release.
  • Model downloads: on by default, never blocking. 'it should download by default and not block deployment or server start or ui interaction.' The server starts and serves normally; a background job (visible in Settings → Admin → Jobs, #315) downloads both models (pinned revision + SHA-256, like calternal-embed's manifest) after start-up, resumable, with retry/backoff. Until they are ready, recordings save as normal and their transcription waits in the queue with the state 'Waiting for the voice model (downloading 42%)'. Offline or blocked networks: transcription stays pending with a clear message; nothing else is affected.
  • Recording UI: the owner points to https://sveltebits.xyz/micro/voice-pill as the recording animation to reuse and refine with calternal's aesthetics (cherry-pick per CLAUDE.md, motion tokens, reduced motion, the #291 spring language). It applies to the shared recorder bar (composer + #303 item preview).
  • VM: 8 vCPU / 16 GiB is the recommended size for auto-transcription; a measurement job confirms it before any promise.
Owner decisions (2026-09-28) on the follow-up research (merged 82843b08): - **Pipeline accepted:** Parakeet TDT 0.6B v2 (sherpa-onnx INT8, CC-BY-4.0 attribution) → **S1-mini by Superwhisper** cleanup on the text (cleaned text by default, a 'Show raw' toggle; Settings → Voice: style / structure / context, defaults semi-formal / prose / general). - V7: **accept S1-mini's licence** (Apache-2.0 + naming clause). Name it exactly 'S1-mini by Superwhisper' in Settings → Voice and in the credits/NOTICE. **Do not bundle the weights** in the AGPL release. - **Model downloads: on by default, never blocking.** 'it should download by default and not block deployment or server start or ui interaction.' The server starts and serves normally; a background job (visible in Settings → Admin → Jobs, #315) downloads both models (pinned revision + SHA-256, like calternal-embed's manifest) after start-up, resumable, with retry/backoff. Until they are ready, recordings save as normal and their transcription waits in the queue with the state 'Waiting for the voice model (downloading 42%)'. Offline or blocked networks: transcription stays pending with a clear message; nothing else is affected. - **Recording UI:** the owner points to https://sveltebits.xyz/micro/voice-pill as the recording animation to reuse and refine with calternal's aesthetics (cherry-pick per CLAUDE.md, motion tokens, reduced motion, the #291 spring language). It applies to the shared recorder bar (composer + #303 item preview). - VM: 8 vCPU / 16 GiB is the recommended size for auto-transcription; a measurement job confirms it before any promise.
Author
Owner

Owner refinement (2026-09-28) on the recording UI: the mic icon expands inline into the animated recording pill. The mic button itself morphs (width morph with the #236/#291 spring, as the tab bar does) into the recording state in place: the waveform/level animation, the elapsed time, Stop/Keep/Discard. No overlay, no sheet, no separate bar, no extra chrome. Owner: 'no over-the-top crap. our design looks great as is.' Take only the voice-pill's core motion idea from sveltebits; keep calternal's existing pill look (the same PillGroup/glass, tokens, icon sizes). Collapse back into the mic on Stop/Keep/Discard. Reduced motion: an instant swap. The same component in the composer tools pill and the Calendar item-preview pill.

Owner refinement (2026-09-28) on the recording UI: **the mic icon expands inline into the animated recording pill**. The mic button itself morphs (width morph with the #236/#291 spring, as the tab bar does) into the recording state in place: the waveform/level animation, the elapsed time, Stop/Keep/Discard. No overlay, no sheet, no separate bar, no extra chrome. Owner: 'no over-the-top crap. our design looks great as is.' Take only the voice-pill's core motion idea from sveltebits; keep calternal's existing pill look (the same PillGroup/glass, tokens, icon sizes). Collapse back into the mic on Stop/Keep/Discard. Reduced motion: an instant swap. The same component in the composer tools pill and the Calendar item-preview pill.
Author
Owner

Research complete in e3dffb544 and 82843b08e (origin/dev); the model comparison and follow-up evaluation are recorded in the merged reports.

Research complete in `e3dffb544` and `82843b08e` (origin/dev); the model comparison and follow-up evaluation are recorded in the merged reports.
kayg closed this issue 2026-10-03 11:55:38 +00:00
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#304
No description provided.