RESEARCH: Sign in with ChatGPT for the voice assistant (Realtime eligibility, client ID for self-hosters) #490

Open
opened 2026-09-30 06:44:18 +00:00 by kayg · 8 comments
Owner

Research: can calternal use the User's ChatGPT subscription ("Sign in with ChatGPT") for the voice assistant, including live voice?

Owner question (2026-09-30): "can we use gpt live from the sign in with chatgpt sub?"
Find out, with primary sources (developers.openai.com/siwc, learn.chatgpt.com/docs/sign-in-with-chatgpt, the OpenAI API docs, the Codex CLI source that uses ChatGPT sign-in):

  1. Which endpoints and models a Sign in with ChatGPT token can call: Responses API only, or also the Realtime API (speech-to-speech), audio transcription, and TTS. Quote the docs; if they are unclear, say so.
  2. Whether an open-source, self-hosted AGPL app can get a client ID (redirect URIs per self-hosted instance? a shared client for the calternal project? each self-hoster brings their own?), the approval process, and the terms that affect us (per-app weekly caps and so on).
  3. If Realtime is not covered: the best alternative, for example the Responses API with streamed audio in and out, or local ASR (Phonon/Parakeet), then Responses for intent, then local TTS. Compare latency and privacy.
  4. How Claude (Anthropic) and Gemini subscriptions compare for the same use, since the owner wants provider choice.
    Deliver docs/research/voice-assistant-provider.md in ASD-STE100 style and a short answer on the issue. No code. Web search on.
## Research: can calternal use the User's ChatGPT subscription ("Sign in with ChatGPT") for the voice assistant, including live voice? Owner question (2026-09-30): "can we use gpt live from the sign in with chatgpt sub?" Find out, with primary sources (developers.openai.com/siwc, learn.chatgpt.com/docs/sign-in-with-chatgpt, the OpenAI API docs, the Codex CLI source that uses ChatGPT sign-in): 1. Which endpoints and models a Sign in with ChatGPT token can call: Responses API only, or also the Realtime API (speech-to-speech), audio transcription, and TTS. Quote the docs; if they are unclear, say so. 2. Whether an open-source, self-hosted AGPL app can get a client ID (redirect URIs per self-hosted instance? a shared client for the calternal project? each self-hoster brings their own?), the approval process, and the terms that affect us (per-app weekly caps and so on). 3. If Realtime is not covered: the best alternative, for example the Responses API with streamed audio in and out, or local ASR (Phonon/Parakeet), then Responses for intent, then local TTS. Compare latency and privacy. 4. How Claude (Anthropic) and Gemini subscriptions compare for the same use, since the owner wants provider choice. Deliver `docs/research/voice-assistant-provider.md` in ASD-STE100 style and a short answer on the issue. No code. Web search on.
Author
Owner

Starting research on branch job/siwc-490 from base 0dc772c3697ea9bd01822c26440c32206d472715. I will check the Sign in with ChatGPT API scope, client registration rules and caps, and compare provider subscription options using primary sources. No code or API behavior will change.

Starting research on branch `job/siwc-490` from base `0dc772c3697ea9bd01822c26440c32206d472715`. I will check the Sign in with ChatGPT API scope, client registration rules and caps, and compare provider subscription options using primary sources. No code or API behavior will change.
Author
Owner

Owner decision: the assistant WILL use GPT Live. Focus the research on HOW: whether a Sign in with ChatGPT token can open a Realtime session (and, if not, the closest path to GPT Live on the User's own subscription), the client-ID and redirect story for self-hosted instances, and the costs and limits. Deprioritise the Claude/Gemini comparison.

Owner decision: the assistant WILL use GPT Live. Focus the research on HOW: whether a Sign in with ChatGPT token can open a Realtime session (and, if not, the closest path to GPT Live on the User's own subscription), the client-ID and redirect story for self-hosted instances, and the costs and limits. Deprioritise the Claude/Gemini comparison.
Author
Owner

Evidence: OpenAI’s SIWC guide labels plan sharing as an open-source client flow, dynamically issues each ChatGPT account/workspace its client_id, requires a 127.0.0.1 callback, and sends plan requests to /v1/responses with store=false and stream=true. Its preview limits state that audio/video input and the transcription API are unsupported. ChatGPT Learn says the user sets each app’s weekly percentage cap and that the cap does not add usage. Google’s Gemini CLI terms state that using its OAuth service with third-party software violates its terms. I am checking provider voice and local ASR options before finalizing the note.

Evidence: OpenAI’s SIWC guide labels plan sharing as an open-source client flow, dynamically issues each ChatGPT account/workspace its client_id, requires a 127.0.0.1 callback, and sends plan requests to /v1/responses with store=false and stream=true. Its preview limits state that audio/video input and the transcription API are unsupported. ChatGPT Learn says the user sets each app’s weekly percentage cap and that the cap does not add usage. Google’s Gemini CLI terms state that using its OAuth service with third-party software violates its terms. I am checking provider voice and local ASR options before finalizing the note.
Author
Owner

Finding: docs/research/voice-transcription.md records the owner decision from #304: voice processing stays local, each voice memo is saved first and transcribed after save, and live text during recording is not needed. The #490 question asks whether ChatGPT plan access can provide live voice. The research note now treats that as a feasibility question and does not treat it as a change to #304. OpenAI SIWC can provide streamed Responses text, but its current limits exclude audio input and transcription; Realtime and TTS are not documented for this token.

Finding: docs/research/voice-transcription.md records the owner decision from #304: voice processing stays local, each voice memo is saved first and transcribed after save, and live text during recording is not needed. The #490 question asks whether ChatGPT plan access can provide live voice. The research note now treats that as a feasibility question and does not treat it as a change to #304. OpenAI SIWC can provide streamed Responses text, but its current limits exclude audio input and transcription; Realtime and TTS are not documented for this token.
Author
Owner

Workspace gate finding (2026-09-30): CARGO_PROFILE_DEV_DEBUG=line-tables-only CARGO_INCREMENTAL=0 CARGO_BUILD_JOBS=4 TMPDIR=<worktree>/target/tmp cargo clippy --workspace --all-targets -- -D warnings exited 101. The existing vendored crate crates/plugins/mail/vendor/async-imap fails because both runtime-async-std and runtime-tokio features are enabled. Rust reports duplicate Read/Write imports, a ReadBuf type mismatch, and the crate's own compile error requiring only one runtime feature. This job changes only docs/research/voice-assistant-provider.md and does not touch mail code.

Workspace gate finding (2026-09-30): `CARGO_PROFILE_DEV_DEBUG=line-tables-only CARGO_INCREMENTAL=0 CARGO_BUILD_JOBS=4 TMPDIR=<worktree>/target/tmp cargo clippy --workspace --all-targets -- -D warnings` exited 101. The existing vendored crate `crates/plugins/mail/vendor/async-imap` fails because both `runtime-async-std` and `runtime-tokio` features are enabled. Rust reports duplicate `Read`/`Write` imports, a `ReadBuf` type mismatch, and the crate's own compile error requiring only one runtime feature. This job changes only `docs/research/voice-assistant-provider.md` and does not touch mail code.
Author
Owner

Workspace test gate finding (2026-09-30): CARGO_PROFILE_DEV_DEBUG=line-tables-only CARGO_INCREMENTAL=0 CARGO_BUILD_JOBS=4 TMPDIR=<worktree>/target/tmp cargo test --workspace exited 101 before tests ran. It fails in the same existing crates/plugins/mail/vendor/async-imap crate because both async-std and Tokio runtime features are enabled. The compiler reports duplicate Read/Write imports and compile_error!("only one of 'runtime-async-std' or 'runtime-tokio' features must be enabled"). This job changes only research documentation.

Workspace test gate finding (2026-09-30): `CARGO_PROFILE_DEV_DEBUG=line-tables-only CARGO_INCREMENTAL=0 CARGO_BUILD_JOBS=4 TMPDIR=<worktree>/target/tmp cargo test --workspace` exited 101 before tests ran. It fails in the same existing `crates/plugins/mail/vendor/async-imap` crate because both async-std and Tokio runtime features are enabled. The compiler reports duplicate `Read`/`Write` imports and `compile_error!("only one of 'runtime-async-std' or 'runtime-tokio' features must be enabled")`. This job changes only research documentation.
Author
Owner

Completed on job/siwc-490.

Short answer: SIWC documents eligible streamed Responses requests for models returned by the signed-in account's /v1/models endpoint. Its preview limits explicitly exclude audio/video input and transcription. The SIWC docs do not establish Realtime, TTS, or audio output from Responses access. The open-source flow issues a client ID for a ChatGPT account/workspace through loopback OAuth; the User sets a weekly per-app cap. The docs do not promise approval for a specific AGPL app. Claude and Gemini do not document the same general third-party subscription flow; Gemini's terms prohibit using Gemini CLI OAuth with third-party tools. Keep issue #304's local voice and post-save transcription decision: SIWC still sends model input to OpenAI.

Research and primary sources: docs/research/voice-assistant-provider.md.

Head: 66aaf70022da8593812b2efe19cf62a8f9cf82f4 (docs: research subscription voice providers). bun.lock is unchanged. The worktree is clean. I merged origin/dev once before gates; the remote advanced again during the long checks, so I did not repeat the merge, as the job rules require.

Gate output:

  • cargo fmt --check: exit 0, no output.

  • cargo clippy --workspace --all-targets -- -D warnings (exit 101):

    error: only one of 'runtime-async-std' or 'runtime-tokio' features must be enabled
      --> crates/plugins/mail/vendor/async-imap/src/lib.rs:83:1
    error: could not compile `async-imap` (lib) due to 14 previous errors
    
  • cargo test --workspace (exit 101; tests did not start):

    error: only one of 'runtime-async-std' or 'runtime-tokio' features must be enabled
      --> crates/plugins/mail/vendor/async-imap/src/lib.rs:83:1
    error: could not compile `async-imap` (lib) due to 8 previous errors; 1 warning emitted
    
  • bun run check (exit 0):

    Text sizes and UI shape values use shared role tokens.
    Loading svelte-check in workspace: /home/kayg/Developer/calternal-wt/siwc-490/apps/web
    Getting Svelte diagnostics...
    svelte-check found 0 errors and 0 warnings
    
  • bun run test --maxWorkers=1 (exit 0):

     Test Files  137 passed (137)
          Tests  887 passed (887)
     Duration  488.42s (transform 30%, environment 25%, import 24%, tests 16%, setup 4%)
    
  • cargo clean (exit 0): Removed 13056 files, 5.0GiB total.

Known gaps and inferences: OpenAI does not clarify whether a browser-based self-hosted Instance fits “locally hosted”, or whether Responses can return audio. There is no calternal end-to-end latency benchmark. A stable ext_agent_host_id per Agent container and a local OAuth helper for remote Instances are recommendations inferred from the docs, not decisions in DESIGN. No code or UI changed.

Completed on `job/siwc-490`. **Short answer:** SIWC documents eligible streamed Responses requests for models returned by the signed-in account's `/v1/models` endpoint. Its preview limits explicitly exclude audio/video input and transcription. The SIWC docs do not establish Realtime, TTS, or audio output from Responses access. The open-source flow issues a client ID for a ChatGPT account/workspace through loopback OAuth; the User sets a weekly per-app cap. The docs do not promise approval for a specific AGPL app. Claude and Gemini do not document the same general third-party subscription flow; Gemini's terms prohibit using Gemini CLI OAuth with third-party tools. Keep issue #304's local voice and post-save transcription decision: SIWC still sends model input to OpenAI. Research and primary sources: `docs/research/voice-assistant-provider.md`. Head: `66aaf70022da8593812b2efe19cf62a8f9cf82f4` (`docs: research subscription voice providers`). `bun.lock` is unchanged. The worktree is clean. I merged `origin/dev` once before gates; the remote advanced again during the long checks, so I did not repeat the merge, as the job rules require. Gate output: - `cargo fmt --check`: exit 0, no output. - `cargo clippy --workspace --all-targets -- -D warnings` (exit 101): ``` error: only one of 'runtime-async-std' or 'runtime-tokio' features must be enabled --> crates/plugins/mail/vendor/async-imap/src/lib.rs:83:1 error: could not compile `async-imap` (lib) due to 14 previous errors ``` - `cargo test --workspace` (exit 101; tests did not start): ``` error: only one of 'runtime-async-std' or 'runtime-tokio' features must be enabled --> crates/plugins/mail/vendor/async-imap/src/lib.rs:83:1 error: could not compile `async-imap` (lib) due to 8 previous errors; 1 warning emitted ``` - `bun run check` (exit 0): ``` Text sizes and UI shape values use shared role tokens. Loading svelte-check in workspace: /home/kayg/Developer/calternal-wt/siwc-490/apps/web Getting Svelte diagnostics... svelte-check found 0 errors and 0 warnings ``` - `bun run test --maxWorkers=1` (exit 0): ``` Test Files 137 passed (137) Tests 887 passed (887) Duration 488.42s (transform 30%, environment 25%, import 24%, tests 16%, setup 4%) ``` - `cargo clean` (exit 0): `Removed 13056 files, 5.0GiB total`. Known gaps and inferences: OpenAI does not clarify whether a browser-based self-hosted Instance fits “locally hosted”, or whether Responses can return audio. There is no calternal end-to-end latency benchmark. A stable `ext_agent_host_id` per Agent container and a local OAuth helper for remote Instances are recommendations inferred from the docs, not decisions in DESIGN. No code or UI changed.
Author
Owner

Hygiene review: job/siwc-490 at 66aaf70022da8593812b2efe19cf62a8f9cf82f4 is not in origin/dev. The research artifact remains unmerged; keeping #490 open.

Hygiene review: `job/siwc-490` at `66aaf70022da8593812b2efe19cf62a8f9cf82f4` is not in `origin/dev`. The research artifact remains unmerged; keeping #490 open.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#490
No description provided.