Skip to content

feat(speech): speech to prompt (dictation) - #853

Open
n0tlu5 wants to merge 7 commits into
xintaofei:mainfrom
n0tlu5:feat/speech-input
Open

n0tlu5 wants to merge 7 commits into
xintaofei:mainfrom
n0tlu5:feat/speech-input

Conversation

@n0tlu5

@n0tlu5 n0tlu5 commented Sep 28, 2026

Copy link
Copy Markdown

Refs #844

First PR of the speech stack for #844; it targets main directly.

What

Dictation for the composer. A mic button records speech and inserts the text into the prompt, using either the browser's speech recognition or an OpenAI-compatible transcription endpoint. It also adds a Speech settings page (engine, language, endpoint and API key; the key never leaves the backend) and microphone entitlements for the desktop app. On plain-http web deployments, browser dictation is reported as unavailable instead of failing on the first click; the cloud engine keeps working.

Commits to review

  • feat(speech): add speech preferences and capability detection
  • feat(speech): add OpenAI-compatible cloud transcription backend
  • feat(desktop): allow microphone capture in app windows
  • feat(chat): add speech input hook with browser and cloud engines
  • feat(chat): add voice dictation button to the composer
  • feat(settings): add Speech settings page
  • fix(speech): report browser dictation unavailable on insecure origins

Checks

pnpm lint ., pnpm test, pnpm build; desktop cargo check, cargo clippy --all-targets --features test-utils -- -D warnings and cargo test --features test-utils; server cargo check, cargo test --lib and cargo clippy -D warnings; codeg-mcp cargo check and clippy. All pass at the tip of the stack, and the behaviour was exercised in web mode in Chromium. No new npm or cargo dependencies.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant