Skip to content

feat(speech): read agent replies aloud - #854

Open
n0tlu5 wants to merge 12 commits into
xintaofei:mainfrom
n0tlu5:feat/speech-read-aloud
Open

n0tlu5 wants to merge 12 commits into
xintaofei:mainfrom
n0tlu5:feat/speech-read-aloud

Conversation

@n0tlu5

@n0tlu5 n0tlu5 commented Sep 28, 2026

Copy link
Copy Markdown

Refs #844

Stacked on #853 (feat/speech-input). GitHub can't base a PR on a fork branch, so this targets main and the diff includes the earlier PRs. Please review only the commits below; I'll rebase once the earlier PR merges.

What

Read-aloud for agent replies: a speaker action on each reply, optional auto-read at turn end (active tab only), browser voices or an OpenAI-compatible speech endpoint, with code blocks and tables replaced by short spoken labels. Playback stops when you type, send, switch tabs or start dictation.

Commits to review

  • feat(speech): add cloud speech synthesis backend
  • feat(speech): extract speakable text from agent replies
  • feat(speech): add read-aloud player
  • feat(chat): add read-aloud action to agent replies
  • feat(chat): auto-read agent replies at turn end

Checks

pnpm lint ., pnpm test, pnpm build; desktop cargo check, cargo clippy --all-targets --features test-utils -- -D warnings and cargo test --features test-utils; server cargo check, cargo test --lib and cargo clippy -D warnings; codeg-mcp cargo check and clippy. All pass at the tip of the stack, and the behaviour was exercised in web mode in Chromium. No new npm or cargo dependencies.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant