Skip to content

feat(voice): hands-free voice mode for the workspace assistant - #856

Open
n0tlu5 wants to merge 21 commits into
xintaofei:mainfrom
n0tlu5:feat/voice-mode
Open

n0tlu5 wants to merge 21 commits into
xintaofei:mainfrom
n0tlu5:feat/voice-mode

Conversation

@n0tlu5

@n0tlu5 n0tlu5 commented Sep 28, 2026

Copy link
Copy Markdown

Refs #844

Stacked on #855 (feat/workspace-assistant). GitHub can't base a PR on a fork branch, so this targets main and the diff includes the earlier PRs. Please review only the commits below; I'll rebase once the earlier PR merges.

What

Hands-free voice mode on top of the assistant: voice activity detection (AudioWorklet, one AudioContext for the whole voice-mode lifetime, closed on exit, so #837 doesn't come back), replies spoken sentence by sentence while they stream, barge-in, fixed voice commands (stop, cancel, confirm, reject), spoken confirmations for the assistant's actions, optional announcements when a session elsewhere finishes or needs permission, a voice orb, a top-bar button and Mod+Shift+J.

Commits to review

  • feat(voice): add voice-mode audio front end and end-of-speech detection
  • feat(voice): speak replies sentence by sentence while they stream
  • feat(voice): add workspace voice mode bound to the assistant session
  • feat(voice): add fixed voice commands and spoken confirmations
  • feat(voice): announce session status across the workspace
  • feat(voice): add voice orb, entry points and assistant settings

Checks

pnpm lint ., pnpm test, pnpm build; desktop cargo check, cargo clippy --all-targets --features test-utils -- -D warnings and cargo test --features test-utils; server cargo check, cargo test --lib and cargo clippy -D warnings; codeg-mcp cargo check and clippy. All pass at the tip of the stack, and the behaviour was exercised in web mode in Chromium. No new npm or cargo dependencies.

n0tlu5 added 21 commits September 27, 2026 17:18
Chromium exposes SpeechRecognition on plain-http origins but every session fails at once with not-allowed, so the mic showed "voice input failed" immediately. Browser dictation is now only offered in a secure context; the cloud engine keeps working over http.

Refs xintaofei#844
Add a single persistent assistant conversation owned by the backend
rather than any tab: assistant_{get,set}_settings, assistant_ensure and
assistant_reset, exposed as Tauri commands and web POST routes.

ensure reuses the live assistant connection (remembered under the
ensure lock, since a fresh spawn is linked to its conversation only on
the first prompt), resumes the stored session when possible, and
spawns the codeg-mcp companion with the new `assistant` feature, which
also turns on `sessions`. Agents without companion support are refused.

Refs xintaofei#844
Wire send_to_session, cancel_session, answer_permission and start_session
behind a Confirm/Cancel card that codeg builds from session state and
registers on the requesting assistant's own connection. Each tool checks
the assistant settings gates, the target's run state and (for
permissions) the pending request before showing any card.

answer_permission maps approve/deny to the pending request's allow_once /
reject_once option id, never an allow_always / reject_always option,
returns unsupported when no one-shot option is offered, and re-checks the
request id after confirmation so a permission already answered on the
target's own card is not answered twice.

Permission prompts for codeg's own assistant tools are auto-allowed once
only on the assistant connection; the codeg card is the real gate.

Refs xintaofei#844
Adds the voice-mode building blocks: a voiceMode preferences section, a
clock-injected energy VAD with an adaptive noise floor and a raised
threshold during playback, a single-AudioContext microphone front end
(AudioWorklet, echo cancellation) that closes its context exactly once
and releases every track even when setup fails, and an utterance
recorder for the browser and cloud engines. The dictation engine
helpers move to src/lib/speech-engines.ts unchanged.

Refs xintaofei#844
Adds a sentence stream that splits live reply deltas at sentence ends
(including CJK punctuation) outside fenced code, speaks one code-omitted
label per fence and never cuts inside a number, plus a streaming queue in
the speech player: beginSpeechStream / enqueueSpeech / endSpeechStream
play segments in order (one-ahead prefetch on the cloud engine) and
onSpeechDrained fires once per finished stream. speak() and stopSpeech()
end a stream without a drained notification.

Refs xintaofei#844
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant