High-performance, low-latency Speech-to-Text for Discord voice channels. Process-isolated architecture ensures the bot never freezes during inference.
| Feature | Implementation |
|---|---|
| Multiprocessing Core | STT runs in an isolated process — bot never freezes |
| Anti-aliased Resampling | torchaudio Kaiser-window filter (48 kHz → 16 kHz) |
| Ring Buffer | 320 ms pre-speech context — first syllable is never cut |
| Silero VAD | State-of-the-art voice activity detection |
| Faster-Whisper | CTranslate2 — 4× faster than OpenAI Whisper |
| Structured Logging | Python logging with module-level loggers |
| Graceful Shutdown | Signal handlers for clean exit |
| Per-User State | Encapsulated UserState dataclass per speaker |
graph TD
subgraph "Main Process — Discord Bot"
A[Discord Gateway] -->|Opus Audio| B["AudioSink<br/>(bot/audio_sink.py)"]
B -->|PCM 48kHz Stereo| C["Resampler<br/>(audio/resampler.py)"]
C -->|PCM 16kHz Mono| D[IPC Audio Queue]
H[IPC Result Queue] -->|JSON| I["ResultHandler<br/>(bot/client.py)"]
end
subgraph "STT Process — Isolated"
D --> E["UserStateManager<br/>(stt/user_state.py)"]
E -->|32ms Frames| F["Silero VAD<br/>(stt/vad.py)"]
F -->|Speech Segments| G["Faster-Whisper<br/>(stt/transcriber.py)"]
G -->|Text| H
end
Discord-Realtime-STT-Bot/
├── main.py # Entry point (graceful shutdown)
├── config.py # Grouped configuration (dataclasses)
├── bot/
│ ├── client.py # Bot setup, commands, result handler
│ └── audio_sink.py # AudioSink + resampling bridge
├── audio/
│ ├── resampler.py # torchaudio anti-aliased resampling
│ └── ring_buffer.py # Generic ring buffer
├── stt/
│ ├── processor.py # STT process main loop
│ ├── vad.py # Silero VAD wrapper
│ ├── transcriber.py # Faster-Whisper wrapper
│ └── user_state.py # Per-user state dataclass + manager
├── utils/
│ └── logging.py # Structured logging config
├── deploy/
│ └── systemd/ # Ubuntu systemd service example
├── requirements.txt
├── .env.example # Environment template
├── .gitignore
└── README.md
- Python 3.10+
- FFmpeg and Opus runtime libraries
- NVIDIA GPU is recommended for low latency, but CPU fallback is supported
sudo apt update
sudo apt install -y python3 python3-venv python3-pip git rsync ffmpeg libopus0 build-essentialgit clone https://github.com/your-repo/discord-stt-bot.git
cd discord-stt-bot
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txtNote:
torch+faster-whispermay exceed 2 GB total.
Create a .env file from the checked-in template:
cp .env.example .env
nano .envAt minimum, set:
DISCORD_TOKEN=your_super_secret_token_hereFor CPU-only Ubuntu servers, start with:
STT_DEVICE=cpu
STT_COMPUTE_TYPE=int8
STT_MODEL_ID=baseFor CUDA deployments, install the PyTorch build that matches your driver/CUDA runtime before installing the rest of the requirements. See the official PyTorch install selector for the correct index URL.
- Enable the Message Content Intent for the bot.
- Invite the bot with permissions to read/send messages and connect/speak in voice channels.
source .venv/bin/activate
python main.pyRuntime settings are loaded from environment variables, usually via .env.
| Group | Setting | Default | Description |
|---|---|---|---|
| Discord | DISCORD_TOKEN |
required | Bot token |
COMMAND_PREFIX |
! |
Prefix command trigger | |
| STT | STT_MODEL_ID |
deepdml/faster-whisper-large-v3-turbo-ct2 |
HuggingFace model ID |
STT_DEVICE |
cuda |
Use cpu if no GPU |
|
STT_COMPUTE_TYPE |
float16 |
Use int8 for CPU |
|
STT_LANGUAGE |
ko |
Transcription language | |
STT_BEAM_SIZE |
1 |
Lower = faster, higher = more accurate | |
| Process | AUDIO_QUEUE_MAXSIZE |
512 |
Backpressure limit for incoming audio |
SPEECH_FLUSH_TIMEOUT_SECONDS |
2.0 |
Flush final speech if Discord stops sending frames | |
USER_TIMEOUT_SECONDS |
60 |
Inactive user cleanup | |
| Audio | input_sample_rate |
48000 |
Discord Opus decoded rate |
output_sample_rate |
16000 |
Whisper input rate |
The repository includes an example unit at deploy/systemd/discord-stt-bot.service.example.
Example layout:
sudo useradd --system --create-home --home-dir /opt/discord-stt-bot discord-stt
sudo rsync -a --exclude .git ./ /opt/discord-stt-bot/
sudo chown -R discord-stt:discord-stt /opt/discord-stt-bot
sudo cp /opt/discord-stt-bot/.env.example /etc/discord-stt-bot.env
sudo nano /etc/discord-stt-bot.env
sudo chmod 600 /etc/discord-stt-bot.env
sudo chown root:root /etc/discord-stt-bot.env
sudo cp /opt/discord-stt-bot/deploy/systemd/discord-stt-bot.service.example /etc/systemd/system/discord-stt-bot.service
sudo systemctl daemon-reload
sudo systemctl enable --now discord-stt-botCheck logs:
journalctl -u discord-stt-bot -fIf using an NVIDIA GPU, make sure the discord-stt user can access the GPU devices and that the installed torch wheel matches the host driver/CUDA runtime.
- Summon:
!joinin any text channel - Speak: Talk — the bot listens to everyone simultaneously
- Dismiss:
!leaveto disconnect - Stop:
Ctrl+Cfor graceful shutdown
Q: The bot joins but doesn't transcribe.
Check console logs. Ensure Silero VAD and Whisper model downloaded successfully.
Q: It's too slow!
Set
device = "cuda"inconfig.py. Runninglarge-v3on CPU is not recommended.
Q: CUDA out of memory?
Switch to a smaller model (
base,small) or usecompute_type = "int8".
Q: I see "Audio queue full" warnings.
STT inference is slower than the incoming voice stream. Use a smaller model, GPU acceleration, or increase
AUDIO_QUEUE_MAXSIZEif the host has enough memory.
Q: The service exits right after startup.
Check
DISCORD_TOKEN, Discord privileged intents, native voice dependencies, and model download/network access injournalctl -u discord-stt-bot -e.
MIT License — free to fork, modify, and use.