Implementation of WaveGrad high-fidelity vocoder from Google Brain in PyTorch.
-
Updated
Jul 7, 2021 - Jupyter Notebook
Implementation of WaveGrad high-fidelity vocoder from Google Brain in PyTorch.
Implementation of "Duration Informed Attention Network for Multimodal Synthesis" paper in PyTorch.
Yet another PyTorch implementation of Tacotron 2 with reduction factor and faster training speed.
Pytorch implementation of Tacotron, a speech synthesis end-to-end generative TTS model.
Extract a target speaker’s clean, non-overlapped speech from multi-speaker audio and export word-safe LJSpeech-style TTS datasets.
crawl4ai for video & audio — turn any YouTube video, podcast, or recording into clean timestamped markdown, or a verified TTS/STT training dataset. Runs locally.
Easy access to speech data across 142 African languages for training TTS and ASR models.
LibriVox dataset for Bulgarian language TTS
[Russian] This script will split audio file on silence, transcript it with google recognition and save it in LJSpeech-1.1 dataset manner.
🎙️ Mixer-TTS for efficient TTS ⚡
QuartzNet implementation for Automatic Speech Recognition task
WaveNet vocoder implementation for speech synthesis task
End-to-end speech-to-text model combining CNN feature extraction + Bidirectional LSTM sequence modeling + CTC loss. Trained on the LJSpeech dataset with a modular codebase for data loading, preprocessing, model building, and training.
🎨 Create stunning portrait animations with Durian, leveraging dual reference guidance for seamless attribute transfer and enhanced visual impact.
Tacotron 2 implementation for text2speech task
A command-line tool for transcribing audio files in a folder to a metadata.csv file, using OpenAI's Whisper.
To associate your repository with the ljspeech topic, visit your repo's landing page and select "manage topics."