1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
LLMRepos·Audio / Speech Models LLM projects
Updated dailyBrowse 72 open-source audio / speech models projects in Models. Compare GitHub stars, recent growth, languages, licenses, and repository activity.
Audio / Speech Models
Browse 72 open-source audio / speech models projects in Models. Compare GitHub stars, recent growth, languages, licenses, and repository activity.
Top Audio / Speech Models repositories
Ranked by current GitHub stars from the latest LLMRepos snapshot.
Showing 40 of 72
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
Instant voice cloning by MIT and MyShell. Audio foundation model.
🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
A TTS model capable of generating ultra-realistic dialogue in one pass.
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages
Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.
A fast, local neural text to speech system
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get started in the field of audio, music, and speech generation research and development.
mozilla/TTS
:robot: :speech_balloon: Deep learning for Text to Speech (Discussion forum: https://discourse.mozilla.org/c/tts)
End-to-End Speech Processing Toolkit
EmotiVoice 😊: a Multi-Voice and Prompt-Controlled TTS Engine
An open source implementation of Microsoft's VALL-E X zero-shot TTS model. Demo is available in https://plachtaa.github.io/vallex/
VITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
High-quality multi-lingual text-to-speech library by MyShell.ai. Support English, Spanish, French, Chinese, Japanese and Korean.
On-device Speech AI for Apple Silicon
StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models
Silero Models: pre-trained text-to-speech models made embarrassingly simple
DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism (SVS & TTS); AAAI 2022; Official code
Foundational model for human-like, expressive TTS
AutoArk/GPA
[AutoArk] GPA (General Purpose Audio) can do ASR, TTS and voice conversion with one tiny model!
🚀 一键部署(含离线整合包)!基于 ChatTTS ,支持流式输出、音色抽卡、长音频生成和分角色朗读。简单易用,无需复杂安装。
MARY TTS -- an open-source, multilingual text-to-speech synthesis system written in pure java
Vietnamese TTS with instant voice cloning • On-device • Real-time CPU inference • 24kHz audio quality • Chuyển văn bản thành giọng nói tiếng Việt • Text to speech tiếng Việt • TTS tiếng Việt
The collection of pre-trained, state-of-the-art AI models for ailia SDK
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
DeepMind's Tacotron-2 Tensorflow implementation
PyTorch implementation of VALL-E(Zero-Shot Text-To-Speech), Reproduced Demo https://lifeiteng.github.io/valle/index.html
Controllable and fast Text-to-Speech for over 7000 languages!
WaveRNN Vocoder + TTS
Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR benchmarks, while also offering outstanding singing lyrics recognition capability.
a free and open source speech synthesizer for Russian and other languages
GPT-SoVITS ONNX Inference Engine & Model Converter
Unofficial Parallel WaveGAN (+ MelGAN & Multi-band MelGAN & HiFi-GAN & StyleMelGAN) with Pytorch