1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
LLMRepos·Models LLM projects
Updated dailyBrowse 144 open-source models LLM projects across Audio / Speech Models, Vision & Multimodal Models, Foundation / LLM Models. Compare stars, growth, languages, licenses, and repository activity.
Subcategories
Top Models repositories
Ranked by current GitHub stars from the latest LLMRepos snapshot.
Showing 40 of 144
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
Instant voice cloning by MIT and MyShell. Audio foundation model.
🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
FinGPT: Open-Source Financial Large Language Models! Revolutionize 🔥 We release the trained model on HuggingFace.
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
A TTS model capable of generating ultra-realistic dialogue in one pass.
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages
Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.
Implementation of DALL-E 2, OpenAI's updated text-to-image synthesis neural network, in Pytorch
A fast, local neural text to speech system
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get started in the field of audio, music, and speech generation research and development.
mozilla/TTS
:robot: :speech_balloon: Deep learning for Text to Speech (Discussion forum: https://discourse.mozilla.org/c/tts)
End-to-End Speech Processing Toolkit
14MB foundation model for tiny devices; phones, wearables, smart home, and robots.
EmotiVoice 😊: a Multi-Voice and Prompt-Controlled TTS Engine
Implementation of Imagen, Google's Text-to-Image Neural Network, in Pytorch
An open source implementation of Microsoft's VALL-E X zero-shot TTS model. Demo is available in https://plachtaa.github.io/vallex/
VITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
Implementation of Dreambooth (https://arxiv.org/abs/2208.12242) with Stable Diffusion
High-quality multi-lingual text-to-speech library by MyShell.ai. Support English, Spanish, French, Chinese, Japanese and Korean.
On-device Speech AI for Apple Silicon
StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models
Silero Models: pre-trained text-to-speech models made embarrassingly simple
Implementation / replication of DALL-E, OpenAI's Text to Image Transformer, in Pytorch
DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism (SVS & TTS); AAAI 2022; Official code
Foundational model for human-like, expressive TTS
Deepchecks: Tests for Continuous Validation of ML Models & Data. Deepchecks is a holistic open-source solution for all of your AI & ML validation needs, enabling to thoroughly test your data and models from research to production.
MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and complex real‑world scenarios, covering stable long‑form speech, multi‑speaker dialogue, voice/character design, environmental sound effects, and real‑time streaming TTS.
Code, Build and Evaluate agents - excellent Model and Skills/MCP/ACP/A2A Support
Collection of apple-native tools for the model context protocol.
High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.
A general fine-tuning kit geared toward image/video/audio diffusion models.
Official Implementation of "Graph of Thoughts: Solving Elaborate Problems with Large Language Models"
A unified evaluation framework for large language models
MARS5 speech model (TTS) from CAMB.AI