LLMRepos·Audio / Speech Models LLM projects

Updated daily

Browse 72 open-source audio / speech models projects in Models. Compare GitHub stars, recent growth, languages, licenses, and repository activity.

Audio / Speech Models

Browse 72 open-source audio / speech models projects in Models. Compare GitHub stars, recent growth, languages, licenses, and repository activity.

72
Repositories
144
Models

Top Audio / Speech Models repositories

Ranked by current GitHub stars from the latest LLMRepos snapshot.

Showing 40 of 72

1 min voice data can also be used to train a good TTS model! (few shot voice cloning)

PythonMIT License+200 stars in 7dupdated 6d ago
45.9k
stars

🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production

PythonMozilla Public License 2.0+39 stars in 7dupdated 738d ago

Instant voice cloning by MIT and MyShell. Audio foundation model.

PythonMIT License+125 stars in 7dupdated 492d ago

🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time

PythonOther+4 stars in 7dupdated 174d ago

An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

PythonOther+346 stars in 7dupdated 6d ago

Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.

PythonApache License 2.0+97 stars in 7dupdated 91d ago

Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.

PythonApache License 2.0+94 stars in 7dupdated 91d ago

Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.

PythonMIT License+103 stars in 7dupdated today
19.4k
stars

A TTS model capable of generating ultra-realistic dialogue in one pass.

PythonApache License 2.0+6 stars in 7dupdated 278d ago
18.3k
stars

A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)

PythonApache License 2.0+157 stars in 7dupdated today
14.4k
stars

Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages

C++Apache License 2.0+145 stars in 7dupdated today

Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.

SwiftMIT License+37 stars in 7dupdated 31d ago
11.3k
stars

A fast, local neural text to speech system

C++MIT License-1 stars in 7dupdated 363d ago
10.3k
stars

Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get started in the field of audio, music, and speech generation research and development.

PythonMIT License+21 stars in 7dupdated 152d ago
10.2k
stars

:robot: :speech_balloon: Deep learning for Text to Speech (Discussion forum: https://discourse.mozilla.org/c/tts)

Jupyter NotebookMozilla Public License 2.0+3 stars in 7dupdated 1,019d ago
9.9k
stars

End-to-End Speech Processing Toolkit

PythonApache License 2.0+10 stars in 7dupdated today

EmotiVoice 😊: a Multi-Voice and Prompt-Controlled TTS Engine

PythonApache License 2.0+6 stars in 7dupdated 741d ago
7.9k
stars

An open source implementation of Microsoft's VALL-E X zero-shot TTS model. Demo is available in https://plachtaa.github.io/vallex/

PythonMIT License-2 stars in 7dupdated 925d ago
7.9k
stars

VITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech

PythonMIT License-2 stars in 7dupdated 993d ago
7.8k
stars

A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.

PythonMIT License+29 stars in 7dupdated 4d ago
7.6k
stars

High-quality multi-lingual text-to-speech library by MyShell.ai. Support English, Spanish, French, Chinese, Japanese and Korean.

PythonMIT License+21 stars in 7dupdated 608d ago
6.3k
stars

StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models

PythonMIT License+3 stars in 7dupdated 745d ago

Silero Models: pre-trained text-to-speech models made embarrassingly simple

Jupyter NotebookOther+13 stars in 7dupdated 24d ago

DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism (SVS & TTS); AAAI 2022; Official code

PythonMIT License+6 stars in 7dupdated 31d ago

Foundational model for human-like, expressive TTS

PythonApache License 2.0+0 stars in 7dupdated 755d ago
2.7k
stars

[AutoArk] GPA (General Purpose Audio) can do ASR, TTS and voice conversion with one tiny model!

PythonApache License 2.0+251 stars in 7dupdated 91d ago

🚀 一键部署(含离线整合包)!基于 ChatTTS ,支持流式输出、音色抽卡、长音频生成和分角色朗读。简单易用,无需复杂安装。

Python+4 stars in 7dupdated 85d ago
2.6k
stars

MARY TTS -- an open-source, multilingual text-to-speech synthesis system written in pure java

JavaOther-1 stars in 7dupdated 584d ago

Vietnamese TTS with instant voice cloning • On-device • Real-time CPU inference • 24kHz audio quality • Chuyển văn bản thành giọng nói tiếng Việt • Text to speech tiếng Việt • TTS tiếng Việt

PythonApache License 2.0+59 stars in 7dupdated today

The collection of pre-trained, state-of-the-art AI models for ailia SDK

PythonOther+17 stars in 7dupdated today
2.4k
stars

HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis

PythonMIT License+1 stars in 7dupdated 758d ago

DeepMind's Tacotron-2 Tensorflow implementation

PythonMIT License-1 stars in 7dupdated 1,145d ago
2.2k
stars

PyTorch implementation of VALL-E(Zero-Shot Text-To-Speech), Reproduced Demo https://lifeiteng.github.io/valle/index.html

PythonApache License 2.0+0 stars in 7dupdated 348d ago

Controllable and fast Text-to-Speech for over 7000 languages!

PythonApache License 2.0-1 stars in 7dupdated 211d ago
2.2k
stars

WaveRNN Vocoder + TTS

PythonMIT License-1 stars in 7dupdated 1,514d ago

Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR benchmarks, while also offering outstanding singing lyrics recognition capability.

PythonApache License 2.0+8 stars in 7dupdated 180d ago
1.8k
stars

a free and open source speech synthesizer for Russian and other languages

C++GNU General Public License v2.0+0 stars in 7dupdated 14d ago

GPT-SoVITS ONNX Inference Engine & Model Converter

PythonMIT License+20 stars in 7dupdated 21d ago

Unofficial Parallel WaveGAN (+ MelGAN & Multi-band MelGAN & HiFi-GAN & StyleMelGAN) with Pytorch

Jupyter NotebookMIT License+0 stars in 7dupdated 855d ago