LLMRepos·Models LLM projects

Updated daily

Browse 144 open-source models LLM projects across Audio / Speech Models, Vision & Multimodal Models, Foundation / LLM Models. Compare stars, growth, languages, licenses, and repository activity.

All categories

Models

Model weights and collections.

144
Repositories
8
Subcategories

Subcategories

Top Models repositories

Ranked by current GitHub stars from the latest LLMRepos snapshot.

Showing 40 of 144

1 min voice data can also be used to train a good TTS model! (few shot voice cloning)

PythonMIT License+200 stars in 7dupdated 6d ago
45.9k
stars

🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production

PythonMozilla Public License 2.0+39 stars in 7dupdated 738d ago

Instant voice cloning by MIT and MyShell. Audio foundation model.

PythonMIT License+125 stars in 7dupdated 492d ago

🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time

PythonOther+4 stars in 7dupdated 174d ago

An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

PythonOther+346 stars in 7dupdated 6d ago

Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.

PythonApache License 2.0+97 stars in 7dupdated 91d ago

Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.

PythonApache License 2.0+94 stars in 7dupdated 91d ago

FinGPT: Open-Source Financial Large Language Models! Revolutionize 🔥 We release the trained model on HuggingFace.

Jupyter NotebookMIT License+45 stars in 7dupdated 22d ago

Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.

PythonMIT License+103 stars in 7dupdated today
19.4k
stars

A TTS model capable of generating ultra-realistic dialogue in one pass.

PythonApache License 2.0+6 stars in 7dupdated 278d ago
14.4k
stars

Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages

C++Apache License 2.0+145 stars in 7dupdated today

Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.

SwiftMIT License+37 stars in 7dupdated 31d ago

Implementation of DALL-E 2, OpenAI's updated text-to-image synthesis neural network, in Pytorch

PythonMIT License+1 stars in 7dupdated 835d ago
11.3k
stars

A fast, local neural text to speech system

C++MIT License-1 stars in 7dupdated 363d ago
10.3k
stars

Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get started in the field of audio, music, and speech generation research and development.

PythonMIT License+21 stars in 7dupdated 152d ago
10.2k
stars

:robot: :speech_balloon: Deep learning for Text to Speech (Discussion forum: https://discourse.mozilla.org/c/tts)

Jupyter NotebookMozilla Public License 2.0+3 stars in 7dupdated 1,019d ago
9.9k
stars

End-to-End Speech Processing Toolkit

PythonApache License 2.0+10 stars in 7dupdated today

14MB foundation model for tiny devices; phones, wearables, smart home, and robots.

PythonApache License 2.0updated today

EmotiVoice 😊: a Multi-Voice and Prompt-Controlled TTS Engine

PythonApache License 2.0+6 stars in 7dupdated 741d ago

Implementation of Imagen, Google's Text-to-Image Neural Network, in Pytorch

PythonMIT License+1 stars in 7dupdated 686d ago
7.9k
stars

An open source implementation of Microsoft's VALL-E X zero-shot TTS model. Demo is available in https://plachtaa.github.io/vallex/

PythonMIT License-2 stars in 7dupdated 925d ago
7.9k
stars

VITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech

PythonMIT License-2 stars in 7dupdated 993d ago
7.8k
stars

A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.

PythonMIT License+29 stars in 7dupdated 4d ago

Implementation of Dreambooth (https://arxiv.org/abs/2208.12242) with Stable Diffusion

Jupyter NotebookMIT License+1 stars in 7dupdated 1,356d ago
7.6k
stars

High-quality multi-lingual text-to-speech library by MyShell.ai. Support English, Spanish, French, Chinese, Japanese and Korean.

PythonMIT License+21 stars in 7dupdated 608d ago
6.3k
stars

StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models

PythonMIT License+3 stars in 7dupdated 745d ago

Silero Models: pre-trained text-to-speech models made embarrassingly simple

Jupyter NotebookOther+13 stars in 7dupdated 24d ago

Implementation / replication of DALL-E, OpenAI's Text to Image Transformer, in Pytorch

PythonMIT License+0 stars in 7dupdated 919d ago

DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism (SVS & TTS); AAAI 2022; Official code

PythonMIT License+6 stars in 7dupdated 31d ago

Foundational model for human-like, expressive TTS

PythonApache License 2.0+0 stars in 7dupdated 755d ago

Deepchecks: Tests for Continuous Validation of ML Models & Data. Deepchecks is a holistic open-source solution for all of your AI & ML validation needs, enabling to thoroughly test your data and models from research to production.

PythonOther+0 stars in 7dupdated 239d ago

MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and complex real‑world scenarios, covering stable long‑form speech, multi‑speaker dialogue, voice/character design, environmental sound effects, and real‑time streaming TTS.

PythonApache License 2.0+20 stars in 7dupdated 29d ago

Code, Build and Evaluate agents - excellent Model and Skills/MCP/ACP/A2A Support

PythonApache License 2.0+9 stars in 7dupdated 1d ago

Collection of apple-native tools for the model context protocol.

TypeScriptMIT License+0 stars in 7dupdated 379d ago

High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.

PythonApache License 2.0+5 stars in 7dupdated today
2.9k
stars

A general fine-tuning kit geared toward image/video/audio diffusion models.

PythonGNU Affero General Public License v3.0+7 stars in 7dupdated today

Official Implementation of "Graph of Thoughts: Solving Elaborate Problems with Large Language Models"

PythonOther+2 stars in 7dupdated 153d ago

A unified evaluation framework for large language models

PythonMIT License-2 stars in 7dupdated 185d ago
2.8k
stars

MARS5 speech model (TTS) from CAMB.AI

Jupyter NotebookGNU Affero General Public License v3.0+0 stars in 7dupdated 753d ago