LLMRepos·Local Runtime / Inference Engine LLM projects

Updated daily

Browse 322 open-source local runtime / inference engine projects in Infrastructure. Compare GitHub stars, recent growth, languages, licenses, and repository activity.

Local Runtime / Inference Engine

Browse 322 open-source local runtime / inference engine projects in Infrastructure. Compare GitHub stars, recent growth, languages, licenses, and repository activity.

322
Repositories
847
Infrastructure

Top Local Runtime / Inference Engine repositories

Ranked by current GitHub stars from the latest LLMRepos snapshot.

Showing 40 of 322

179.4k
stars

Get up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.

GoMIT License+541 stars in 7dupdated today
129.8k
stars

The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface.

PythonGNU General Public License v3.0+1.6k stars in 7dupdated today
125.5k
stars

LLM inference in C/C++

C++MIT License+1.1k stars in 7dupdated today
48.7k
stars

LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.

GoMIT License+126 stars in 7dupdated today
36.1k
stars

VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning

PythonApache License 2.0+286 stars in 7dupdated 12d ago
33.9k
stars

Hundreds of models & providers. One command to find what runs on your hardware.

RustMIT License+1.5k stars in 7dupdated today
32.5k
stars

AirLLM 70B inference with single 4GB GPU

Jupyter NotebookApache License 2.0+964 stars in 7dupdated today
28.9k
stars

Find secrets with Gitleaks 🔑

GoMIT License+154 stars in 7dupdated 5d ago

TimesFM (Time Series Foundation Model) is a pretrained time-series foundation model developed by Google Research for time-series forecasting.

PythonApache License 2.0+232 stars in 7dupdated 41d ago

🚀 AI 全自动短视频引擎 | AI Fully Automated Short Video Engine

PythonApache License 2.0+420 stars in 7dupdated 71d ago

Distribute and run LLMs with a single file.

C++Other+66 stars in 7dupdated 3d ago

Distribute and run LLMs with a single file.

C++Other+66 stars in 7dupdated 3d ago

本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)

HTMLApache License 2.0+48 stars in 7dupdated 36d ago
20.6k
stars

LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

PythonApache License 2.0+1.5k stars in 7dupdated today

A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations

PythonApache License 2.0+41 stars in 7dupdated today
18.3k
stars

Deploy real projects from GitHub or your AI coding agent, then keep them running with AI-powered operations.

TypeScriptOther+5 stars in 7dupdated today
18.3k
stars

A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)

PythonApache License 2.0+157 stars in 7dupdated today
16k
stars

MNN: A blazing-fast, lightweight inference engine battle-tested by Alibaba, powering high-performance on-device LLMs and Edge AI.

C++Apache License 2.0+71 stars in 7dupdated today
13.6k
stars

20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.

PythonApache License 2.0+6 stars in 7dupdated 7d ago

Talk to any LLM with hands-free voice interaction, voice interruption, and Live2D taking face running locally across platforms

PythonOther+138 stars in 7dupdated 101d ago

MiMo Code: Where Models and Agents Co-Evolve

TypeScriptMIT License+91 stars in 7dupdated today
11.5k
stars

Run GGUF models easily with a KoboldAI UI. One File. Zero Install.

C++GNU Affero General Public License v3.0+73 stars in 7dupdated today
10.9k
stars

A self-hosted, offline, ChatGPT-like chatbot. Powered by Llama 2. 100% private, with no data leaving your device. New: Code Llama support!

TypeScriptMIT License+3 stars in 7dupdated 853d ago

Official inference library for Mistral models

Jupyter NotebookApache License 2.0-5 stars in 7dupdated 69d ago

OpenVINO™ is an open source toolkit for optimizing and deploying AI inference

C++Apache License 2.0+46 stars in 7dupdated today

Astrid is a portable, capability-secure operating system for composable software.

RustApache License 2.0-32 stars in 7dupdated today

High-speed Large Language Model Serving for Local Deployment

C++MIT License+24 stars in 7dupdated 105d ago
8.9k
stars

Accelerate local LLM inference and finetuning (LLaMA, Mistral, ChatGLM, Qwen, DeepSeek, Mixtral, Gemma, Phi, MiniCPM, Qwen-VL, MiniCPM-V, etc.) on Intel XPU (e.g., local PC with iGPU and NPU, discrete GPU such as Arc, Flex and Max); seamlessly integrate with llama.cpp, Ollama, HuggingFace, LangChain, LlamaIndex, vLLM, DeepSpeed, Axolotl, etc.

PythonApache License 2.0-2 stars in 7dupdated 208d ago
8.8k
stars

The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!

PythonApache License 2.0+11 stars in 7dupdated 3d ago
8.3k
stars

Run frontier LLMs and VLMs locally on Qualcomm devices across NPU, GPU, and CPU with a few lines of code

RustBSD 3-Clause "New" or "Revised" License+23 stars in 7dupdated today
7.7k
stars

Own your AI. The native macOS harness for AI agents -- any model, persistent memory, autonomous execution, cryptographic identity. Built in Swift. Fully offline. Open source.

SwiftMIT License+58 stars in 7dupdated today
7k
stars

Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability, agent orchestration, and guardrails so you stay focused on your agents core logic.

RustApache License 2.0+13 stars in 7dupdated 5d ago

TypeScript AI agent orchestration framework with dynamic workflows. Describe the goal, not the graph: a coordinator plans the task DAG at runtime and runs it on any LLM (Claude, ChatGPT, Gemini, DeepSeek, or local models).

TypeScriptMIT License+40 stars in 7dupdated today
6.5k
stars

Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.

PythonMIT License+138 stars in 7dupdated 10d ago

A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.

CApache License 2.0+454 stars in 7dupdated 17d ago

Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook

SwiftApache License 2.0+189 stars in 7dupdated today

FlashInfer: Kernel Library for LLM Serving

PythonApache License 2.0+55 stars in 7dupdated today
5.8k
stars

Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes

GoApache License 2.0+27 stars in 7dupdated 1d ago

⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.

RustApache License 2.0+50 stars in 7dupdated 4d ago