Get up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
LLMRepos·Local Runtime / Inference Engine LLM projects
Updated dailyBrowse 322 open-source local runtime / inference engine projects in Infrastructure. Compare GitHub stars, recent growth, languages, licenses, and repository activity.
Local Runtime / Inference Engine
Browse 322 open-source local runtime / inference engine projects in Infrastructure. Compare GitHub stars, recent growth, languages, licenses, and repository activity.
Top Local Runtime / Inference Engine repositories
Ranked by current GitHub stars from the latest LLMRepos snapshot.
Showing 40 of 322
The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface.
LLM inference in C/C++
LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
Hundreds of models & providers. One command to find what runs on your hardware.
AirLLM 70B inference with single 4GB GPU
Find secrets with Gitleaks 🔑
TimesFM (Time Series Foundation Model) is a pretrained time-series foundation model developed by Google Research for time-series forecasting.
🚀 AI 全自动短视频引擎 | AI Fully Automated Short Video Engine
Distribute and run LLMs with a single file.
Distribute and run LLMs with a single file.
本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)
jundot/omlx
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations
Deploy real projects from GitHub or your AI coding agent, then keep them running with AI-powered operations.
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
alibaba/MNN
MNN: A blazing-fast, lightweight inference engine battle-tested by Alibaba, powering high-performance on-device LLMs and Edge AI.
20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.
Talk to any LLM with hands-free voice interaction, voice interruption, and Live2D taking face running locally across platforms
MiMo Code: Where Models and Agents Co-Evolve
Run GGUF models easily with a KoboldAI UI. One File. Zero Install.
A self-hosted, offline, ChatGPT-like chatbot. Powered by Llama 2. 100% private, with no data leaving your device. New: Code Llama support!
Official inference library for Mistral models
OpenVINO™ is an open source toolkit for optimizing and deploying AI inference
Astrid is a portable, capability-secure operating system for composable software.
High-speed Large Language Model Serving for Local Deployment
Accelerate local LLM inference and finetuning (LLaMA, Mistral, ChatGLM, Qwen, DeepSeek, Mixtral, Gemma, Phi, MiniCPM, Qwen-VL, MiniCPM-V, etc.) on Intel XPU (e.g., local PC with iGPU and NPU, discrete GPU such as Arc, Flex and Max); seamlessly integrate with llama.cpp, Ollama, HuggingFace, LangChain, LlamaIndex, vLLM, DeepSpeed, Axolotl, etc.
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
Run frontier LLMs and VLMs locally on Qualcomm devices across NPU, GPU, and CPU with a few lines of code
Own your AI. The native macOS harness for AI agents -- any model, persistent memory, autonomous execution, cryptographic identity. Built in Swift. Fully offline. Open source.
Open-source implementation of AlphaEvolve
Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability, agent orchestration, and guardrails so you stay focused on your agents core logic.
TypeScript AI agent orchestration framework with dynamic workflows. Describe the goal, not the graph: a coordinator plans the task DAG at runtime and runs it on any LLM (Claude, ChatGPT, Gemini, DeepSeek, or local models).
Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
FlashInfer: Kernel Library for LLM Serving
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.