Accelerate local LLM inference and finetuning (LLaMA, Mistral, ChatGLM, Qwen, DeepSeek, Mixtral, Gemma, Phi, MiniCPM, Qwen-VL, MiniCPM-V, etc.) on Intel XPU (e.g., local PC with iGPU and NPU, discrete GPU such as Arc, Flex and Max); seamlessly integrate with llama.cpp, Ollama, HuggingFace, LangChain, LlamaIndex, vLLM, DeepSpeed, Axolotl, etc.
LLMRepos·Inference Optimization LLM projects
Updated dailyBrowse 171 open-source inference optimization projects in Infrastructure. Compare GitHub stars, recent growth, languages, licenses, and repository activity.
Inference Optimization
Browse 171 open-source inference optimization projects in Infrastructure. Compare GitHub stars, recent growth, languages, licenses, and repository activity.
Top Inference Optimization repositories
Ranked by current GitHub stars from the latest LLMRepos snapshot.
Showing 40 of 171
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
LLM speculative inference server for consumer & heterogeneous hardware
Fast ML inference & training for ONNX models in Rust
Run Mixtral-8x7B models in Colab or consumer desktops
Instant, controllable, local pre-trained AI models in Rust
⚡ Build your chatbot within minutes on your favorite device; offer SOTA compression techniques for LLMs; run LLMs efficiently on Intel Platforms⚡
MoBA: Mixture of Block Attention for Long-Context LLMs
TokenSpeed is a speed-of-light LLM inference engine.
Run a 1-billion parameter LLM on a $10 board with 256MB RAM
一款简单易用和高性能的AI部署框架 | An Easy-to-Use and High-Performance AI Deployment Framework
Run LLMs on AMD Ryzen™ AI NPUs in minutes. Just like Ollama - but purpose-built and deeply optimized for the AMD NPUs.
Run LLMs on AMD Ryzen™ AI NPUs in minutes. Just like Ollama - but purpose-built and deeply optimized for the AMD NPUs.
Quickly find bottlenecks in Rust - one profiler for CPU, memory, SQL, HTTP, I/O and async code.
trymirai/uzu
A high-performance inference engine for AI models
:metal: TT-NN operator library, and TT-Metalium low level kernel programming model.
Fast Multimodal LLM on Mobile Devices
INT4/INT5/INT8 and FP16 inference on CPU for RWKV language model
xLLM-AI/xllm
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)
Benchmarking Deep Learning operations on different hardware
From-scratch PyTorch implementation of Google's TurboQuant (ICLR 2026) for LLM KV cache compression. 5x compression at 3-bit with 99.5% attention fidelity.
TurboOCR, >200 img/s OmnidocBench. TensorRT FP16, PP-OCRv6, HTTP + gRPC
Run Kimi K3 locally on CPU with ~55GB measured runtime RAM. A single-file Linux inference server powered by cPilot Runtime.
Run Kimi K3 locally on CPU with ~55GB measured runtime RAM. A single-file Linux inference server powered by cPilot Runtime.
A throughput-oriented high-performance serving framework for LLMs
Model router for agentic systems. Routes every prompt to the right model in <50ms. Cut costs 40-70% with just an endpoint change.
Machine learning compiler based on MLIR for Sophgo TPU.
A LLM semantic caching system aiming to enhance user experience by reducing response time via cached query-result pairs.
SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.
MindSpore + 🤗Huggingface: Run any Transformers/Diffusers model on MindSpore with seamless compatibility and acceleration.
A high-performance ML model serving framework, offers dynamic batching and CPU/GPU pipelines to fully exploit your compute machine
Open deep learning compiler stack for Kendryte AI accelerators ✨
Suno AI's Bark model in C/C++ for fast text-to-speech generation
NVIDIA/cuvs
cuVS - a library for vector search and clustering on the GPU
vMLX - JANGTQ Uber Compressed MLX Models - L2 Disk Cache (survives restart) + L1 Paged (super fast ttft) + Hybrid SSM Scheduler + Cont Batching + etc!
Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agent mode, and tool calling.
Run full Kimi K3 on a single device. And an OpenAI-compatible API server for local chat and coding agents.