LLMRepos·Production Serving & Deployment LLM projects

Updated daily

Browse 219 open-source production serving & deployment projects in Infrastructure. Compare GitHub stars, recent growth, languages, licenses, and repository activity.

Production Serving & Deployment

Browse 219 open-source production serving & deployment projects in Infrastructure. Compare GitHub stars, recent growth, languages, licenses, and repository activity.

219
Repositories
847
Infrastructure

Top Production Serving & Deployment repositories

Ranked by current GitHub stars from the latest LLMRepos snapshot.

Showing 40 of 219

89.9k
stars

A high-throughput and memory-efficient inference and serving engine for LLMs

PythonApache License 2.0+618 stars in 7dupdated today
57.2k
stars

The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthropic, OpenAI, VertexAI, vLLM, Nvidia NIM]

PythonOther+612 stars in 7dupdated today

Wrap Antigravity, ChatGPT Codex, Claude Code, Grok Build as an OpenAI/Gemini/Claude/Codex compatible API service, allowing you to enjoy the free Gemini 3.1 Pro, GPT 5.6 Series, Grok 4.5, Claude model through API

GoMIT License+1k stars in 7dupdated today

A unified AI model hub for aggregation & distribution. It supports cross-converting various LLMs into OpenAI-compatible, Claude-compatible, or Gemini-compatible formats. A centralized gateway for personal and enterprise model management.

GoGNU Affero General Public License v3.0+762 stars in 7dupdated today
39.2k
stars

Sub2API 一站式开源中转服务,让 Claude、Openai 、Gemini、Grok订阅统一接入,支持拼车共享,更高效分摊成本,原生工具无缝使用。

GoGNU Lesser General Public License v3.0+1.8k stars in 7dupdated today
32.4k
stars

SGLang is a high-performance serving framework for large language models and multimodal models.

PythonApache License 2.0+410 stars in 7dupdated today

Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.

PythonMIT License+103 stars in 7dupdated today
18.3k
stars

Deploy real projects from GitHub or your AI coding agent, then keep them running with AI-powered operations.

TypeScriptOther+5 stars in 7dupdated today
14.5k
stars

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.

PythonOther+67 stars in 7dupdated today
12.8k
stars

A blazing fast AI Gateway with integrated guardrails. Route to 1,600+ LLMs, 50+ AI Guardrails with 1 fast & friendly API.

TypeScriptMIT License+68 stars in 7dupdated 91d ago
12.5k
stars

Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.

PythonApache License 2.0+17 stars in 7dupdated today
11.3k
stars

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

PythonApache License 2.0+156 stars in 7dupdated today

The Triton Inference Server provides an optimized cloud and edge inferencing solution.

PythonBSD 3-Clause "New" or "Revised" License+12 stars in 7dupdated today

The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.

PythonApache License 2.0+21 stars in 7dupdated today

Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.

PythonApache License 2.0+9 stars in 7dupdated today
9.3k
stars

🚀 Next Gen Multi-tenant AI One-Stop Solution. Builtin Admin & Billing System. Enterprise-Grade Unified LLM Gateway Support for 200+ Models And 35+ Providers, Load Balacing w/ Priority-base Routing, Cost Management, Chat Share, Cloud Sync, Credit/Subscription Billing, All File Parsing, Web Search, Built-in Model Cache.

TypeScriptApache License 2.0+11 stars in 7dupdated 165d ago
9.2k
stars

One Postgres for your application data, full-text search, vector retrieval, and aggregations. Home of the pg_search extension.

RustGNU Affero General Public License v3.0updated today

LMDeploy is a toolkit for compressing, deploying, and serving LLMs.

PythonApache License 2.0+7 stars in 7dupdated today
7.8k
stars

A Datacenter Scale Distributed Inference Serving Framework

RustOther+60 stars in 7dupdated today
7.5k
stars

Fastest enterprise AI gateway (50x faster than LiteLLM) with adaptive load balancer, cluster mode, guardrails, 1000+ models support & <100 µs overhead at 5k RPS.

GoApache License 2.0+171 stars in 7dupdated today

⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.

RustApache License 2.0+50 stars in 7dupdated 4d ago
5.7k
stars

High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale

RustApache License 2.0+11 stars in 7dupdated 4d ago

📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉

PythonGNU General Public License v3.0+23 stars in 7dupdated 10d ago

Dockerized OpenAI-compatible wrapper for Kokoro-82M text-to-speech w/multiplatform CPU, AMD, NVIDIA GPU PyTorch; multi-speaker, voice-mixing, auto-stitching, caption timestamps, SSML, readalong web UI

PythonApache License 2.0+28 stars in 7dupdated today
5.2k
stars

Structured data extraction, instruction calling and agentic workflows with ML, LLM and Vision LLM

PythonGNU General Public License v3.0+1 stars in 7dupdated today
5.1k
stars

⚡️ Open-source AI Gateway — Use any SDK to call 100+ LLMs. Built-in failover, load balancing, cost control & end-to-end tracing.

GoOther+61 stars in 7dupdated today

Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.

PythonApache License 2.0+381 stars in 7dupdated 3d ago
4.7k
stars

Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.

PythonApache License 2.0+381 stars in 7dupdated 3d ago
4.6k
stars

The python library for real-time communication

JavaScriptMIT License-1 stars in 7dupdated 224d ago
4.6k
stars

PostgreSQL ➕ REST, low-code, simplify and accelerate development, ⚡ instant, realtime, high-performance on any Postgres application, existing or new, MCP server

GoMIT License+2 stars in 7dupdated 3d ago
4.5k
stars

learn LLM inference system on Apple Silicon for systems engineers: build a tiny vLLM + Qwen

PythonApache License 2.0+17 stars in 7dupdated today

Official Agnes AI gateway and model catalog for OpenAI-compatible text, image, video, and agent workflows.

+918 stars in 7dupdated 6d ago

The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.

PythonOther+43 stars in 7dupdated today

Open-source backend-as-a-service. Postgres, auth, storage, functions, AI gateway, MCP.

TypeScriptApache License 2.0+106 stars in 7dupdated today

Run MCP stdio servers over SSE and SSE over stdio. AI gateway.

TypeScriptMIT License+7 stars in 7dupdated 319d ago

LLM speculative inference server for consumer & heterogeneous hardware

C++Apache License 2.0+23 stars in 7dupdated today

Community maintained hardware plugin for vLLM on Ascend

C++Apache License 2.0+40 stars in 7dupdated today