A high-throughput and memory-efficient inference and serving engine for LLMs
LLMRepos·Production Serving & Deployment LLM projects
Updated dailyBrowse 219 open-source production serving & deployment projects in Infrastructure. Compare GitHub stars, recent growth, languages, licenses, and repository activity.
Production Serving & Deployment
Browse 219 open-source production serving & deployment projects in Infrastructure. Compare GitHub stars, recent growth, languages, licenses, and repository activity.
Top Production Serving & Deployment repositories
Ranked by current GitHub stars from the latest LLMRepos snapshot.
Showing 40 of 219
The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthropic, OpenAI, VertexAI, vLLM, Nvidia NIM]
Wrap Antigravity, ChatGPT Codex, Claude Code, Grok Build as an OpenAI/Gemini/Claude/Codex compatible API service, allowing you to enjoy the free Gemini 3.1 Pro, GPT 5.6 Series, Grok 4.5, Claude model through API
A unified AI model hub for aggregation & distribution. It supports cross-converting various LLMs into OpenAI-compatible, Claude-compatible, or Gemini-compatible formats. A centralized gateway for personal and enterprise model management.
Sub2API 一站式开源中转服务,让 Claude、Openai 、Gemini、Grok订阅统一接入,支持拼车共享,更高效分摊成本,原生工具无缝使用。
SGLang is a high-performance serving framework for large language models and multimodal models.
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Deploy real projects from GitHub or your AI coding agent, then keep them running with AI-powered operations.
Nano vLLM
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
A blazing fast AI Gateway with integrated guardrails. Route to 1,600+ LLMs, 50+ AI Guardrails with 1 fast & friendly API.
Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
The Triton Inference Server provides an optimized cloud and edge inferencing solution.
Large Language Model Text Generation Inference
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.
coaidev/coai
🚀 Next Gen Multi-tenant AI One-Stop Solution. Builtin Admin & Billing System. Enterprise-Grade Unified LLM Gateway Support for 200+ Models And 35+ Providers, Load Balacing w/ Priority-base Routing, Cost Management, Chat Share, Cloud Sync, Credit/Subscription Billing, All File Parsing, Web Search, Built-in Model Cache.
One Postgres for your application data, full-text search, vector retrieval, and aggregations. Home of the pg_search extension.
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
A Datacenter Scale Distributed Inference Serving Framework
Fastest enterprise AI gateway (50x faster than LiteLLM) with adaptive load balancer, cluster mode, guardrails, 1000+ models support & <100 µs overhead at 5k RPS.
⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉
Dockerized OpenAI-compatible wrapper for Kokoro-82M text-to-speech w/multiplatform CPU, AMD, NVIDIA GPU PyTorch; multi-speaker, voice-mixing, auto-stitching, caption timestamps, SSML, readalong web UI
Structured data extraction, instruction calling and agentic workflows with ML, LLM and Vision LLM
⚡️ Open-source AI Gateway — Use any SDK to call 100+ LLMs. Built-in failover, load balancing, cost control & end-to-end tracing.
A blazing fast inference solution for text embeddings models
Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.
Osmantic/ODS
Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.
The python library for real-time communication
prest/prest
PostgreSQL ➕ REST, low-code, simplify and accelerate development, ⚡ instant, realtime, high-performance on any Postgres application, existing or new, MCP server
learn LLM inference system on Apple Silicon for systems engineers: build a tiny vLLM + Qwen
Official Agnes AI gateway and model catalog for OpenAI-compatible text, image, video, and agent workflows.
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
Open-source backend-as-a-service. Postgres, auth, storage, functions, AI gateway, MCP.
Run MCP stdio servers over SSE and SSE over stdio. AI gateway.
LLM speculative inference server for consumer & heterogeneous hardware
Community maintained hardware plugin for vLLM on Ascend