Run frontier AI locally.
LLMRepos·Distributed Inference LLM projects
Updated dailyBrowse 62 open-source distributed inference projects in Infrastructure. Compare GitHub stars, recent growth, languages, licenses, and repository activity.
Distributed Inference
Browse 62 open-source distributed inference projects in Infrastructure. Compare GitHub stars, recent growth, languages, licenses, and repository activity.
Top Distributed Inference repositories
Ranked by current GitHub stars from the latest LLMRepos snapshot.
Showing 40 of 62
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
Making large AI models cheaper, faster and more accessible
Easy-to-use and powerful LLM and SLM library with awesome model zoo.
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.
🌸 Run LLMs at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading
Build, Manage and Deploy AI/ML Systems
Running large language models on a single GPU for throughput-oriented scenarios.
A Datacenter Scale Distributed Inference Serving Framework
FlashInfer: Kernel Library for LLM Serving
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
FEDML - The unified and scalable ML library for large-scale distributed training, model serving, and federated learning. FEDML Launch, a cross-cloud scheduler, further enables running any AI jobs on any GPU cloud or on-premise cluster. Built on this library, TensorOpera AI (https://TensorOpera.ai) is your generative AI platform at scale.
cubestudio开源云原生一站式机器学习/深度学习/大模型AI平台/MaaS/mlops/人工智能平台/训推平台,算法全链路流程,多租户,算力租赁平台,token中转,拖拉拽任务流pipeline编排,多机多卡分布式训练,超参搜索,推理服务,VGPU虚拟化,云边端协同,边缘计算,自动化标注平台,deepseek等大模型sft微调/奖励模型/强化学习训练,vllm/ollama/mindie大模型多机推理,私有知识库llmops智能体,AI模型市场,支持国产异构算力调度,昇腾/寒武纪/海光/摩尔/沐曦等,支持ib/roce/RDMA,信创支持
Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kubernetes, and bare metal.
UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)
Parallax is a distributed model serving framework that lets you build your own AI cluster anywhere
RayLLM - LLMs on Ray (Archived). Read README for more info.
SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.
🚀 Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support
SEED RL: Scalable and Efficient Deep-RL with Accelerated Central Inference. Implements IMPALA and R2D2 algorithms in TF2 with SEED's architecture.
Agentic RL on Any Harness at Scale
GPU worker client for the Talos network. Pairs with your Talos account, serves open-model inference jobs over a WebSocket, and reports uptime for payouts.
[CVPR 2024 Highlight] DistriFusion: Distributed Parallel Inference for High-Resolution Diffusion Models
run-ai/genv
GPU environment and cluster management with LLM support
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
Trajectory Planner in Multi-Agent and Dynamic Environments
An open framework to simulate and deploy perception-based PX4/ArduPilot drone swarms with ROS2, YOLO, LiDAR, NVIDIA Jetson
A high-performance inference system for large language models, designed for production environments.
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Flink Agents is an Agentic AI framework based on Apache Flink
InternEvo is an open-sourced lightweight training framework aims to support model pre-training without the need for extensive dependencies.
Accelerate inference without tears
Super-Efficient RLHF Training of LLMs with Parameter Reallocation
ClearML Agent - MLOps/LLMOps made easy. MLOps/LLMOps scheduler & orchestration solution
llm-d Router: The intelligent entry point for inference requests
rlops/rlix
Run more RL experiments. Wait less for GPUs.
thushan/olla
High-performance lightweight proxy and load balancer for LLM infrastructure. Intelligent routing, automatic failover and unified model discovery across local and remote inference backends.