LLMRepos·Distributed Inference LLM projects

Updated daily

Browse 62 open-source distributed inference projects in Infrastructure. Compare GitHub stars, recent growth, languages, licenses, and repository activity.

Distributed Inference

Browse 62 open-source distributed inference projects in Infrastructure. Compare GitHub stars, recent growth, languages, licenses, and repository activity.

62
Repositories
847
Infrastructure

Top Distributed Inference repositories

Ranked by current GitHub stars from the latest LLMRepos snapshot.

Showing 40 of 62

47k
stars

Run frontier AI locally.

PythonApache License 2.0+148 stars in 7dupdated 62d ago
43.6k
stars

Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.

PythonApache License 2.0+67 stars in 7dupdated today

Making large AI models cheaper, faster and more accessible

PythonApache License 2.0+0 stars in 7dupdated today

Easy-to-use and powerful LLM and SLM library with awesome model zoo.

PythonApache License 2.0+5 stars in 7dupdated 93d ago

The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.

PythonApache License 2.0+21 stars in 7dupdated today

🌸 Run LLMs at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading

PythonMIT License+15 stars in 7dupdated 716d ago
10.2k
stars

Build, Manage and Deploy AI/ML Systems

PythonApache License 2.0+16 stars in 7dupdated 3d ago

Running large language models on a single GPU for throughput-oriented scenarios.

PythonApache License 2.0-3 stars in 7dupdated 666d ago
7.8k
stars

A Datacenter Scale Distributed Inference Serving Framework

RustOther+60 stars in 7dupdated today

FlashInfer: Kernel Library for LLM Serving

PythonApache License 2.0+55 stars in 7dupdated today
5.5k
stars

A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.

PythonApache License 2.0+43 stars in 7dupdated today
4.1k
stars

FEDML - The unified and scalable ML library for large-scale distributed training, model serving, and federated learning. FEDML Launch, a cross-cloud scheduler, further enables running any AI jobs on any GPU cloud or on-premise cluster. Built on this library, TensorOpera AI (https://TensorOpera.ai) is your generative AI platform at scale.

PythonApache License 2.0+3 stars in 7dupdated 300d ago

cubestudio开源云原生一站式机器学习/深度学习/大模型AI平台/MaaS/mlops/人工智能平台/训推平台,算法全链路流程,多租户,算力租赁平台,token中转,拖拉拽任务流pipeline编排,多机多卡分布式训练,超参搜索,推理服务,VGPU虚拟化,云边端协同,边缘计算,自动化标注平台,deepseek等大模型sft微调/奖励模型/强化学习训练,vllm/ollama/mindie大模型多机推理,私有知识库llmops智能体,AI模型市场,支持国产异构算力调度,昇腾/寒武纪/海光/摩尔/沐曦等,支持ib/roce/RDMA,信创支持

PythonOther+18 stars in 7dupdated 8d ago
2.2k
stars

Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kubernetes, and bare metal.

PythonMozilla Public License 2.0+5 stars in 7dupdated today
1.5k
stars

UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)

C++Apache License 2.0+2 stars in 7dupdated 1d ago

Parallax is a distributed model serving framework that lets you build your own AI cluster anywhere

PythonApache License 2.0+0 stars in 7dupdated 54d ago

RayLLM - LLMs on Ray (Archived). Read README for more info.

+0 stars in 7dupdated 530d ago

SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.

PythonApache License 2.0+101 stars in 7dupdated today

🚀 Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support

PythonApache License 2.0+24 stars in 7dupdated today

SEED RL: Scalable and Efficient Deep-RL with Accelerated Central Inference. Implements IMPALA and R2D2 algorithms in TF2 with SEED's architecture.

PythonApache License 2.0+0 stars in 7dupdated 1,364d ago
769
stars

GPU worker client for the Talos network. Pairs with your Talos account, serves open-model inference jobs over a WebSocket, and reports uptime for payouts.

PythonMIT License-21 stars in 7dupdated 47d ago

[CVPR 2024 Highlight] DistriFusion: Distributed Parallel Inference for High-Resolution Diffusion Models

PythonMIT License+0 stars in 7dupdated 630d ago
659
stars

GPU environment and cluster management with LLM support

PythonGNU Affero General Public License v3.0+0 stars in 7dupdated 830d ago

Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2

RustApache License 2.0+6 stars in 7dupdated today
615
stars

Trajectory Planner in Multi-Agent and Dynamic Environments

C++BSD 3-Clause "New" or "Revised" License-1 stars in 7dupdated 1,356d ago

An open framework to simulate and deploy perception-based PX4/ArduPilot drone swarms with ROS2, YOLO, LiDAR, NVIDIA Jetson

C++MIT License+10 stars in 7dupdated today

A high-performance inference system for large language models, designed for production environments.

C++Apache License 2.0+0 stars in 7dupdated 248d ago
495
stars

Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton

GoApache License 2.0+2 stars in 7dupdated today
472
stars

Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.

RustApache License 2.0+10 stars in 7dupdated today
472
stars

Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.

RustApache License 2.0+10 stars in 7dupdated today

JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).

PythonApache License 2.0+1 stars in 7dupdated 231d ago

Flink Agents is an Agentic AI framework based on Apache Flink

JavaApache License 2.0+0 stars in 7dupdated today

InternEvo is an open-sourced lightweight training framework aims to support model pre-training without the need for extensive dependencies.

PythonApache License 2.0+0 stars in 7dupdated 368d ago

Super-Efficient RLHF Training of LLMs with Parameter Reallocation

PythonApache License 2.0+0 stars in 7dupdated 487d ago

ClearML Agent - MLOps/LLMOps made easy. MLOps/LLMOps scheduler & orchestration solution

PythonApache License 2.0+0 stars in 7dupdated 7d ago

llm-d Router: The intelligent entry point for inference requests

GoApache License 2.0+9 stars in 7dupdated today
296
stars

Run more RL experiments. Wait less for GPUs.

PythonApache License 2.0+1 stars in 7dupdated 8d ago
284
stars

High-performance lightweight proxy and load balancer for LLM infrastructure. Intelligent routing, automatic failover and unified model discovery across local and remote inference backends.

GoApache License 2.0+5 stars in 7dupdated 3d ago