Universal LLM Deployment Engine with ML Compilation
LLMRepos·Web / Edge / On-Device Runtime LLM projects
Updated dailyBrowse 93 open-source web / edge / on-device runtime projects in Infrastructure. Compare GitHub stars, recent growth, languages, licenses, and repository activity.
Web / Edge / On-Device Runtime
Browse 93 open-source web / edge / on-device runtime projects in Infrastructure. Compare GitHub stars, recent growth, languages, licenses, and repository activity.
Top Web / Edge / On-Device Runtime repositories
Ranked by current GitHub stars from the latest LLMRepos snapshot.
Showing 40 of 93
High-performance In-browser LLM Inference Engine
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages
Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.
A fast, local neural text to speech system
Astrid is a portable, capability-secure operating system for composable software.
Run frontier LLMs and VLMs locally on Qualcomm devices across NPU, GPU, and CPU with a few lines of code
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
On-device Speech AI for Apple Silicon
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
MimiClaw: Harness on a $5 chip. No OS(Linux). No Node.js. No Mac mini. No Raspberry Pi. No VPS. Hardware agents OS.
RuVector is a High Performance, Real-Time, Self-Learning Ai, Vector GNN, Memory DB built in Rust.
pykeio/ort
Fast ML inference & training for ONNX models in Rust
Run Mixtral-8x7B models in Colab or consumer desktops
Instant, controllable, local pre-trained AI models in Rust
tnm/zclaw
Your personal AI assistant at all-in 888KiB (~35KB in app code). Running on an ESP32. GPIO, cron, custom tools, memory, and more.
Run a 1-billion parameter LLM on a $10 board with 256MB RAM
Run LLMs on AMD Ryzen™ AI NPUs in minutes. Just like Ollama - but purpose-built and deeply optimized for the AMD NPUs.
Run LLMs on AMD Ryzen™ AI NPUs in minutes. Just like Ollama - but purpose-built and deeply optimized for the AMD NPUs.
trymirai/uzu
A high-performance inference engine for AI models
Fast Multimodal LLM on Mobile Devices
Golem Cloud is the agent-native platform for building AI agents and distributed applications that never lose state, never duplicate work, and never require you to build infrastructure.
Run local LLMs like llama, deepseek-distill, kokoro and more inside your browser
Own your compute, own your intelligence. Time to build your Personal AI Data Center.
Run Kimi K3 locally on CPU with ~55GB measured runtime RAM. A single-file Linux inference server powered by cPilot Runtime.
Tensorlake is a serverless runtime for sandboxes and deploying background agentic applications
Gemma Gem runs Google's Gemma 4 model entirely on-device via WebGPU — no API keys, no cloud, no data leaving your machine.
Sudoless Apple Silicon system monitor (native SwiftUI GUI) with ANE / Media Engine / memory-bandwidth tracking
The simplest and lowest-cost AI integration solution. If you like this project, please give it a Star~ | 最简单、最低成本的AI接入方案。喜欢本项目的话点个 Star 吧~
🎤 Lobe TTS - A high-quality & reliable TTS/STT library for Server and Browser
Run full Kimi K3 on a single device. And an OpenAI-compatible API server for local chat and coding agents.
Sandboxed runtime for programming languages and WASI binaries. Works in the browser, on your server, or via MCP.
A modular Swift SDK for audio processing with MLX on Apple Silicon
a lightweight LLM model inference framework
Run the official Stable Diffusion releases in a Docker container with txt2img, img2img, depth2img, pix2pix, upscale4x, and inpaint.
Lightweight, cross-platform process sandboxing powered by OpenAI Codex's runtime. Sandbox any command with file, network, and credential controls.
Offline Speech Recognition with OpenAI Whisper and TensorFlow Lite for Android
usemoss/moss
The retrieval layer for production AI systems. Lightning-fast (<10ms) search without vector databases. Built for browser, edge, on-device, and cloud.
An open framework to simulate and deploy perception-based PX4/ArduPilot drone swarms with ROS2, YOLO, LiDAR, NVIDIA Jetson
Swiftlet is a Swift and Metal runtime that runs large Qwen Mixture-of-Experts models locally on Apple devices by streaming expert weights from storage, enabling 35B and 80B models to run with low RAM, including on iPhone.