LLMRepos·Web / Edge / On-Device Runtime LLM projects

Updated daily

Browse 93 open-source web / edge / on-device runtime projects in Infrastructure. Compare GitHub stars, recent growth, languages, licenses, and repository activity.

Web / Edge / On-Device Runtime

Browse 93 open-source web / edge / on-device runtime projects in Infrastructure. Compare GitHub stars, recent growth, languages, licenses, and repository activity.

93
Repositories
847
Infrastructure

Top Web / Edge / On-Device Runtime repositories

Ranked by current GitHub stars from the latest LLMRepos snapshot.

Showing 40 of 93

23.1k
stars

Universal LLM Deployment Engine with ML Compilation

PythonApache License 2.0+19 stars in 7dupdated 7d ago
18.6k
stars

High-performance In-browser LLM Inference Engine

TypeScriptApache License 2.0+21 stars in 7dupdated 20d ago
14.4k
stars

Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages

C++Apache License 2.0+145 stars in 7dupdated today

Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.

SwiftMIT License+37 stars in 7dupdated 31d ago
11.3k
stars

A fast, local neural text to speech system

C++MIT License-1 stars in 7dupdated 363d ago

Astrid is a portable, capability-secure operating system for composable software.

RustApache License 2.0-32 stars in 7dupdated today
8.3k
stars

Run frontier LLMs and VLMs locally on Qualcomm devices across NPU, GPU, and CPU with a few lines of code

RustBSD 3-Clause "New" or "Revised" License+23 stars in 7dupdated today

A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.

PythonMIT License+29 stars in 7dupdated 4d ago

Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook

SwiftApache License 2.0+189 stars in 7dupdated today
5.7k
stars

MimiClaw: Harness on a $5 chip. No OS(Linux). No Node.js. No Mac mini. No Raspberry Pi. No VPS. Hardware agents OS.

CMIT License+20 stars in 7dupdated 3d ago
4.4k
stars

RuVector is a High Performance, Real-Time, Self-Learning Ai, Vector GNN, Memory DB built in Rust.

RustMIT License+19 stars in 7dupdated today
2.5k
stars

Fast ML inference & training for ONNX models in Rust

RustApache License 2.0+19 stars in 7dupdated 1d ago

Run Mixtral-8x7B models in Colab or consumer desktops

PythonMIT License-1 stars in 7dupdated 868d ago
2.2k
stars

Instant, controllable, local pre-trained AI models in Rust

RustApache License 2.0+1 stars in 7dupdated 2d ago
2.2k
stars

Your personal AI assistant at all-in 888KiB (~35KB in app code). Running on an ESP32. GPIO, cron, custom tools, memory, and more.

CMIT License+7 stars in 7dupdated 99d ago
1.9k
stars

Run a 1-billion parameter LLM on a $10 board with 256MB RAM

CMIT License+11 stars in 7dupdated 183d ago

Run LLMs on AMD Ryzen™ AI NPUs in minutes. Just like Ollama - but purpose-built and deeply optimized for the AMD NPUs.

C++MIT License+39 stars in 7dupdated 4d ago
1.8k
stars

Run LLMs on AMD Ryzen™ AI NPUs in minutes. Just like Ollama - but purpose-built and deeply optimized for the AMD NPUs.

C++MIT License+40 stars in 7dupdated 4d ago
1.7k
stars

A high-performance inference engine for AI models

RustMIT License+4 stars in 7dupdated today
1.5k
stars

Golem Cloud is the agent-native platform for building AI agents and distributed applications that never lose state, never duplicate work, and never require you to build infrastructure.

RustOther-14 stars in 7dupdated today

Run local LLMs like llama, deepseek-distill, kokoro and more inside your browser

TypeScriptMIT License+1 stars in 7dupdated 34d ago

Own your compute, own your intelligence. Time to build your Personal AI Data Center.

MIT License+100 stars in 7dupdated 3d ago

Run Kimi K3 locally on CPU with ~55GB measured runtime RAM. A single-file Linux inference server powered by cPilot Runtime.

HTMLOtherupdated 24d ago

Tensorlake is a serverless runtime for sandboxes and deploying background agentic applications

PythonApache License 2.0+5 stars in 7dupdated today

Gemma Gem runs Google's Gemma 4 model entirely on-device via WebGPU — no API keys, no cloud, no data leaving your machine.

TypeScriptApache License 2.0+4 stars in 7dupdated 87d ago

Sudoless Apple Silicon system monitor (native SwiftUI GUI) with ANE / Media Engine / memory-bandwidth tracking

SwiftMIT License+55 stars in 7dupdated 6d ago

The simplest and lowest-cost AI integration solution. If you like this project, please give it a Star~ | 最简单、最低成本的AI接入方案。喜欢本项目的话点个 Star 吧~

CMIT License+0 stars in 7dupdated 227d ago
803
stars

🎤 Lobe TTS - A high-quality & reliable TTS/STT library for Server and Browser

TypeScriptMIT License+2 stars in 7dupdated 175d ago

Run full Kimi K3 on a single device. And an OpenAI-compatible API server for local chat and coding agents.

RustOther+16 stars in 7dupdated 18d ago
770
stars

Sandboxed runtime for programming languages and WASI binaries. Works in the browser, on your server, or via MCP.

TypeScriptMIT License+0 stars in 7dupdated 24d ago

A modular Swift SDK for audio processing with MLX on Apple Silicon

SwiftMIT License+6 stars in 7dupdated 8d ago

a lightweight LLM model inference framework

C++Apache License 2.0+0 stars in 7dupdated 870d ago

Run the official Stable Diffusion releases in a Docker container with txt2img, img2img, depth2img, pix2pix, upscale4x, and inpaint.

PythonGNU Affero General Public License v3.0+0 stars in 7dupdated 970d ago
707
stars

Lightweight, cross-platform process sandboxing powered by OpenAI Codex's runtime. Sandbox any command with file, network, and credential controls.

RustApache License 2.0+4 stars in 7dupdated 99d ago

Offline Speech Recognition with OpenAI Whisper and TensorFlow Lite for Android

C++MIT License+3 stars in 7dupdated 159d ago
668
stars

The retrieval layer for production AI systems. Lightning-fast (<10ms) search without vector databases. Built for browser, edge, on-device, and cloud.

PythonBSD 2-Clause "Simplified" License+9 stars in 7dupdated today

An open framework to simulate and deploy perception-based PX4/ArduPilot drone swarms with ROS2, YOLO, LiDAR, NVIDIA Jetson

C++MIT License+10 stars in 7dupdated today

Swiftlet is a Swift and Metal runtime that runs large Qwen Mixture-of-Experts models locally on Apple devices by streaming expert weights from storage, enabling 35B and 80B models to run with low RAM, including on iPhone.

SwiftApache License 2.0+16 stars in 7dupdated today