LLMRepos·Voice / Vision I/O LLM projects

Updated daily

Browse 89 open-source voice / vision i/o projects in AI Engineering. Compare GitHub stars, recent growth, languages, licenses, and repository activity.

Voice / Vision I/O

Browse 89 open-source voice / vision i/o projects in AI Engineering. Compare GitHub stars, recent growth, languages, licenses, and repository activity.

89
Repositories
4,956
AI Engineering

Top Voice / Vision I/O repositories

Ranked by current GitHub stars from the latest LLMRepos snapshot.

Showing 40 of 89

Community plugin to control Blender 3D with any LLM of your choice

PythonMIT License+293 stars in 7dupdated today
14.7k
stars

Open Source framework for voice agents, multimodal apps, and realtime AI. Maintained by Daily and the community.

PythonBSD 2-Clause "Simplified" License+432 stars in 7dupdated today
13.3k
stars

The fastest browser for AI agents to run browser automation, built for sharing your logged-in browser state with your AI agents, like Codex or Claude Code, without disturbing you. Zero cost, zero config.

JavaScriptMIT License+1.7k stars in 7dupdated today

Enterprise-grade, local-first Agent Workbench for people and agent teams. A unified multi-engine workspace for Codex Harness, DeepSeek Harness, and OpenCode, with unified plugins and Skills, multi-agent projects and tasks, and editable code, documents, presentations, design, and video.

TypeScriptOther+625 stars in 7dupdated today
3.6k
stars

The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | 全网最强 DeepSeek Harness 外挂视觉插件,为 DeepSeek、GLM 等纯文本模型外挂视觉能力,粘贴图片即得结构化 JSON 证据(OCR、版面、语义)。

TypeScriptMIT License+791 stars in 7dupdated today

Hold a key, speak, release — AI-polished text appears at your cursor in any app. Open-source voice input for macOS & Windows. (按住快捷键说话,松开即得润色后的文字)

RustMIT License+108 stars in 7dupdated today

Make any agent harness multimodal-native.

HTMLApache License 2.0+87 stars in 7dupdated today

⭐ All-in-one AI companion! Super Agent Party = Self hosted neuro sama + openclaw! ⭐ 全能AI伴侣!超级智能体派对 = 自托管neuro sama + openclaw!

JavaScriptGNU Affero General Public License v3.0+9 stars in 7dupdated 1d ago

SUSI.AI server backend - the Artificial Intelligence server for personal assistants https://susi.ai

JavaGNU Lesser General Public License v2.1-1 stars in 7dupdated 1,833d ago

LLM Agent Framework in ComfyUI includes MCP sever, Omost,GPT-sovits, ChatTTS,GOT-OCR2.0, and FLUX prompt nodes,access to Feishu,discord,and adapts to all llms with similar openai / aisuite interfaces, such as o1,ollama, gemini, grok, qwen, GLM, deepseek, kimi,doubao. Adapted to local llms, vlm, gguf such as llama-3.3 Janus-Pro, Linkage graphRAG

PythonGNU Affero General Public License v3.0+4 stars in 7dupdated 26d ago

A realtime voice runtime that keeps Agents talking, working, and present. Real-time Voice Runtime for AI Agents

JavaScriptApache License 2.0+60 stars in 7dupdated today

Codex skill for converting slide images, PDFs, and image-based PPTX files into editable PowerPoint decks.

PythonMIT License+109 stars in 7dupdated 27d ago

Let Claude (or any LLM) actually watch a video — scene-aware, deduplicated frames + transcript, from a URL or local file. Runs locally, MIT.

PythonMIT License+28 stars in 7dupdated 2d ago
1.9k
stars

Connect your browser to AI models. Just use Dia on Chrome, Arc or Firefox.

JavaScriptMIT License+2 stars in 7dupdated today

A comprehensive ComfyUI integration for Microsoft's VibeVoice text-to-speech model, enabling high-quality single and multi-speaker voice synthesis directly within your ComfyUI workflows.

PythonMIT License+5 stars in 7dupdated 187d ago
1.3k
stars

Natural voice conversations with Claude Code

PythonMIT License+7 stars in 7dupdated today

Give Claude the ability to watch and understand videos — Claude Code plugin with frame extraction and multimodal audio analysis

TypeScriptMIT License+23 stars in 7dupdated 18d ago
1.2k
stars

PhoneClaw turns phones into local AI agent runtimes with on-device models, native mobile Skills, LiveLand, and optional Mac Gateway inference.

SwiftApache License 2.0+2 stars in 7dupdated 18d ago

A ComfyUI custom node integration for local multi-engine multi-language Text-to-Speech and Voice Conversion. Supports: RVC, Echo-TTS, Qwen3-TTS, Cozy Voice 3, Step Audio EditX, IndexTTS-2, Chatterbox (classic and multilingual), F5-TTS, Higgs Audio 2, 3, and VibeVoice with unlimited text length, SRT timing, Character support, and many audio tools

PythonOther+12 stars in 7dupdated 3d ago

AI speech toolkit for Apple Silicon — ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and CoreML

SwiftApache License 2.0+21 stars in 7dupdated 2d ago

The Self-Coding System for Your App — Alan AI SDK for Cordova

Ruby+0 stars in 7dupdated 489d ago

Microsoft Text-to-Speech API sample code in several languages, part of Cognitive Services.

C#Other+0 stars in 7dupdated 17d ago

提供产品级IOCR自定义模板识别,以图搜图,人像搜索等,免费,可商用,Java AI 人工智能一站式解决方案,为工作减负,为产品研发加速。项目类别包括:以及AI SDK,web应用等。

JavaApache License 2.0+0 stars in 7dupdated 111d ago

Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.

JavaScriptMIT License+319 stars in 7dupdated today

SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.

PythonApache License 2.0+101 stars in 7dupdated today

Vonage REST API client for PHP. API support for SMS, Voice, Text-to-Speech, Numbers, Verify (2FA) and more.

PHPApache License 2.0-1 stars in 7dupdated 40d ago
925
stars

TongFlow — Multimodal GenAI Studio

TypeScriptGNU Affero General Public License v3.0+85 stars in 7dupdated today

Voice-first local agent orchestration runtime for auditable DAG workflows.

TypeScriptMIT License+18 stars in 7dupdated today
847
stars

Magick is a cutting-edge toolkit for a new kind of AI builder. Make Magick with us!

TypeScriptOther+0 stars in 7dupdated 426d ago

🔥🔥 LLaVA++: Extending LLaVA with Phi-3 and LLaMA-3 (LLaVA LLaMA-3, LLaVA Phi-3)

Python+1 stars in 7dupdated 385d ago

[dsh]为纯文本模型设计更强大的视觉工具箱:一行安装使用、粘贴图片直接识别、多张图片问答、截图到前端UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, grounding, pixel diff, Artifacts, and Web UI.

TypeScriptMIT License+193 stars in 7dupdated today

A real-time interactive Omni Avatar built on LiveKit, which allows you to seamlessly integrate with any open source Avatar components (real-time model, visual, voice, memory, search, etc.).

PythonApache License 2.0+2 stars in 7dupdated 5d ago

Rapida is an open-source, end-to-end voice AI orchestration platform for building real-time conversational voice agents with audio streaming, STT, TTS, VAD, multi-channel integration, agent state management, and observability.

GoOther+5 stars in 7dupdated today

Agentic ADK is an Agent application development framework launched by Alibaba International AI Business, based on Google-ADK and Ali-LangEngine.

JavaApache License 2.0+1 stars in 7dupdated 246d ago

Rust Agent Development Kit (ADK-Rust): Build AI agents in Rust with modular components for models, tools, memory, realtime voice, and more. ADK-Rust is a flexible framework for developing AI agents with simplicity and power. Model-agnostic, deployment-agnostic, optimized for frontier AI models. Includes support for real-time voice agents.

RustOther+12 stars in 7dupdated 1d ago

The Self-Coding System for Your App — Alan AI SDK for React Native

Ruby+0 stars in 7dupdated 191d ago

An MCP Server for Android running on the phone, optmized for token usage, supports also files downloads and cloudflare and ngrok automated tunnelling.

KotlinMIT License+74 stars in 7dupdated 4d ago

JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents

TypeScriptApache License 2.0+99 stars in 7dupdated 22d ago

A general framework for distilling human-created multimodal resources into reusable, executable skills that AI agents can browse, compose, and run, validated across diverse domains including web, PowerPoint, Excel, Blender, CAD, Unreal Engine 5, and REAPER-based music production.

PythonMIT License+27 stars in 7dupdated 38d ago