Community plugin to control Blender 3D with any LLM of your choice
LLMRepos·Voice / Vision I/O LLM projects
Updated dailyBrowse 89 open-source voice / vision i/o projects in AI Engineering. Compare GitHub stars, recent growth, languages, licenses, and repository activity.
Voice / Vision I/O
Browse 89 open-source voice / vision i/o projects in AI Engineering. Compare GitHub stars, recent growth, languages, licenses, and repository activity.
Top Voice / Vision I/O repositories
Ranked by current GitHub stars from the latest LLMRepos snapshot.
Showing 40 of 89
Open Source framework for voice agents, multimodal apps, and realtime AI. Maintained by Daily and the community.
The fastest browser for AI agents to run browser automation, built for sharing your logged-in browser state with your AI agents, like Codex or Claude Code, without disturbing you. Zero cost, zero config.
Enterprise-grade, local-first Agent Workbench for people and agent teams. A unified multi-engine workspace for Codex Harness, DeepSeek Harness, and OpenCode, with unified plugins and Skills, multi-agent projects and tasks, and editable code, documents, presentations, design, and video.
The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | 全网最强 DeepSeek Harness 外挂视觉插件,为 DeepSeek、GLM 等纯文本模型外挂视觉能力,粘贴图片即得结构化 JSON 证据(OCR、版面、语义)。
Hold a key, speak, release — AI-polished text appears at your cursor in any app. Open-source voice input for macOS & Windows. (按住快捷键说话,松开即得润色后的文字)
Make any agent harness multimodal-native.
⭐ All-in-one AI companion! Super Agent Party = Self hosted neuro sama + openclaw! ⭐ 全能AI伴侣!超级智能体派对 = 自托管neuro sama + openclaw!
SUSI.AI server backend - the Artificial Intelligence server for personal assistants https://susi.ai
LLM Agent Framework in ComfyUI includes MCP sever, Omost,GPT-sovits, ChatTTS,GOT-OCR2.0, and FLUX prompt nodes,access to Feishu,discord,and adapts to all llms with similar openai / aisuite interfaces, such as o1,ollama, gemini, grok, qwen, GLM, deepseek, kimi,doubao. Adapted to local llms, vlm, gguf such as llama-3.3 Janus-Pro, Linkage graphRAG
A realtime voice runtime that keeps Agents talking, working, and present. Real-time Voice Runtime for AI Agents
Codex skill for converting slide images, PDFs, and image-based PPTX files into editable PowerPoint decks.
Let Claude (or any LLM) actually watch a video — scene-aware, deduplicated frames + transcript, from a URL or local file. Runs locally, MIT.
Connect your browser to AI models. Just use Dia on Chrome, Arc or Firefox.
A comprehensive ComfyUI integration for Microsoft's VibeVoice text-to-speech model, enabling high-quality single and multi-speaker voice synthesis directly within your ComfyUI workflows.
Natural voice conversations with Claude Code
Give Claude the ability to watch and understand videos — Claude Code plugin with frame extraction and multimodal audio analysis
PhoneClaw turns phones into local AI agent runtimes with on-device models, native mobile Skills, LiveLand, and optional Mac Gateway inference.
A ComfyUI custom node integration for local multi-engine multi-language Text-to-Speech and Voice Conversion. Supports: RVC, Echo-TTS, Qwen3-TTS, Cozy Voice 3, Step Audio EditX, IndexTTS-2, Chatterbox (classic and multilingual), F5-TTS, Higgs Audio 2, 3, and VibeVoice with unlimited text length, SRT timing, Character support, and many audio tools
AI speech toolkit for Apple Silicon — ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and CoreML
The Self-Coding System for Your App — Alan AI SDK for Cordova
Microsoft Text-to-Speech API sample code in several languages, part of Cognitive Services.
提供产品级IOCR自定义模板识别,以图搜图,人像搜索等,免费,可商用,Java AI 人工智能一站式解决方案,为工作减负,为产品研发加速。项目类别包括:以及AI SDK,web应用等。
Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.
Botium Speech Processing
SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.
Vonage REST API client for PHP. API support for SMS, Voice, Text-to-Speech, Numbers, Verify (2FA) and more.
TongFlow — Multimodal GenAI Studio
Voice-first local agent orchestration runtime for auditable DAG workflows.
Magick is a cutting-edge toolkit for a new kind of AI builder. Make Magick with us!
🔥🔥 LLaVA++: Extending LLaVA with Phi-3 and LLaMA-3 (LLaVA LLaMA-3, LLaVA Phi-3)
[dsh]为纯文本模型设计更强大的视觉工具箱:一行安装使用、粘贴图片直接识别、多张图片问答、截图到前端UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, grounding, pixel diff, Artifacts, and Web UI.
A real-time interactive Omni Avatar built on LiveKit, which allows you to seamlessly integrate with any open source Avatar components (real-time model, visual, voice, memory, search, etc.).
Rapida is an open-source, end-to-end voice AI orchestration platform for building real-time conversational voice agents with audio streaming, STT, TTS, VAD, multi-channel integration, agent state management, and observability.
Agentic ADK is an Agent application development framework launched by Alibaba International AI Business, based on Google-ADK and Ali-LangEngine.
Rust Agent Development Kit (ADK-Rust): Build AI agents in Rust with modular components for models, tools, memory, realtime voice, and more. ADK-Rust is a flexible framework for developing AI agents with simplicity and power. Model-agnostic, deployment-agnostic, optimized for frontier AI models. Includes support for real-time voice agents.
The Self-Coding System for Your App — Alan AI SDK for React Native
An MCP Server for Android running on the phone, optmized for token usage, supports also files downloads and cloudflare and ngrok automated tunnelling.
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents
A general framework for distilling human-created multimodal resources into reusable, executable skills that AI agents can browse, compose, and run, validated across diverse domains including web, PowerPoint, Excel, Blender, CAD, Unreal Engine 5, and REAPER-based music production.