Implementation of DALL-E 2, OpenAI's updated text-to-image synthesis neural network, in Pytorch
LLMRepos·Vision & Multimodal Models LLM projects
Updated dailyBrowse 44 open-source vision & multimodal models projects in Models. Compare GitHub stars, recent growth, languages, licenses, and repository activity.
Vision & Multimodal Models
Browse 44 open-source vision & multimodal models projects in Models. Compare GitHub stars, recent growth, languages, licenses, and repository activity.
Top Vision & Multimodal Models repositories
Ranked by current GitHub stars from the latest LLMRepos snapshot.
Showing 40 of 44
Implementation of Imagen, Google's Text-to-Image Neural Network, in Pytorch
Implementation of Dreambooth (https://arxiv.org/abs/2208.12242) with Stable Diffusion
Implementation / replication of DALL-E, OpenAI's Text to Image Transformer, in Pytorch
Just playing with getting VQGAN+CLIP running locally, rather than having to use colab.
A simple command line tool for text to image generation, using OpenAI's CLIP and a BigGAN. Technique was originally created by https://twitter.com/advadnoun
The collection of pre-trained, state-of-the-art AI models for ailia SDK
Custom Diffusion: Multi-Concept Customization of Text-to-Image Diffusion (CVPR 2023)
Autoregressive Model Beats Diffusion: 🦙 Llama for Scalable Image Generation
[ICML 2024] Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs (RPG)
Text-to-Image generation. The repo for NeurIPS 2021 paper "CogView: Mastering Text-to-Image Generation via Transformers".
[ECCV 2024] The official implementation of paper "BrushNet: A Plug-and-Play Image Inpainting Model with Decomposed Dual-Branch Diffusion"
WiFi-3D-Fusion is an open-source research project that leverages WiFi CSI signals and deep learning to estimate 3D human pose, fusing wireless sensing with computer vision techniques for next-generation spatial awareness.
[ICCV 2025 Best Paper] Official repository for BrickGPT, the first approach for generating physically stable toy brick models from text prompts.
Generate images from texts. In Russian
[CVPR 2025 Oral]Infinity ∞ : Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
Adapting Meta AI's Segment Anything to Downstream Tasks with Adapters and Prompts
[ICCV 2025] 🔥🔥 UNO: A Universal Customization Method for Both Single and Multi-Subject Conditioning
[ECCV 2024 Oral] DriveLM: Driving with Graph Visual Question Answering
CogView4, CogView3-Plus and CogView3(ECCV 2024)
Text2Room generates textured 3D meshes from a given text prompt using 2D text-to-image models (ICCV2023).
Official Pytorch Implementation for "MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation" presenting "MultiDiffusion" (ICML 2023)
Pretrained model hub for Keras 3.
official code repo for paper "CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers"
Implementation of Muse: Text-to-Image Generation via Masked Generative Transformers, in Pytorch
[CVPR'23] OpenScene: 3D Scene Understanding with Open Vocabularies
🔥🔥 LLaVA++: Extending LLaVA with Phi-3 and LLaMA-3 (LLaVA LLaMA-3, LLaVA Phi-3)
Personal Project: MPP-Qwen14B & MPP-Qwen-Next(Multimodal Pipeline Parallel based on Qwen-LM). Support [video/image/multi-image] {sft/conversations}. Don't let the poverty limit your imagination! Train your own 8B/14B LLaVA-training-like MLLM on RTX3090/4090 24GB.
This repo contains the code for "VLM2Vec / MMEB" [ICLR 2025], "VLM2Vec-V2 / MMEB-V2" [TMLR 2026], and "MMEB-V3" [COLM 2026]
HunyuanImage-2.1: An Efficient Diffusion Model for High-Resolution (2K) Text-to-Image Generation
[SIGGRAPH Asia 2026] 4DAnyone: Create Anyone in 4D from a Casual Monocular Video
Official implementation of OneDiffusion paper (CVPR 2025)
[ICCV 2023] A latent space for stochastic diffusion models
基于Stable Diffusion优化的AI绘画模型。支持输入中英文文本,可生成多种现代艺术风格的高质量图像。| An optimized text-to-image model based on Stable Diffusion. Both Chinese and English text inputs are available to generate images. The model can generate high-quality images in several modern art styles.
The Clay Foundation Model - An open source AI model and interface for Earth
Vector Hub - Library for easy discovery, and consumption of State-of-the-art models to turn data into vectors. (text2vec, image2vec, video2vec, graph2vec, bert, inception, etc)
akanimax/T2F
T2F: text to face generation using Deep Learning
Official implementation of "En3D: An Enhanced Generative Model for Sculpting 3D Humans from 2D Synthetic Data", CVPR 2024; 3D Avatar Generation and Animation
Implementation of Parti, Google's pure attention-based text-to-image neural network, in Pytorch
Bria-AI/FIBO
FIBO is a SOTA, first open-source, JSON-native text-to-image model built for controllable, predictable, and legally safe image generation.