LLMRepos·Vision & Multimodal Models LLM projects

Updated daily

Browse 44 open-source vision & multimodal models projects in Models. Compare GitHub stars, recent growth, languages, licenses, and repository activity.

Vision & Multimodal Models

Browse 44 open-source vision & multimodal models projects in Models. Compare GitHub stars, recent growth, languages, licenses, and repository activity.

44
Repositories
144
Models

Top Vision & Multimodal Models repositories

Ranked by current GitHub stars from the latest LLMRepos snapshot.

Showing 40 of 44

Implementation of DALL-E 2, OpenAI's updated text-to-image synthesis neural network, in Pytorch

PythonMIT License+1 stars in 7dupdated 835d ago

Implementation of Imagen, Google's Text-to-Image Neural Network, in Pytorch

PythonMIT License+1 stars in 7dupdated 686d ago

Implementation of Dreambooth (https://arxiv.org/abs/2208.12242) with Stable Diffusion

Jupyter NotebookMIT License+1 stars in 7dupdated 1,356d ago

Implementation / replication of DALL-E, OpenAI's Text to Image Transformer, in Pytorch

PythonMIT License+0 stars in 7dupdated 919d ago

Just playing with getting VQGAN+CLIP running locally, rather than having to use colab.

PythonOther-2 stars in 7dupdated 1,422d ago

A simple command line tool for text to image generation, using OpenAI's CLIP and a BigGAN. Technique was originally created by https://twitter.com/advadnoun

PythonMIT License+1 stars in 7dupdated 1,660d ago

The collection of pre-trained, state-of-the-art AI models for ailia SDK

PythonOther+17 stars in 7dupdated today

Custom Diffusion: Multi-Concept Customization of Text-to-Image Diffusion (CVPR 2023)

PythonOther+0 stars in 7dupdated 92d ago

Autoregressive Model Beats Diffusion: 🦙 Llama for Scalable Image Generation

PythonMIT License+0 stars in 7dupdated 739d ago

[ICML 2024] Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs (RPG)

Jupyter NotebookMIT License+0 stars in 7dupdated 569d ago
1.8k
stars

Text-to-Image generation. The repo for NeurIPS 2021 paper "CogView: Mastering Text-to-Image Generation via Transformers".

PythonApache License 2.0+0 stars in 7dupdated 1,064d ago

[ECCV 2024] The official implementation of paper "BrushNet: A Plug-and-Play Image Inpainting Model with Decomposed Dual-Branch Diffusion"

PythonOther+2 stars in 7dupdated 615d ago

WiFi-3D-Fusion is an open-source research project that leverages WiFi CSI signals and deep learning to estimate 3D human pose, fusing wireless sensing with computer vision techniques for next-generation spatial awareness.

PythonOther+5 stars in 7dupdated 363d ago

[ICCV 2025 Best Paper] Official repository for BrickGPT, the first approach for generating physically stable toy brick models from text prompts.

PythonMIT License+4 stars in 7dupdated 95d ago

Generate images from texts. In Russian

Jupyter NotebookApache License 2.0+0 stars in 7dupdated 1,322d ago

[CVPR 2025 Oral]Infinity ∞ : Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis

PythonMIT License+0 stars in 7dupdated 130d ago

Adapting Meta AI's Segment Anything to Downstream Tasks with Adapters and Prompts

PythonMIT License+2 stars in 7dupdated 99d ago
1.4k
stars

[ICCV 2025] 🔥🔥 UNO: A Universal Customization Method for Both Single and Multi-Subject Conditioning

PythonApache License 2.0+1 stars in 7dupdated 346d ago

[ECCV 2024 Oral] DriveLM: Driving with Graph Visual Question Answering

HTMLApache License 2.0-3 stars in 7dupdated 418d ago
1.1k
stars

CogView4, CogView3-Plus and CogView3(ECCV 2024)

PythonApache License 2.0+0 stars in 7dupdated 513d ago

Text2Room generates textured 3D meshes from a given text prompt using 2D text-to-image models (ICCV2023).

PythonMIT License+2 stars in 7dupdated 1,013d ago

Official Pytorch Implementation for "MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation" presenting "MultiDiffusion" (ICML 2023)

Jupyter Notebook+1 stars in 7dupdated 1,068d ago

Pretrained model hub for Keras 3.

PythonApache License 2.0+1 stars in 7dupdated today
953
stars

official code repo for paper "CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers"

PythonApache License 2.0-2 stars in 7dupdated 1,482d ago

Implementation of Muse: Text-to-Image Generation via Masked Generative Transformers, in Pytorch

PythonMIT License+0 stars in 7dupdated 907d ago

[CVPR'23] OpenScene: 3D Scene Understanding with Open Vocabularies

PythonApache License 2.0-2 stars in 7dupdated 1,032d ago

🔥🔥 LLaVA++: Extending LLaVA with Phi-3 and LLaMA-3 (LLaVA LLaMA-3, LLaVA Phi-3)

Python+1 stars in 7dupdated 385d ago
685
stars

Personal Project: MPP-Qwen14B & MPP-Qwen-Next(Multimodal Pipeline Parallel based on Qwen-LM). Support [video/image/multi-image] {sft/conversations}. Don't let the poverty limit your imagination! Train your own 8B/14B LLaVA-training-like MLLM on RTX3090/4090 24GB.

Jupyter Notebook-1 stars in 7dupdated 532d ago

This repo contains the code for "VLM2Vec / MMEB" [ICLR 2025], "VLM2Vec-V2 / MMEB-V2" [TMLR 2026], and "MMEB-V3" [COLM 2026]

PythonApache License 2.0+4 stars in 7dupdated 1d ago

HunyuanImage-2.1: An Efficient Diffusion Model for High-Resolution (2K) Text-to-Image Generation​

PythonOther+0 stars in 7dupdated 314d ago

[SIGGRAPH Asia 2026] 4DAnyone: Create Anyone in 4D from a Casual Monocular Video

PythonApache License 2.0updated today

Official implementation of OneDiffusion paper (CVPR 2025)

PythonOther+0 stars in 7dupdated 618d ago

[ICCV 2023] A latent space for stochastic diffusion models

PythonOther+0 stars in 7dupdated 967d ago

基于Stable Diffusion优化的AI绘画模型。支持输入中英文文本,可生成多种现代艺术风格的高质量图像。| An optimized text-to-image model based on Stable Diffusion. Both Chinese and English text inputs are available to generate images. The model can generate high-quality images in several modern art styles.

MIT License+0 stars in 7dupdated 1,264d ago

The Clay Foundation Model - An open source AI model and interface for Earth

PythonApache License 2.0+2 stars in 7dupdated 105d ago

Vector Hub - Library for easy discovery, and consumption of State-of-the-art models to turn data into vectors. (text2vec, image2vec, video2vec, graph2vec, bert, inception, etc)

+0 stars in 7dupdated 735d ago
546
stars

T2F: text to face generation using Deep Learning

PythonMIT License+0 stars in 7dupdated 1,563d ago
543
stars

Official implementation of "En3D: An Enhanced Generative Model for Sculpting 3D Humans from 2D Synthetic Data", CVPR 2024; 3D Avatar Generation and Animation

PythonApache License 2.0+0 stars in 7dupdated 637d ago

Implementation of Parti, Google's pure attention-based text-to-image neural network, in Pytorch

PythonMIT License+0 stars in 7dupdated 990d ago
334
stars

FIBO is a SOTA, first open-source, JSON-native text-to-image model built for controllable, predictable, and legally safe image generation.

PythonOther+5 stars in 7dupdated 229d ago