Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
blanchefort
's Collections
Medical
VLA models
Audio
Translate
OCR
OmniModels
Edge models
Video encoders
Judge
Datasets for Embodied
Ru text encoders
Text2Image
VLMs
VLMs
updated
Mar 2
Upvote
-
Sort: Collection
Qwen/Qwen2-VL-7B-Instruct
Image-Text-to-Text
•
8B
•
Updated
Feb 6, 2025
•
1.52M
•
1.28k
NVEagle/Eagle-X5-13B-Chat
Image-Text-to-Text
•
15B
•
Updated
Sep 16, 2024
•
27
•
28
internlm/internlm-xcomposer2d5-7b
Visual Question Answering
•
Updated
Jul 22, 2024
•
886
•
210
AIRI-Institute/OmniFusion
Updated
Apr 10, 2024
•
59
OpenGVLab/InternVideo2_chat_8B_HD
Video-Text-to-Text
•
8B
•
Updated
Dec 18, 2024
•
1
•
18
OpenGVLab/InternVideo2-Chat-8B
Video-Text-to-Text
•
8B
•
Updated
Oct 10, 2024
•
60
•
26
zai-org/cogvlm2-video-llama3-chat
Text Generation
•
13B
•
Updated
Jul 24, 2024
•
129
•
56
nyu-visionx/cambrian-34b
Text Generation
•
35B
•
Updated
Jun 28, 2024
•
13
•
27
zai-org/cogvlm-base-490-hf
Text Generation
•
18B
•
Updated
Nov 20, 2023
•
109
•
7
zai-org/cogvlm-chat-hf
Text Generation
•
18B
•
Updated
Dec 19, 2023
•
696
•
199
zai-org/cogvlm-grounding-generalist-hf
Text Generation
•
18B
•
Updated
Dec 11, 2023
•
357
•
16
Qwen/Qwen-VL
Text Generation
•
Updated
Jan 25, 2024
•
9.78k
•
285
liuhaotian/llava-v1.5-7b
Image-Text-to-Text
•
Updated
May 8, 2024
•
248k
•
558
LanguageBind/MoE-LLaVA-Phi2-2.7B-4e-384
Text Generation
•
6B
•
Updated
Feb 1, 2024
•
19
•
32
LanguageBind/Video-LLaVA-7B-hf
Image-Text-to-Text
•
7B
•
Updated
May 16, 2024
•
15.3k
•
50
openvla/openvla-7b-prismatic
Image-Text-to-Text
•
Updated
Jul 9, 2024
•
323
•
9
openvla/openvla-7b-finetuned-libero-object
Image-Text-to-Text
•
8B
•
Updated
Oct 9, 2024
•
4.17k
•
2
openvla/openvla-7b-finetuned-libero-10
Image-Text-to-Text
•
8B
•
Updated
Oct 9, 2024
•
4.35k
•
7
IntelLabs/LlavaOLMoBitnet1B
Updated
Aug 30, 2024
•
3
•
31
mistral-community/pixtral-12b-240910
Image-Text-to-Text
•
Updated
Oct 1, 2024
•
18
•
380
LanguageBind/MoE-LLaVA-StableLM-1.6B-4e
Text Generation
•
3B
•
Updated
Feb 1, 2024
•
39
•
8
llava-hf/LLaVA-NeXT-Video-7B-hf
Video-Text-to-Text
•
7B
•
Updated
Nov 11, 2025
•
198k
•
126
Qwen/Qwen-VL-Chat
Text Generation
•
Updated
Jan 25, 2024
•
21.9k
•
384
LanguageBind/Video-LLaVA-7B
Text Generation
•
7B
•
Updated
Apr 9, 2024
•
525
•
89
LanguageBind/LanguageBind_Image
Zero-Shot Image Classification
•
Updated
Feb 1, 2024
•
17.7k
•
12
LanguageBind/LanguageBind_Video
Zero-Shot Image Classification
•
Updated
Feb 1, 2024
•
2.13k
•
4
llava-hf/llava-1.5-13b-hf
Image-Text-to-Text
•
13B
•
Updated
Jan 27, 2025
•
11.4k
•
35
llava-hf/llava-1.5-7b-hf
Image-Text-to-Text
•
7B
•
Updated
Jun 6, 2025
•
3.15M
•
368
FreedomIntelligence/LongLLaVA-53B-A13B
Image-Text-to-Text
•
52B
•
Updated
Nov 28, 2024
•
48
•
20
meta-llama/Llama-3.2-11B-Vision
Image-Text-to-Text
•
11B
•
Updated
Sep 27, 2024
•
8.76k
•
597
BAAI/Emu3-VisionTokenizer
Feature Extraction
•
0.3B
•
Updated
Oct 8, 2024
•
2.51k
•
63
openbmb/MiniCPM-V-2_6
Image-Text-to-Text
•
8B
•
Updated
Jun 13, 2025
•
21.9k
•
1.06k
openbmb/MiniCPM-V
Visual Question Answering
•
3B
•
Updated
Jan 15, 2025
•
623
•
209
openbmb/MiniCPM-V-2
Visual Question Answering
•
3B
•
Updated
Jan 15, 2025
•
77.5k
•
501
openbmb/MiniCPM-Llama3-V-2_5
Image-Text-to-Text
•
9B
•
Updated
Jan 15, 2025
•
18.1k
•
1.41k
nvidia/NVLM-D-72B
Image-Text-to-Text
•
79B
•
Updated
Jan 14, 2025
•
179k
•
776
vikhyatk/moondream2
Image-Text-to-Text
•
2B
•
Updated
Sep 23, 2025
•
2.62M
•
1.43k
allenai/Molmo-72B-0924
Image-Text-to-Text
•
73B
•
Updated
Oct 9, 2025
•
5.45k
•
299
allenai/MolmoE-1B-0924
Image-Text-to-Text
•
Updated
Apr 24, 2025
•
1.49k
•
157
allenai/Molmo-7B-D-0924
Image-Text-to-Text
•
8B
•
Updated
Dec 15, 2025
•
22k
•
566
allenai/Molmo-7B-O-0924
Image-Text-to-Text
•
8B
•
Updated
Oct 9, 2025
•
1.27k
•
164
deepseek-ai/Janus-1.3B
Any-to-Any
•
2B
•
Updated
Jan 27, 2025
•
2.14k
•
598
neulab/Pangea-7B
8B
•
Updated
Oct 24, 2024
•
3.04k
•
133
neulab/Pangea-7B-hf
8B
•
Updated
Oct 28, 2025
•
1.33k
•
13
BAAI/Aquila-VL-2B-llava-qwen
Visual Question Answering
•
2B
•
Updated
24 days ago
•
127
•
61
mistralai/Pixtral-Large-Instruct-2411
Updated
Jun 2
•
67
•
434
google/paligemma2-10b-pt-224
Image-Text-to-Text
•
10B
•
Updated
Dec 5, 2024
•
1.08k
•
10
google/paligemma2-3b-pt-224
Image-Text-to-Text
•
3B
•
Updated
Dec 5, 2024
•
16.8k
•
176
vidore/colqwen2-v1.0
Visual Document Retrieval
•
Updated
Jun 5, 2025
•
38.7k
•
121
deepseek-ai/Janus-Pro-7B
Any-to-Any
•
Updated
Feb 1, 2025
•
13.2k
•
3.64k
deepseek-ai/Janus-Pro-1B
Any-to-Any
•
Updated
Feb 1, 2025
•
14.5k
•
482
nvidia/Eagle2-9B
Image-Text-to-Text
•
9B
•
Updated
Jan 28, 2025
•
292
•
63
openbmb/MiniCPM-o-2_6
Any-to-Any
•
9B
•
Updated
Oct 5, 2025
•
324k
•
1.3k
DAMO-NLP-SG/VideoLLaMA3-7B
Video-Text-to-Text
•
8B
•
Updated
Sep 2, 2025
•
9.32k
•
76
DAMO-NLP-SG/VideoLLaMA3-2B
Video-Text-to-Text
•
2B
•
Updated
Sep 3, 2025
•
1.84k
•
21
ATH-MaaS/Ovis2-8B
Image-Text-to-Text
•
9B
•
Updated
Aug 15, 2025
•
1.05k
•
74
Qwen/Qwen3-VL-2B-Thinking
Image-Text-to-Text
•
2B
•
Updated
Oct 20, 2025
•
79.6k
•
115
LiquidAI/LFM2-VL-3B
Image-Text-to-Text
•
3B
•
Updated
Mar 30
•
12.9k
•
140
facebook/sam3
Mask Generation
•
0.9B
•
Updated
Nov 20, 2025
•
2.29M
•
2.59k
stepfun-ai/Step3-VL-10B-FP8
Image-Text-to-Text
•
Updated
Feb 4
•
830
•
10
nvidia/llama-nemotron-colembed-vl-3b-v2
Visual Document Retrieval
•
4B
•
Updated
May 20
•
1.83k
•
22
Upvote
-
Sort: Collection
Share collection
View history
Collection guide
Browse collections