By use case · Vision
Best local vision (multimodal) models
These models read images as input, so you can ask about a screenshot or a photo, not only text. 6 open vision (multimodal) models are ranked by quality below. The strongest is Qwen3-VL 30B-A3B (~20.4 GB at Q4_K_M); the lightest is Qwen3-VL 4B (~4.4 GB).
Models that accept images as input alongside text.
- 1 ~20.4 GBQwen3-VL 30B-A3B31.1B MoE · runs on 13/43 devices
- 2 ~9 GBLlama 3.2 Vision 11B10.7B · runs on 29/43 devices
- 3 ~7.3 GBQwen3-VL 8B8.77B · runs on 33/43 devices
- 4 ~7.1 GBQwen2.5-VL 7B8.29B · runs on 33/43 devices
- 5 ~4.4 GBQwen3-VL 4B4.44B · runs on 43/43 devices
- 6 ~4.4 GBQwen2.5-VL 3B3.75B · runs on 43/43 devices
Tagged by each model's stated purpose. The memory figure is what it needs at Q4_K_M, and the device beside it is the lightest tracked machine that fits it. "Runs on N devices" counts the 43 tracked devices that fit it at Q4_K_M.
FAQ
What is the best local vision (multimodal) model?
Qwen3-VL 30B-A3B leads on quality among the open vision (multimodal) models tracked here. It needs ~20.4 GB at Q4_K_M, so the lightest hardware that runs it is Apple M5 (32GB). Pick by what fits your memory using the list above.
What is the smallest vision (multimodal) model that runs on a laptop?
Qwen3-VL 4B is the lightest, at ~4.4 GB at Q4_K_M, so any device with 8 GB or more can load it.
How were these vision (multimodal) models chosen?
They are open-weight models whose own design targets vision (multimodal) (by name, family or model card). Memory figures are computed at Q4_K_M and sourced; see the methodology page.
Memory is computed at Q4_K_M; catalog updated 2026-10-05. See methodology.