Text + image to text models, tools and agents
332 in the catalogue today. Every one with what it costs, who sells it and where it stands. 332 in all, a hundred to a page; this is page 4 of 4.
A picture and a question in, an answer in words out. This is what makes a model useful on a screenshot, a chart, a scanned page or a photograph of a whiteboard, and it is now the default shape for a capable general model.
- kosmos-2.5Microsoft
- liftDatalab
- llama-joycaption-beta-one-hf-llavafancyfeast
- llava-1.5-7b-hfLlava Hugging Face
- llava-v1.5-13bliuhaotian
- llava-v1.5-7bliuhaotian19th of 26
- llava-v1.6-34bliuhaotian
- llava-v1.6-mistral-7bliuhaotian10th of 26
- llava-v1.6-mistral-7b-hfLlava Hugging Face
- llava-v1.6-vicuna-7bliuhaotian13th of 26
- medgemma-1.5-4b-itGoogle
- medgemma-27b-itGoogle
- medgemma-4b-itGoogle
- medgemma-4b-ptGoogle
- moondream2vikhyatk
- moondream3-previewmoondream
- olmOCR-7B-0225-previewAi2
- paligemma-3b-pt-224Google
- paligemma2-3b-pt-224Google
- pixtral-12b-240910Unofficial Mistral Community [deprecated]
- proxy-lite-3bconvergence-ai
- qwen3-vl-235b-a22bAlibaba$0.7→$2.8per Mtok in / out2 selling
- shieldgemma-2-4b-itGoogle
- smol-visionmerve
- step3stepfun-ai
- t5gemma-2-270m-270mGoogle
- t5gemma-2-4b-4bGoogle
- translategemma-12b-itGoogle
- translategemma-27b-itGoogle
- translategemma-4b-itGoogle
- vilt-b32-finetuned-vqadandelin
- xgen-mm-phi3-mini-instruct-r-v1Salesforce