Mage-VL
Mage-VL is a codec-native, proactive-streaming multimodal foundation model for image and video understanding, whose visual encoder is trained entirely from scratch at a compact 4B scale.
text + image → text · made by Microsoft
Sold by
Nobody in the catalogue publishes a price for this yet.
About
Mage-VL — a text model from Microsoft.
It takes text and images and returns text. It was published in July 2026. The catalogue files it under chat. Its weights are published under APACHE-2.0, so you may run it on your own machine, or buy it from whoever serves it cheapest.
Every current figure
- Maker
- Microsoft
- Register
- model
- Takes
- text + image
- Returns
- text
- Published
- July 2026
- Parameters
- 4.7 billion · read from its own weights
- Licence
- apache-2.0
Known as 1 name
microsoft/Mage-VL