Phi-3.5-vision-instruct
Phi-3.5-vision is a lightweight, state-of-the-art open multimodal model built upon datasets which include - synthetic data and filtered publicly available websites - with a focus on very high-quality, reasoning dense data both on text and vision.
text + image → text · made by Microsoft
Sold by 1 way
| Seller | Lane | Rate |
|---|---|---|
| Microsoft Azure AI cloud | standard | $0.13per Mtok in$0.52per Mtok outraw.githubusercontent.com · read 2026-09-16 |
Measured 2 standings
| Place | Board | Metric | Score |
|---|---|---|---|
| 1stof 26 | Science qa — Epoch AI | Score | 91.3 |
| 151stof 152 | LMArena · Vision | Score (Elo) | 920 |
About
Phi-3.5-vision-instruct — a text model from Microsoft, sold by one company from $0.13 in and $0.52 out per million tokens, placed 1st of 26 on Science qa — Epoch AI.
It takes text and images and returns text. It was published in August 2024. The catalogue files it under chat. Its weights are published under MIT, so you may run it on your own machine, or buy it from whoever serves it cheapest. Only Microsoft Azure AI sells it, at $0.13 in and $0.52 out per million tokens. It has been measured on 2 boards, and stands best at 1st of 26 on Science qa — Epoch AI.
Every current figure
- Maker
- Microsoft
- Register
- model
- Takes
- text + image
- Returns
- text
- Published
- August 2024
- Parameters
- 4.1 billion · read from its own weights
- Licence
- mit
- Sellers
- 1
- Maker's own price
- not read
- Price
- $0.13 in and $0.52 out per million tokens — Microsoft Azure AI
- Boards
- 2
- Best place
- 1st of 26 — Science qa — Epoch AI
Known as 1 name
microsoft/Phi-3.5-vision-instruct