Pass IndexThe State of AISign in

Phi-3.5-vision-instruct

Phi-3.5-vision is a lightweight, state-of-the-art open multimodal model built upon datasets which include - synthetic data and filtered publicly available websites - with a focus on very high-quality, reasoning dense data both on text and vision.

text + image → text · made by Microsoft

$0.13$0.52per Mtok in / out
Microsoft Azure AI

Sold by 1 way

SellerLaneRate
Microsoft Azure AI cloudstandard$0.13per Mtok in$0.52per Mtok outraw.githubusercontent.com · read 2026-09-16

Measured 2 standings

PlaceBoardMetricScore
1stof 26Science qa — Epoch AIScore91.3
151stof 152LMArena · VisionScore (Elo)920

About

Phi-3.5-vision-instruct — a text model from Microsoft, sold by one company from $0.13 in and $0.52 out per million tokens, placed 1st of 26 on Science qa — Epoch AI.

It takes text and images and returns text. It was published in August 2024. The catalogue files it under chat. Its weights are published under MIT, so you may run it on your own machine, or buy it from whoever serves it cheapest. Only Microsoft Azure AI sells it, at $0.13 in and $0.52 out per million tokens. It has been measured on 2 boards, and stands best at 1st of 26 on Science qa — Epoch AI.

Every current figure

Maker
Microsoft
Register
model
Takes
text + image
Returns
text
Published
August 2024
Parameters
4.1 billion · read from its own weights
Licence
mit
Sellers
1
Maker's own price
not read
Price
$0.13 in and $0.52 out per million tokens — Microsoft Azure AI
Boards
2
Best place
1st of 26 — Science qa — Epoch AI

Known as 1 name

microsoft/Phi-3.5-vision-instruct