Phi-4-multimodal-instruct
Phi-4-multimodal-instruct is a lightweight open multimodal foundation model that leverages the language, vision, and speech research and datasets used for Phi-3.5 and 4.0 models.
text → text · made by Microsoft
$0.05→$0.1per Mtok in / out
DeepInfra · 2 sellers
Sold by 2 ways
| Seller | Lane | Rate |
|---|---|---|
| Microsoft Azure AI cloud | standard | $0.08per Mtok in$0.32per Mtok out$4per Mtok in, audioraw.githubusercontent.com · read 2026-09-16 |
| DeepInfra aggregator | standard | $0.05per Mtok in$0.1per Mtok outapi.deepinfra.com · read 2026-08-25 |
Measured 1 standing
| Place | Board | Metric | Score |
|---|---|---|---|
| 24thof 65 | Open ASR Leaderboard · English short-form | avg (Average WER %) | 5.1 |
About
Phi-4-multimodal-instruct — a text model from Microsoft, sold by 2 companies from $0.05 in and $0.1 out per million tokens, placed 24th of 65 on Open ASR Leaderboard · English short-form.
It takes text and returns text, with a context window of 131,072 tokens. The catalogue files it under chat. Two companies sell it. The cheapest is $0.05 in and $0.1 out per million tokens at DeepInfra. It stands 24th of 65 on Open ASR Leaderboard · English short-form.
Every current figure
- Maker
- Microsoft
- Register
- model
- Takes
- text
- Returns
- text
- Context
- 131,072 tokens
- Licence
- not read
- Sellers
- 2
- Maker's own price
- not read
- Price
- $0.05 in and $0.1 out per million tokens — DeepInfra
- Boards
- 1
- Best place
- 24th of 65 — Open ASR Leaderboard · English short-form
Known as 1 name
microsoft/Phi-4-multimodal-instruct