Dolphin
Dolphin (Document Image Parsing via Heterogeneous Anchor Prompting) is a novel multimodal document image parsing model that follows an analyze-then-parse paradigm.
text + image → text · made by ByteDance
Sold by
Nobody in the catalogue publishes a price for this yet.
About
Dolphin — a text model from ByteDance.
It takes text and images and returns text. It was published in May 2025. The catalogue files it under extract. Its weights are published under MIT, so you may run it on your own machine, or buy it from whoever serves it cheapest.
Every current figure
- Maker
- ByteDance
- Register
- model
- Takes
- text + image
- Returns
- text
- Published
- May 2025
- Parameters
- 398 million · read from its own weights
- Licence
- mit
Known as 1 name
ByteDance/Dolphin