vilt-b32-finetuned-vqa
Vision-and-Language Transformer (ViLT) model fine-tuned on VQAv2.
text + image → text · made by dandelin
Sold by
Nobody in the catalogue publishes a price for this yet.
About
vilt-b32-finetuned-vqa — a text model from dandelin.
It takes text and images and returns text. It was published in March 2022. The catalogue files it under chat. Its weights are published under APACHE-2.0, so you may run it on your own machine, or buy it from whoever serves it cheapest.
Every current figure
- Maker
- dandelin
- Register
- model
- Takes
- text + image
- Returns
- text
- Published
- March 2022
- Licence
- apache-2.0
Known as 1 name
dandelin/vilt-b32-finetuned-vqa