Pass IndexThe State of AISign in

vit-base-patch16-224

Vision Transformer (ViT) model pre-trained on ImageNet-21k (14 million images, 21,843 classes) at resolution 224x224, and fine-tuned on ImageNet 2012 (1 million images, 1,000 classes) at resolution 224x224.

image → text · made by Google

Sold by

Nobody in the catalogue publishes a price for this yet.

About

vit-base-patch16-224 — a text model from Google.

It takes images and returns text. It was published in March 2022. The catalogue files it under chat. Its weights are published under APACHE-2.0, so you may run it on your own machine, or buy it from whoever serves it cheapest.

Every current figure

Maker
Google
Register
model
Takes
image
Returns
text
Published
March 2022
Parameters
87 million · read from its own weights
Licence
apache-2.0

Known as 1 name

google/vit-base-patch16-224