Pass IndexThe State of AISign in

Multimodal

A model that takes or returns more than text — images, audio, video, files.

how it works · the vocabulary

visionomni

Modalities are priced apart because they cost apart: an image becomes hundreds or thousands of tokens depending on its size, and audio is usually billed by the second instead. This site records what each thing takes and returns as its own field for that reason.

In the catalogue

Nearby

AgentAgent memoryAlignmentAutoregressiveBM25ChunkingCold startComputer use