Vision-language model
A model that reads images alongside text — screenshots, documents, charts, photographs.
how it works · the vocabulary
VLMmultimodal modelvision
Images are converted into tokens, so a high-resolution page can cost as much as several thousand words. Most document processing that used to need OCR and a layout engine is now one call to one of these.