Pass IndexThe State of AISign in

Gemini Robotics ER 2 Preview

A vision-language model (VLM) for robotics built on Gemini 3.5 Flash, with improved spatial reasoning, video moment finding, multi-robot orchestration, and multi-step tool use.

text + image + audio + video → text · made by Google

$1$5per Mtok in / out
Google

Sold by 2 ways

SellerLaneRate
Google 0 ways$1per Mtok in$5per Mtok out$0.1per Mtok cached
Google apinot on the seller's list todaybatch$1per Mtok in$5per Mtok out$0.1per Mtok cachedai.google.dev · read 2026-08-24
Google 0 ways$2per Mtok in$10per Mtok out$0.2per Mtok cached
Google apinot on the seller's list todaystandard$2per Mtok in$10per Mtok out$0.2per Mtok cachedai.google.dev · read 2026-08-24

About

Gemini Robotics ER 2 Preview — a text model from Google.

It takes text, images, audio and video and returns text, with a context window of 131,072 tokens. The catalogue files it under chat.

Every current figure

Maker
Google
Register
model
Takes
text + image + audio + video
Returns
text
Context
131,072 tokens
Licence
not read

Known as 1 name

gemini-robotics-er-2-preview