Gemini Robotics ER 2 Preview
A vision-language model (VLM) for robotics built on Gemini 3.5 Flash, with improved spatial reasoning, video moment finding, multi-robot orchestration, and multi-step tool use.
text + image + audio + video → text · made by Google
$1→$5per Mtok in / out
Google
Sold by 2 ways
| Seller | Lane | Rate |
|---|---|---|
| Google 0 ways | $1per Mtok in$5per Mtok out$0.1per Mtok cached | |
| Google apinot on the seller's list today | batch | $1per Mtok in$5per Mtok out$0.1per Mtok cachedai.google.dev · read 2026-08-24 |
| Google 0 ways | $2per Mtok in$10per Mtok out$0.2per Mtok cached | |
| Google apinot on the seller's list today | standard | $2per Mtok in$10per Mtok out$0.2per Mtok cachedai.google.dev · read 2026-08-24 |
About
Gemini Robotics ER 2 Preview — a text model from Google.
It takes text, images, audio and video and returns text, with a context window of 131,072 tokens. The catalogue files it under chat.
Every current figure
- Maker
- Register
- model
- Takes
- text + image + audio + video
- Returns
- text
- Context
- 131,072 tokens
- Licence
- not read
Known as 1 name
gemini-robotics-er-2-preview