GLM-4.6V-Flash
GLM-4.6V-Flash is Z.ai's lightweight (roughly 9-10B parameter) open-weight vision-language model, the small companion to the 106B GLM-4.6V, aimed at local deployment and low-latency use.
text + image → text · made by Z.ai
$0.3→$0.9per Mtok in / out
Hugging Face Inference Providers
Sold by 2 ways
| Seller | Lane | Rate |
|---|---|---|
| Hugging Face Inference Providers 2 ways | $0.3per Mtok in$0.9per Mtok out | |
| Hugging Face Inference Providers aggregator | standard | $0.3per Mtok in$0.9per Mtok outmodels.dev · read 2026-09-16 |
| Hugging Face Inference Providers aggregator | novita | $0.3per Mtok in$0.9per Mtok outrouter.huggingface.co · read 2026-08-25 |
About
GLM-4.6V-Flash — a text model from Z.ai, sold by one company from $0.3 in and $0.9 out per million tokens.
It takes text and images and returns text, with a context window of 128,000 tokens. It was published in December 2025, trained on material up to 2025. Its sellers say it can reason step by step and call a tool. The catalogue files it under chat. Only Hugging Face Inference Providers sells it, at $0.3 in and $0.9 out per million tokens. Beside the standard rate there is a separately routed lane.
Every current figure
- Maker
- Z.ai
- Register
- model
- Takes
- text + image
- Returns
- text
- Context
- 128,000 tokens
- Longest answer
- 64,000 tokens
- Published
- December 2025
- Knowledge to
- 2025
- Licence
- not read
- Sellers
- 1
- Maker's own price
- not read
- Price
- $0.3 in and $0.9 out per million tokens — Hugging Face Inference Providers
Known as 1 name
zai-org/GLM-4.6V-Flash