NVIDIA Nemotron 3.5 Lightning 30B A3B
A hybrid Mixture-of-Experts LLM from NVIDIA with 30B total / 3B active parameters, using interleaved Mamba-2, MoE and attention layers, intended primarily for customization and post-training such as fine-tuning, RL, domain adaptation and quantization.
text → text · made by NVIDIA
Sold by 3 ways
| Seller | Lane | Rate |
|---|---|---|
| Thinking Machines Lab api | training | $0.2per Mtok in$0.49per Mtok out$0.039per Mtok cachedtinker-docs.thinkingmachines.ai · read 2026-08-25 |
| Thinking Machines Lab api | training-256k | $0.26per Mtok in$0.66per Mtok out$0.052per Mtok cachedtinker-docs.thinkingmachines.ai · read 2026-08-25 |
| Fireworks AI aggregator | standard | $0.05per Mtok in$0.2per Mtok out$0.01per Mtok cacheddocs.fireworks.ai · read 2026-08-24 |
About
NVIDIA Nemotron 3.5 Lightning 30B A3B — a text model from NVIDIA, sold by 2 companies from $0.05 in and $0.2 out per million tokens.
It takes text and returns text, with a context window of 1,048,576 tokens. It was published in August 2026. Its sellers say it can reason step by step and call a tool. The catalogue files it under chat. Two companies sell it. The cheapest is $0.05 in and $0.2 out per million tokens at Fireworks AI. Beside the standard rate there is a separately routed lane.
Every current figure
- Maker
- NVIDIA
- Register
- model
- Takes
- text
- Returns
- text
- Context
- 1,048,576 tokens
- Longest answer
- 262,144 tokens
- Published
- August 2026
- Parameters
- 32 billion
- Licence
- not read
- Sellers
- 2
- Maker's own price
- not read
- Price
- $0.05 in and $0.2 out per million tokens — Fireworks AI
Known as 3 names
fireworks/nemotron-lightning-3p5-30b-a3bnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16:peft:262144