Pass IndexThe State of AISign in

NVIDIA Nemotron 3.5 Lightning 30B A3B

A hybrid Mixture-of-Experts LLM from NVIDIA with 30B total / 3B active parameters, using interleaved Mamba-2, MoE and attention layers, intended primarily for customization and post-training such as fine-tuning, RL, domain adaptation and quantization.

text → text · made by NVIDIA

$0.05$0.2per Mtok in / out
Fireworks AI · 2 sellers

Sold by 3 ways

SellerLaneRate
Thinking Machines Lab apitraining$0.2per Mtok in$0.49per Mtok out$0.039per Mtok cachedtinker-docs.thinkingmachines.ai · read 2026-08-25
Thinking Machines Lab apitraining-256k$0.26per Mtok in$0.66per Mtok out$0.052per Mtok cachedtinker-docs.thinkingmachines.ai · read 2026-08-25
Fireworks AI aggregatorstandard$0.05per Mtok in$0.2per Mtok out$0.01per Mtok cacheddocs.fireworks.ai · read 2026-08-24

About

NVIDIA Nemotron 3.5 Lightning 30B A3B — a text model from NVIDIA, sold by 2 companies from $0.05 in and $0.2 out per million tokens.

It takes text and returns text, with a context window of 1,048,576 tokens. It was published in August 2026. Its sellers say it can reason step by step and call a tool. The catalogue files it under chat. Two companies sell it. The cheapest is $0.05 in and $0.2 out per million tokens at Fireworks AI. Beside the standard rate there is a separately routed lane.

Every current figure

Maker
NVIDIA
Register
model
Takes
text
Returns
text
Context
1,048,576 tokens
Longest answer
262,144 tokens
Published
August 2026
Parameters
32 billion
Licence
not read
Sellers
2
Maker's own price
not read
Price
$0.05 in and $0.2 out per million tokens — Fireworks AI

Known as 3 names

fireworks/nemotron-lightning-3p5-30b-a3bnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16:peft:262144