Pass IndexThe State of AISign in

Llama-3.3-70B-Instruct-Turbo

Llama 3.3-70B Turbo is a highly optimized version of the Llama 3.3-70B model, utilizing FP8 quantization to deliver significantly faster inference speeds with a minor trade-off in accuracy.

text → text · no single maker — the market carries it

$0.1$0.32per Mtok in / out
DeepInfra · 3 sellers

Sold by 4 ways

SellerLaneRate
Requesty aggregatornot on the seller's list todaydeepinfra flex$0.072per Mtok in$0.23per Mtok outrouter.requesty.ai · read 2026-08-25
Requesty aggregatordeepinfra$0.12per Mtok in$0.3per Mtok out$0.12per Mtok cachedrouter.requesty.ai · read 2026-09-15
DeepInfra aggregatorstandard$0.1per Mtok in$0.32per Mtok outapi.deepinfra.com · read 2026-08-25
Together AI aggregatorstandard$1.04per Mtok in$1.04per Mtok outraw.githubusercontent.com · read 2026-09-16

Measured 1 standing

PlaceBoardMetricScore
5thof 105Vectara Hallucination LeaderboardHallucination Rate4.1

About

Llama-3.3-70B-Instruct-Turbo — a text model, sold by 3 companies from $0.1 in and $0.32 out per million tokens, placed 5th of 105 on Vectara Hallucination Leaderboard.

It takes text and returns text, with a context window of 131,072 tokens. It was published in December 2024. Its sellers say it can call a tool. The catalogue files it under translate. Three companies sell it. The cheapest is $0.1 in and $0.32 out per million tokens at DeepInfra. Beside the standard rate there is a separately routed lane. It stands 5th of 105 on Vectara Hallucination Leaderboard.

Every current figure

Register
model
Takes
text
Returns
text
Context
131,072 tokens
Longest answer
16,384 tokens
Published
December 2024
Parameters
70 billion · read from its own name
Sellers
3
Price
$0.1 in and $0.32 out per million tokens — DeepInfra
Boards
1
Best place
5th of 105 — Vectara Hallucination Leaderboard

Known as 3 names

meta-llama/Llama-3.3-70B-Instruct-Turbodeepinfra/meta-llama/Llama-3.3-70B-Instruct-Turbodeepinfra/meta-llama/Llama-3.3-70B-Instruct-Turbo:flex