Llama-3.3-70B-Instruct-Turbo
Llama 3.3-70B Turbo is a highly optimized version of the Llama 3.3-70B model, utilizing FP8 quantization to deliver significantly faster inference speeds with a minor trade-off in accuracy.
text → text · no single maker — the market carries it
Sold by 4 ways
| Seller | Lane | Rate |
|---|---|---|
| Requesty aggregatornot on the seller's list today | deepinfra flex | $0.072per Mtok in$0.23per Mtok outrouter.requesty.ai · read 2026-08-25 |
| Requesty aggregator | deepinfra | $0.12per Mtok in$0.3per Mtok out$0.12per Mtok cachedrouter.requesty.ai · read 2026-09-15 |
| DeepInfra aggregator | standard | $0.1per Mtok in$0.32per Mtok outapi.deepinfra.com · read 2026-08-25 |
| Together AI aggregator | standard | $1.04per Mtok in$1.04per Mtok outraw.githubusercontent.com · read 2026-09-16 |
Measured 1 standing
| Place | Board | Metric | Score |
|---|---|---|---|
| 5thof 105 | Vectara Hallucination Leaderboard | Hallucination Rate | 4.1 |
About
Llama-3.3-70B-Instruct-Turbo — a text model, sold by 3 companies from $0.1 in and $0.32 out per million tokens, placed 5th of 105 on Vectara Hallucination Leaderboard.
It takes text and returns text, with a context window of 131,072 tokens. It was published in December 2024. Its sellers say it can call a tool. The catalogue files it under translate. Three companies sell it. The cheapest is $0.1 in and $0.32 out per million tokens at DeepInfra. Beside the standard rate there is a separately routed lane. It stands 5th of 105 on Vectara Hallucination Leaderboard.
Every current figure
- Register
- model
- Takes
- text
- Returns
- text
- Context
- 131,072 tokens
- Longest answer
- 16,384 tokens
- Published
- December 2024
- Parameters
- 70 billion · read from its own name
- Sellers
- 3
- Price
- $0.1 in and $0.32 out per million tokens — DeepInfra
- Boards
- 1
- Best place
- 5th of 105 — Vectara Hallucination Leaderboard
Known as 3 names
meta-llama/Llama-3.3-70B-Instruct-Turbodeepinfra/meta-llama/Llama-3.3-70B-Instruct-Turbodeepinfra/meta-llama/Llama-3.3-70B-Instruct-Turbo:flex