Pass · production benchmark

999 real messages,
ten languages, one Desk.

Three user profiles ran in parallel against production Pass. Every prompt was unique, so nothing came from cache. Then the same questions went to Anthropic's Opus 5 — today's frontier model — and an independent judge scored both, blind.
2026-08-19997 delivered · 2 network errors0 cache hits10 languages
6.68 s
median answer
14.9 s
9 out of 10 under
$0.000299
per message, all-in
997/999
delivered

Three user profiles

Personal
Trips, recipes, family messages, explanations for kids, quick translations.
Median5.81 sOpus 20.7 s
Cost / msg$0.000199Opus $0.0160
Quality80.2Opus 88.1
80× cheaper3.6× faster91% quality
Business
Client emails, quarterly agendas, risk analysis, hiring trade-offs, reports.
Median7.44 sOpus 25.1 s
Cost / msg$0.000273Opus $0.0197
Quality78.6Opus 85.1
72× cheaper3.4× faster92% quality
Coder
Algorithms, SQL, API reviews — plus answers drawn from the account's knowledge base.
Median6.93 sOpus 23.4 s
Cost / msg$0.000424Opus $0.0242
Quality81.6Opus 87.9
57× cheaper3.4× faster93% quality

How the Desk routed the work

72%answered directly

The Desk's small model reads every message — in any language — and decides on the spot: answer it itself, or hand it to the row of the preset that fits. It handled 725 of 997 messages alone and dispatched 272.

answered directly 725
code 89
default 73
analysis 71
longform 31
small 8

Where the time went

35
<3 s
389
3–6 s
332
6–10 s
142
10–15 s
80
15–30 s
19
30 s+

Ten languages, no language rules

LanguageMedianQuality vs Opus 5
Hindi6.4 s
109%
Russian6.2 s
101%
Japanese6.9 s
95%
Arabic6.5 s
92%
French7.3 s
92%
Chinese7.8 s
91%
English7.0 s
91%
Spanish7.1 s
87%
Portuguese6.6 s
85%
German5.8 s
85%

Routing decisions were made on messages written in English, Chinese, Hindi, Spanish, Arabic, French, Portuguese, Russian, Japanese and German — with no language-specific rules anywhere in the system.

Pass V1.0 vs Opus 5

Pass V1.0
$0.000299
per message · $0.30 per 1000
6.68 s
median answer
80.1
blind quality score
VS
Opus 5 · frontier
$0.019924
per message · $19.92 per 1000
23.31 s
median answer
87.0
blind quality score
66.7× cheaper
$0.30 vs $19.92 per 1000 messages
3.5× faster
6.68 s vs 23.31 s median
92.0%
of the quality, blind-scored
264
answer pairs judged

On everyday questions Pass is level with the frontier and sometimes ahead; it gives ground on demanding, structured work. That is what the remaining 8% is — bought back at one 67th of the price. Model selection follows the published preset.

Appendix — orchestration vs inference

Every step of the pipeline is metered separately. The Desk's routing decision is orchestration; the model that writes the answer is inference. Keeping them apart means the inference figure is the true price of the answer itself, while the cost of running the pipeline stays on its own line. As the pipeline grows into multi-step jobs, every step will be metered and summed the same way.

LayerCallsTotalPer messageShare
Orchestration999$0.0161$0.0000165.4%
Inference997$0.2816$0.00028294.6%
Total$0.2977$0.000299100%

Orchestration adds 5.4% on top of inference — under two hundredths of a cent per message — and it is what lets a small, cheap model handle three quarters of the traffic instead of an expensive one. All response times quoted above already include this step, end to end.