The Desk's small model reads every message — in any language — and decides on the spot: answer it itself, or hand it to the row of the preset that fits. It handled 725 of 997 messages alone and dispatched 272.
| Language | Median | Quality vs Opus 5 | |
|---|---|---|---|
| Hindi | 6.4 s | 109% | |
| Russian | 6.2 s | 101% | |
| Japanese | 6.9 s | 95% | |
| Arabic | 6.5 s | 92% | |
| French | 7.3 s | 92% | |
| Chinese | 7.8 s | 91% | |
| English | 7.0 s | 91% | |
| Spanish | 7.1 s | 87% | |
| Portuguese | 6.6 s | 85% | |
| German | 5.8 s | 85% |
Routing decisions were made on messages written in English, Chinese, Hindi, Spanish, Arabic, French, Portuguese, Russian, Japanese and German — with no language-specific rules anywhere in the system.
On everyday questions Pass is level with the frontier and sometimes ahead; it gives ground on demanding, structured work. That is what the remaining 8% is — bought back at one 67th of the price. Model selection follows the published preset.
Every step of the pipeline is metered separately. The Desk's routing decision is orchestration; the model that writes the answer is inference. Keeping them apart means the inference figure is the true price of the answer itself, while the cost of running the pipeline stays on its own line. As the pipeline grows into multi-step jobs, every step will be metered and summed the same way.
| Layer | Calls | Total | Per message | Share |
|---|---|---|---|---|
| Orchestration | 999 | $0.0161 | $0.000016 | 5.4% |
| Inference | 997 | $0.2816 | $0.000282 | 94.6% |
| Total | — | $0.2977 | $0.000299 | 100% |
Orchestration adds 5.4% on top of inference — under two hundredths of a cent per message — and it is what lets a small, cheap model handle three quarters of the traffic instead of an expensive one. All response times quoted above already include this step, end to end.