Three user profiles ran against production Pass — every request compiled by the small model: route, goal, language, what to pull. A 90-message sample went to Anthropic's Opus 5, and an independent judge scored both blind, two ways.
The judge (gemini-3.6-flash, judge-v1 prompt) scored blind twice. Scoring each answer alone on a 0–100 rubric — correctness, then completeness, then clarity — Pass averages 87.2 to Opus's 78.8. Asked head-to-head which answer an expert would prefer, the judge picked Opus 55–23. Both metrics are real: the rubric rewards being right and clear; the head-to-head rewards depth and length. Both answers are truncated to 1400 characters for the judge, and Opus answers run longer (median 2002 vs 1185 chars), so truncation costs Opus more in the rubric and buys it preference in the duel.
Every request produced a Job: the route (analysis 104 · longform 98 · code 50 · small 11 · answered itself 37), the goal in the compiler's own words (300 of 300), and the language — ten languages, thirty messages each, no per-language rules anywhere in the system.
300 unique prompts — 100 per profile, 10 per language — against production, cache off. A 90-pair sample (3 random per profile×language) went to Opus 5 with identical prompts, max_tokens 2000. Pass costs are metered per request in the ledger (inference only; Pass's $0.01 service fee is separate); Opus costs are OpenRouter's billed usage. 78 of 90 pairs returned complete verdicts. A repeat run the same day reproduced the picture: quality 112%, duel 28–56, 24 minutes wall-clock for the full pipeline. The previous benchmark is archived at /models/results-2026-08-19.