Closed evaluation models, tools and agents
8 in the catalogue today, and the list grows as the market does. Every one with what it costs, who sells it and where it stands.
Tools for judging what a model produced: scoring runs, tracing a chain of calls, catching a regression before a user does. Priced per trace or per evaluation. Most use a model as the judge, which is worth knowing, because it means your evaluation has a bill and an opinion of its own.
Models whose weights are not published and which are bought through somebody's API. You get the maker's infrastructure, their scale and their uptime, and no way to run the thing yourself or to keep it if it is withdrawn. Most of the strongest models are here, so this is less a choice than a fact about the market; the choice is which seller you buy the same model from, and the catalogue holds their prices side by side.
- AlphaEvolveGoogle
- Grok 3 MinixAI21st of 108· 17 boards$0.25→$1.27per Mtok in / out3 selling
- Grounding with your dataGoogle
- Qwen 3.5 PlusAlibaba76th of 290· 16 boards$0.4→$2.5per Mtok in / out3 selling
- Reka Flash ResearchReka$0.025–$0.06per call1 selling
- Sonar Deep ResearchPerplexity$2→$8per Mtok in / out2 selling
- clip-vit-large-patch14-336OpenAI$0.0005per second1 selling
- nemotron-lightning-3.5-30b-a3b$0.05→$0.2per Mtok in / out, fireworks lane1 selling