Pass IndexThe State of AISign in

Closed evaluation models, tools and agents

8 in the catalogue today, and the list grows as the market does. Every one with what it costs, who sells it and where it stands.

Tools for judging what a model produced: scoring runs, tracing a chain of calls, catching a regression before a user does. Priced per trace or per evaluation. Most use a model as the judge, which is worth knowing, because it means your evaluation has a bill and an opinion of its own.

Models whose weights are not published and which are bought through somebody's API. You get the maker's infrastructure, their scale and their uptime, and no way to run the thing yourself or to keep it if it is withdrawn. Most of the strongest models are here, so this is less a choice than a fact about the market; the choice is which seller you buy the same model from, and the catalogue holds their prices side by side.

Wider

Closed models, tools and agentsEvaluation models, tools and agents