Open-weight evaluation models
11 in the catalogue today. Every one with what it costs, who sells it and where it stands.
Tools for judging what a model produced: scoring runs, tracing a chain of calls, catching a regression before a user does. Priced per trace or per evaluation. Most use a model as the judge, which is worth knowing, because it means your evaluation has a bill and an opinion of its own.
Weights published under a licence that lets you run them where you like and sell what you build — Apache 2.0, MIT and their kin impose little beyond keeping the notice. This is the list to start from if you need the model on your own hardware, in your own region, or simply want no third party between you and it. The trade is that hosting is now your problem, and the prices shown beside each one are what somebody else charges to do it for you.
- Aion-RP 1.0 (8B)AionLabs$0.8→$1.6per Mtok in / out2 selling
- DeepSeek R1 0528 Qwen3 8BDeepSeek196th of 290· 3 boards$0.06→$0.09per Mtok in / out1 selling
- DeepSeek-R1-Distill-Llama-8BDeepSeek$0.2→$0.2per Mtok in / out3 selling
- GPT OSS Safeguard 120BOpenAI$0.15→$0.6per Mtok in / out3 selling
- LocoOperator-4BLocoreMind
- Olmo 3 32B ThinkAllen Institute for AI (Ai2)
- PaddleOCR-VL-0.9Bpaddlepaddle$0.14→$0.8per Mtok in / out1 selling
- Qwen3 30B A3B Thinking 2507Alibaba132nd of 290· 12 boards$0.2→$2.4per Mtok in / out4 selling
- Qwythos-9B-Claude-Mythos-5-1Mempero-ai
- fable-tracesAliesTaha
- phi-4-reasoning-plusMicrosoft$0.07→$0.35per Mtok in / out2 selling