Evaluation models, tools and agents
44 in the catalogue today. Every one with what it costs, who sells it and where it stands.
Tools for judging what a model produced: scoring runs, tracing a chain of calls, catching a regression before a user does. Priced per trace or per evaluation. Most use a model as the judge, which is worth knowing, because it means your evaluation has a bill and an opinion of its own.
- AI review runLingo.dev$0.01per call1 selling
- Aion-RP 1.0 (8B)AionLabs$0.8→$1.6per Mtok in / out2 selling
- AlphaEvolveGoogle
- DeepSeek R1 0528 Qwen3 8BDeepSeek196th of 290· 3 boards$0.06→$0.09per Mtok in / out1 selling
- DeepSeek V3 (Turbo)DeepSeek$2→$8per Mtok in / out3 selling
- DeepSeek-Coder-V2-InstructDeepSeek1st of 170$1.2→$1.2per Mtok in / out1 selling
- DeepSeek-Coder-V2-Lite-InstructDeepSeek10th of 170$0.5→$0.5per Mtok in / out1 selling
- DeepSeek-R1-Distill-Llama-8BDeepSeek$0.2→$0.2per Mtok in / out3 selling
- Eval explanationsPatronus AI$0.01per call1 selling
- GPT OSS Safeguard 120BOpenAI$0.15→$0.6per Mtok in / out3 selling
- Grok 3 MinixAI21st of 108· 17 boards$0.25→$1.27per Mtok in / out3 selling
- Grounding with your dataGoogle
- Isaac 0.2 1BPerceptron$0.15→$1.25per Mtok in / out1 selling
- Isaac 0.2 2B PreviewPerceptron$0.15→$1.25per Mtok in / out1 selling
- LLM evaluationConfident AI$0.05→$0.4per Mtok in / out1 selling
- LMUnitContextual AI$3per Mtok in1 selling
- Large evaluatorPatronus AI$0.02per call1 selling
- Llama Guard 3 11B VisionMeta$0.35→$0.35per Mtok in / out2 selling
- LocoOperator-4BLocoreMind
- Nemotron Nano 9B V2NVIDIA234th of 302$0.04→$0.16per Mtok in / out5 selling
- Nemotron-3-Nano-Omni-30B-TEENVIDIA$0.025→$0.098per Mtok in / out1 selling
- Nemotron-Content-Safety-3.5NVIDIA$0.2→$0.2per Mtok in / out1 selling
- Olmo 3 32B ThinkAllen Institute for AI (Ai2)
- PaddleOCR-VL-0.9Bpaddlepaddle$0.14→$0.8per Mtok in / out1 selling
- Phi-4-reasoningMicrosoft$0.12→$0.5per Mtok in / out1 selling
- Qwable-v1lordx64
- Qwen 3.5 PlusAlibaba76th of 290· 16 boards$0.4→$2.5per Mtok in / out3 selling
- Qwen2.5-Omni-3BAlibaba
- Qwen3 30B A3B Thinking 2507Alibaba132nd of 290· 12 boards$0.2→$2.4per Mtok in / out4 selling
- Qwythos-9B-Claude-Mythos-5-1Mempero-ai
- Reka Flash ResearchReka$0.025–$0.06per call1 selling
- ScoresBraintrust$0.0025per call1 selling
- Small evaluatorPatronus AI$0.01per call1 selling
- Sonar Deep ResearchPerplexity$2→$8per Mtok in / out2 selling
- clip-vit-large-patch14-336OpenAI$0.0005per second1 selling
- ctxl-rerank-v2-instruct-multilingualContextual AI$0.05per Mtok in1 selling
- deepsearch-v1Jina AI$0.05per Mtok in1 selling
- deepseek-chatDeepSeek22nd of 72· 7 boards$0.26→$1.03per Mtok in / out3 selling
- fable-tracesAliesTaha
- llama-nemotron-rerank-vl-1b-v2NVIDIA$0.01per Mtok in1 selling
- nemotron-lightning-3.5-30b-a3b$0.05→$0.2per Mtok in / out, fireworks lane1 selling
- phi-4-reasoning-plusMicrosoft$0.07→$0.35per Mtok in / out2 selling
- relace-compactRelace$0.2→$0.2per Mtok in / out1 selling
- traceLangSmith$0.0005–$0.005per call1 selling