Conditionally licensed evaluation models
7 in the catalogue today, and the list grows as the market does. Every one with what it costs, who sells it and where it stands.
Tools for judging what a model produced: scoring runs, tracing a chain of calls, catching a regression before a user does. Priced per trace or per evaluation. Most use a model as the judge, which is worth knowing, because it means your evaluation has a bill and an opinion of its own.
Weights you can download, under a licence that asks for something in return: an acceptable-use policy, a naming requirement, a revenue ceiling above which you must ask, a restriction on training other models. The Llama and Gemma families are the familiar cases. Read the actual licence before you build on one of these — the condition is usually easy to meet and occasionally fatal, and which it is depends on your product rather than on the model.
- DeepSeek-Coder-V2-InstructDeepSeek1st of 170$1.2→$1.2per Mtok in / out1 selling
- DeepSeek-Coder-V2-Lite-InstructDeepSeek10th of 170$0.5→$0.5per Mtok in / out1 selling
- Llama Guard 3 11B VisionMeta$0.35→$0.35per Mtok in / out2 selling
- Nemotron Nano 9B V2NVIDIA234th of 302$0.04→$0.16per Mtok in / out5 selling
- Qwable-v1lordx64
- Qwen2.5-Omni-3BAlibaba
- llama-nemotron-rerank-vl-1b-v2NVIDIA$0.01per Mtok in1 selling