Pass IndexThe State of AISign in

LLM as judge

Using a strong model to grade another model's answers against a rubric.

how it is measured · the vocabulary

model gradingrubric

Cheap, fast, and biased in known ways: it favours longer answers, its own family's style, and whichever answer it sees first. Usable when the biases are controlled for, misleading when they are not.

Nearby

ARC-AGIBenchmarkBenchmark contaminationEloEval setGPQAHumanEvalLMArena