Pass IndexThe State of AISign in

pass@k

The chance that at least one of k attempts is correct.

how it is measured · the vocabulary

k samples

pass@1 is what a user experiences. A large gap between pass@1 and pass@10 means the model can solve the problem but not reliably, which is exactly the case where letting an agent retry is worth the money.

Nearby

ARC-AGIBenchmarkBenchmark contaminationEloEval setGPQAHumanEvalLLM as judge