Pass IndexThe State of AISign in

HumanEval

164 small Python functions the model must write from a docstring, graded by running the tests.

how it is measured · the vocabulary

pass@1

Historically important, now saturated and too easy to be informative. Quoted scores above 90 tell you almost nothing about whether a model can work in a real codebase.

Nearby

ARC-AGIBenchmarkBenchmark contaminationEloEval setGPQALLM as judgeLMArena