HumanEval
164 small Python functions the model must write from a docstring, graded by running the tests.
how it is measured · the vocabulary
pass@1
Historically important, now saturated and too easy to be informative. Quoted scores above 90 tell you almost nothing about whether a model can work in a real codebase.