Pass IndexThe State of AISign in

SWE-bench

Real GitHub issues from real repositories: the model must produce a patch that makes the project's own tests pass.

how it is measured · the vocabulary

SWE-bench Verified

The most credible coding benchmark because it is graded by execution rather than by resemblance. "Verified" is the human-checked subset of 500 tasks that most figures refer to.

In the catalogue

Nearby

ARC-AGIBenchmarkBenchmark contaminationEloEval setGPQAHumanEvalLLM as judge