pass@k
The chance that at least one of k attempts is correct.
how it is measured · the vocabulary
k samples
pass@1 is what a user experiences. A large gap between pass@1 and pass@10 means the model can solve the problem but not reliably, which is exactly the case where letting an agent retry is worth the money.