Red teaming
Deliberately attacking your own system to find what makes it misbehave before somebody else does.
how it works · the vocabulary
adversarial testing
For models this means hunting for prompts that produce harmful output; for agents it means hunting for inputs that make them take actions they should not. The second is now the larger risk and gets the smaller share of the attention.