Pass IndexThe State of AISign in

Red teaming

Deliberately attacking your own system to find what makes it misbehave before somebody else does.

how it works · the vocabulary

adversarial testing

For models this means hunting for prompts that produce harmful output; for agents it means hunting for inputs that make them take actions they should not. The second is now the larger risk and gets the smaller share of the attention.

In the catalogue

Nearby

AgentAgent memoryAlignmentAutoregressiveBM25ChunkingCold startComputer use