Alignment
The work of making a model's behaviour match what its makers intend.
how it works · the vocabulary
safety trainingconstitutional AI
In practice it covers refusing harmful requests, telling the truth about uncertainty, and following instructions rather than reinterpreting them. It is a training outcome, not a filter bolted on afterwards, which is why models differ in character as much as in capability.