RLHF
Training a model on human preferences: people rank two answers, a reward model learns the pattern, and the model is tuned to score well against it.
how it works · the vocabulary
reinforcement learning from human feedback
It is what makes a model helpful and polite rather than merely fluent, and it is also where refusal behaviour and house style come from. Expensive, because it needs human judgement at scale.