DPO
A cheaper way to train on preferences that skips the reward model and optimises the model directly against pairs of better and worse answers.
how it works · the vocabulary
direct preference optimisationORPOKTO
Most open-weight models released since 2024 are aligned this way or with one of its descendants, because it is simpler and needs far less compute than RLHF.