Pass IndexThe State of AISign in

DPO

A cheaper way to train on preferences that skips the reward model and optimises the model directly against pairs of better and worse answers.

how it works · the vocabulary

direct preference optimisationORPOKTO

Most open-weight models released since 2024 are aligned this way or with one of its descendants, because it is simpler and needs far less compute than RLHF.

Nearby

AgentAgent memoryAlignmentAutoregressiveBM25ChunkingCold startComputer use