InstructGPT
Instruction tuning — the model starts following orders.
Models were taught to answer the way a person finds useful: plain words, on point, following the request. People compared candidate answers, and the model was tuned on those comparisons.
Training on human preference turned a language model into an interlocutor. Natural language became the universal interface to computing.
Instruction tuning — the model starts following orders.
The fastest technology adoption in history.
Training against a written set of principles instead of hand-labelling every answer.
Human-level results on professional exams move AI out of the toy category.
A model trained to please tends to agree. Teaching it to push back is an open alignment problem.
Replacing "did the answer feel good" with "did the result actually work".