To the line
Capabilities on the main line · POST-TRAINING

Following human intent

In one viewIn 2022, models became much better at understanding requests and responding through dialogue. They were tuned on instruction examples and human evaluations.

ChatGPT turned a language model from a research tool into a mainstream interface. Confident errors and the tendency to agree with users remained major limitations.

StatusPASSED
TypeCapabilities on the main line
Marker2022
Events in dossier7
Development chronology

Researched

2022-03

InstructGPT

About this eventInstruction tuning — the model starts following orders.

InstructGPT combined demonstrations of instructions with human preference feedback and showed that a smaller aligned model could be more useful than a larger base model. It shifted model development beyond pretraining toward post-training for response format, refusals, style and user intent.

Source: arXiv
2022-11

ChatGPT

About this eventA dialogue interface made instruction following a mainstream way to use a language model.

Source: OpenAI
2022-12

Constitutional AI

About this eventTraining against a written set of principles instead of hand-labelling every answer.

Source: arXiv
2023-03

GPT-4

About this eventHuman-level results on professional exams move AI out of the toy category.

Source: OpenAI
2024-03

Claude 3

About this eventAnthropic's family established sustained frontier competition across reasoning, coding and image analysis.

Source: Anthropic

In progress

сейчас

Sycophancy as a defect

About this eventA model trained to please tends to agree. Teaching it to push back is an open alignment problem.

Planned

впереди

Learning from outcomes

About this eventReplacing "did the answer feel good" with "did the result actually work".

Sources and research

Primary material behind this dossier: papers, lab publications and official reports.

Directions
Capabilities on the main line