Back to section
Scientific discovery · Feb 2026

OpenAI internal model

First Proof research problems used as an eval

During training, an internal model was tested on ten new research-level problems. After expert feedback, OpenAI judged at least five proof attempts likely correct.

Important, but not a clean autonomous eval: people selected attempts and requested clarifications. One attempt initially judged promising was later withdrawn as incorrect.

Sources
OpenAI · First Proof submissionsOpen primary source