The paper showed that intermediate textual steps can materially improve multi-step problem solving in sufficiently large models. This is not proof of genuine internal reasoning, but it supplied a practical signal: extra computation at answer time can become a separate axis of capability improvement.
Source: arXivChain of thought
About this eventAsking the model to think step by step raised accuracy — the first hint that thinking at answer time pays.