Reasoning as a separate mode
About this evento1 showed that extra inference-time compute can materially improve hard problem solving.
This begins the shift from assistant to agent. Today these systems still need human oversight: they accumulate errors, lose the goal on long tasks and learn little from their own experience.
About this evento1 showed that extra inference-time compute can materially improve hard problem solving.
About this eventAnthropic opened a standard for connecting models to data and work tools.
About this eventThe Responses API and Agents SDK combined reasoning, tools, orchestration and observability.
About this eventMETR proposed measuring agents by the human task duration they complete at a given reliability.
About this eventGoogle opened a protocol for agents from different vendors to exchange tasks and results.
About this eventAgentKit turned agent applications into an engineering layer with versions, evaluations and traces.
About this eventAnthropic released a model for long-running coding and knowledge-work tasks designed around hours or days of agentic work rather than a single answer.
Fable 5 matters less as a new peak on short tests than as a change in the unit of work: it was designed for sustained processes involving tools, files and intermediate verification. Anthropic temporarily restricted the model after launch and later restored global access, underlining that longer autonomy requires dependable safeguards as well as capability.
Source 1: Anthropic · запускSource 2: Anthropic · карточка моделиSource 3: Anthropic · повторный запускAbout this eventAnthropic updated its mainstream model for coding, agent workflows and everyday knowledge work.
Sonnet 5 shows the other side of agentic progress: long-running workflows must not only be possible on a flagship but fast and affordable enough for routine use. The event therefore belongs on the autonomy line even though the model remains general-purpose rather than a standalone agent.
Source 1: Anthropic · анонсSource 2: Anthropic · документация моделейAbout this eventChatGPT Work carries out longer tasks across files, apps and finished deliverables.
About this eventThe longer the action chain, the greater the chance of accumulating an error or misunderstanding the goal.
About this eventAn agent needs to retain experience across tasks and improve without full model retraining.
About this eventA system completes a long task independently while its plan, actions and outcome remain auditable.
Primary material behind this dossier: papers, lab publications and official reports.