An agent feeds the output of one step into the next, so small mistakes compound along a long trajectory. Reliability therefore needs more than a stronger model: explicit state checks, action limits, audit logs and a safe way to stop the process.
Source: METR · длинные задачиSilent failures
About this eventThe core problem of long tasks: failing without signalling failure — confidently returning a wrong result.