Autonomous task horizons are lengthening
METR proposes measuring agents by the duration of human work they can complete at a stated success rate. Its longitudinal results show rapid growth in the length of software tasks frontier systems can finish under test conditions.
FUURAA original conceptual visualWhat the evidence indicates
A concise reading of the source
METR proposes measuring agents by the duration of human work they can complete at a stated success rate. Its longitudinal results show rapid growth in the length of software tasks frontier systems can finish under test conditions.
FUURAA interpretation
Why this could matter
Task duration offers a more operational signal than isolated question answering, but it remains dependent on benchmark design and success thresholds.
How to read this signal
Documented development
The underlying event, report or finding has been published. Its future consequences may still be uncertain.



