By mid-August, OpenAI counted 3.1 agent-workdays for each human workday in its research organization. This measures runtime, not an equivalent amount of useful research. People still choose priorities and decide which results to pursue.

Pachocki describes a shrinking view into the models’ work: reasoning mixes with tool use and conversations, models can manipulate that reasoning, and more capability works without a written trace. He calls for shared safety requirements and voluntary slowdowns. That is his assessment, not a published count of undetected incidents.