Executive summary
OpenAI’s September 6 report says coding agents now handle research assignments lasting days, under human direction. Alongside it, chief scientist Jakub Pachocki warned that monitoring the models’ written reasoning is becoming less reliable.
3.1×Agent runtime per human workday; not output
>50%Successful 4–8 hour tasks needed human intervention
2028March target for a full AI researcher; not achieved
What happened
By mid-August, OpenAI counted 3.1 agent-workdays for each human workday in its research organization. This measures runtime, not an equivalent amount of useful research. People still choose priorities and decide which results to pursue.
Pachocki describes a shrinking view into the models’ work: reasoning mixes with tool use and conversations, models can manipulate that reasoning, and more capability works without a written trace. He calls for shared safety requirements and voluntary slowdowns. That is his assessment, not a published count of undetected incidents.
The 2050 museum label
The intern was helping build next year's intern.
What this does—and doesn’t—show
These are OpenAI’s internal measurements, not an independent audit. More than half of successful tasks estimated at four to eight human hours needed intervention. The report does not establish runaway self-improvement, and the monitoring warning does not mean all oversight has stopped working.