OpenAI’s research agents show how frontier labs could compound their advantage
OpenAI’s internal agent report shows how early model access and large inference budgets could compound research advantages, with costs and limits.

OpenAI says its coding agents now log 3.1 workdays of runtime for every human workday across its research operation, and that it has reached its September goal of building an automated research intern. Its September 6 internal research report, covered in The Rundown’s September 8 newsletter, offers a look at the resources behind that push.
The figures point to a potential competitive advantage for frontier labs. Early access to powerful models gives them time to develop research workflows, while large inference budgets let them put those workflows into practice at scale. OpenAI’s report also shows why the size of that advantage remains difficult to measure.
What OpenAI’s agents are doing
By mid-August, median daily researcher inference consumption exceeded $600 at API prices, with the 90th percentile above $7,000, according to OpenAI. Those figures price the workload at published API rates. They do not disclose what OpenAI actually pays internally to run it.
The 3.1 agent-workdays figure similarly needs context. OpenAI converts aggregate runtime into eight-hour days. That measures how long agents run, rather than how much useful research they produce. Its researcher population also includes people in research support roles.
OpenAI describes its achieved intern milestone as handling well-defined tasks under human direction, including assignments lasting several days. A fully automated researcher remains a target for March 2028. The current report gives a more bounded picture of automation, with researchers assigning work and stepping in when needed.
Troubleshooting and monitoring research runs have grown as agent activities. OpenAI also reports better results on longer, more complex tasks, but more than half of successful tasks lasting four to eight hours required human intervention. Its success charts exclude uncertain outcomes, which limits how broadly readers can interpret the results.
August experiments per active experimenter reached their highest level since tracking began in January 2025. Compute availability also increased, making it difficult to isolate how much of that growth came from agents. These are OpenAI’s observations about its own operation, without a controlled comparison against competitors.
Why it matters
The competitive implication is substantial. Labs that develop frontier models can put them to work internally before outside researchers gain access, then spend heavily enough on inference to make those capabilities part of everyday research. If that work helps improve subsequent models and research methods, the advantage could compound. OpenAI’s report makes that possibility more concrete, while leaving its scale unmeasured.
Astra’s release puts the timing in perspective. OpenAI announced GPT-6 Astra on September 3, beginning with selected organizations and planning broader paid-plan and API availability over subsequent days. By September 8, its rollout had begun. The announcement alone does not establish how far that rollout had progressed.
Public access may spread the capability without immediately erasing experience accumulated inside a lab. Researchers who have already learned how to assign tasks, preserve useful context, and manage interventions may have a practical head start. OpenAI describes experimental Codex support for keeping notes across context windows and searching earlier context. Those features offer a concrete way to make longer assignments more workable, although the report does not isolate Astra’s contribution to overall research gains.
For competing labs and smaller research organizations, the challenge therefore extends beyond obtaining a model. They would need to develop effective workflows, fund repeated runs, and provide enough human oversight to turn agent activity into useful experiments. OpenAI’s spending figures illustrate the intensity of its deployment. They establish neither the necessary budget for another organization nor the return on each inference dollar.
Secure operation adds another burden. In an August 18 disclosure, OpenAI reported a two-week pause in reinforcement learning training on its latest models intended for deployment while it strengthened safeguards. It estimated monitoring overhead at roughly 20% of the inference compute being monitored, with substantial variation by workload. That supports the writer’s resource argument by showing that operating powerful agents also demands monitoring, infrastructure, and engineering investment.
Greater capability can bring interruptions, too. OpenAI’s Astra announcement says production checks can slow, pause, or stop legitimate work. The competitive payoff depends on how reliably labs convert all that activity into useful research. Privileged access and deep budgets could reinforce one another, but the report provides no measured lead over rivals or quantified rate of compounding.
Sources & further reading
This story builds on reporting from The Rundown newsletter on September 8, 2026.