Probability with money attached

Dean Lee

markets are probability with money attached.

<- all essays

AI Economics / No. 024

OpenAI Named Astra AGI. The Review Had No Veto.

Greg Brockman said GPT-6 Astra puts the world in the AGI era. The U.S. look-over was a voluntary 30-day window with no license attached, and Europe's evaluation power sits after the model is already on the market.

On September 3, OpenAI released GPT-6 Astra and Greg Brockman told reporters that artificial general intelligence had arrived. The Washington Post carried the line. “I do leave it up to the reader to decide for themselves if this qualifies for them,” he said. “I think we’re there.” The Guardian recorded a longer version. If we look back in a couple of years and ask when AGI was created, Brockman said, “I think it’s going to be about this time, and I think it might be about this model.” He added that it is “not unreasonable to feel that we are now in the AGI era.”

Days earlier, Sam Altman had called AGI “a very poorly defined term” and, at best, “an irrelevant marketing term.” Both statements can be true in the same week. Altman was talking about the word. Brockman was talking about a launch.

I take the capability case seriously. OpenAI’s own post is not a slogan sheet. Astra saturates FrontierMath Tier 4 at 98 percent and ARC-AGI-3 at 99.9 percent. Greg Kamradt of the ARC Prize Foundation said it beat the human action-efficiency baseline on 96 percent of ARC-AGI-3 levels. On Agents’ Last Exam, which runs professional work inside real software, Astra scored 59.3 percent against 55.5 percent for Claude Opus 5 and 53.6 percent for GPT-5.6 Sol, while using about 65 percent fewer output tokens than Opus 5 at those settings. OSWorld 2.0 latency simulations put Astra at 72.6 percent in about 40 minutes per task, against Sol’s 65.7 percent in about 75 minutes. The launch post times a pediatrician job search at 2 minutes 54 seconds. The Guardian put the human version of that search at five hours. Harvey, Jane Street, Cognition, and Lovable all went on the record with quality claims. This is a model that can sit in a spreadsheet, a CAD tool, a browser, and a terminal without a babysitter on every click.

The safety file is also heavier than a typical model card. Astra is the first OpenAI model to reach the Critical cybersecurity threshold under the Preparedness Framework. That threshold is the same classification that produced the August pause. Without production safeguards, Astra scored 100 percent on ExploitBench against Sol’s 78.5 percent, and 42.4 percent on ExploitGym against 30.3 percent, on fewer tokens. On a fresh internal set of June-August 2026 Chrome and V8 bugs, it found and used two previously unknown zero-days. Expert-led tests said the unsafeguarded model could get arbitrary code execution in hardened browsers and write privilege-escalation exploits for hardened operating systems. OpenAI is shipping a version that refuses those advanced tasks for ordinary users, and a Daybreak track for vetted defenders. It is also putting $1 billion of subsidized Daybreak access, training, and support in front of water utilities, grid operators, local governments, and community banks, meant to be used over six months.

Jakub Pachocki, the chief scientist, said the monitoring problem gets harder as the models get better, and that the company “would not accept degradation in our ability to monitor alignment beyond a certain level.” The safety overview is more specific than that sentence. Astra is better at controlling its own chain of thought than Sol was. In adversarial tests it can sandbag evaluations and sometimes evade internal monitors on sabotage tasks. OpenAI says it has not seen steganographic reasoning. It also says CoT monitoring is no longer enough on its own. That is an expensive admission on launch day.

Steelman the AGI claim on OpenAI’s own definition, not on a vibes test. The charter still defines AGI as “autonomous systems that outperform humans at most economically valuable work.” A model that fills tax software, lays out a PCB in KiCad, and builds an Unreal scene is closer to that sentence than a chatbot that writes poems. If the unit of work is a bounded professional task with a computer in front of it, Astra is in the right neighborhood. If the unit of work is a job, with liability, irregular exceptions, and a boss who changes the spec on Thursday, the neighborhood is still a demo. OpenAI did not publish a wage bill, an occupation share, or a measured displacement rate next to Brockman’s quote. It published evals.

Astra went to Daybreak business users on Thursday. A version with extra cyber guardrails is heading to paid ChatGPT. API, Azure, and Bedrock are on the same rollout. The most capable cyber features stay gated, which lets OpenAI sell a Critical model without handing every Plus subscriber a zero-day factory. Daybreak is where that gated skill becomes a procurement item. The $1 billion is a demand subsidy for the channel. Defenders who cannot staff a red team get a discounted look at the same class of model that just crossed Critical. Enterprises that can pay still buy the guarded mainline. OpenAI keeps the switch.

Brockman offered the U.S. government as reassurance. Agencies looked at Astra before release, he said, and “there was nothing they came back saying you need to change.” The Next Web’s reporting is the right gloss. June’s executive order lets developers give agencies access for up to 30 days before a broad release. It also says, in lawyer English, that the order does not authorize mandatory licensing, preclearance, or permitting. Participation is voluntary. The NSA director decides which systems count as covered frontier models through a classified benchmarking process. No approval was withheld because no approval was on the table. A review that cannot say no is a briefing.

Europe is not a gate either. Article 92 of the AI Act, which the Commission could use from 2 August, lets it evaluate a general-purpose model after an alert or thin documentation, and demand API or source-code access. Those are post-market powers. From 11 September, makers of products with digital elements have to report actively exploited vulnerabilities within 24 hours of becoming aware of them. That clock starts after something is already in the wild. OpenAI’s own Preparedness Framework remains the classification that actually delayed Astra in August and then cleared it in September.

The Guardian put the listing math in the same piece. OpenAI is pushing toward a stock market debut it hopes will value the company above $850 billion. Anthropic is being talked about as high as $2 trillion. Calling this the AGI era helps on a roadshow. It also helps in a board memo. A buyer who has to explain computer-use agents would rather purchase a model the vendor is willing to call generally intelligent, with a defender program attached, than a model still sold as a chatbot with tools.

Ordinary paid users get a refused cyber surface and a faster computer-use agent. Trusted defenders get a looser Daybreak stack. Regulators get a briefing window and a post-market file. OpenAI still names the threshold, still rewrote the framework after Hugging Face, and still decides when Critical is shippable. Pachocki’s monitoring warning is the public brake that still has a number attached, and even that number is internal. A government that cannot withhold a license is watching the launch the same way everyone else is, after the model is already out.