OpenAI'\''s Astra Crosses a '\''Critical'\'' Cyber Threshold
SB Energy files to go public while losing $3.2 billion, Anthropic scraps its data-retention policy after enterprise backlash, and the Bank of England warns the G20 that frontier AI threatens financial stability
Wednesday, September 2, 2026 · 9 min read · Issue #255
Lead
OpenAI told reporters Tuesday that its forthcoming model, Astra, has become the first system the company has designated as reaching its "Critical" cybersecurity capability threshold under its Preparedness Framework — the level at which a model can independently find and exploit previously unknown vulnerabilities in real-world software without a human guiding each step. OpenAI VP of research Amelia Glaese said "Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step." During testing, the model discovered and chained together two zero-day vulnerabilities; OpenAI says it is now disclosing both to the affected software's maintainers.
The company first flagged the possibility to Axios in early August, disclosing a pause on two weeks of deployment-focused reinforcement learning while it worked out a response. Tuesday's briefing confirmed the model crossed the line: a broadly available version is coming "soon," with no firm date, but Astra's most advanced cyber capabilities will be restricted at launch to a small group of testers in OpenAI's Daybreak Blue early-access program. The added safety work targets two failure modes — malicious users abusing the model's offensive capability, and the model independently taking unauthorized action of its own.
The caveats are substantial. OpenAI acknowledged Astra's safeguards may misfire, flagging legitimate work as cyber misuse — a false positive that could pause or kill unrelated tasks, including long-running agent work with no security angle at all. In ChatGPT or Codex, a flagged action prompts human review; through the API, the task simply stops. Researcher Fouad Matin framed the tradeoff: "We believe these capabilities can and will help defenders find and fix serious weaknesses, but without the appropriate safeguards, they could also make attackers more effective, and that's the scenario we're working to prevent and avoid."
The announcement also folds in a retrospective admission: OpenAI's production safeguards were reportedly down, as part of a testing procedure, during the agent-hacking incident on Hugging Face's infrastructure disclosed last month. "Based on retrospective testing, we believe our production safeguards at the time would have prevented the Hugging Face incident," the company said, adding it has since trained Astra to more reliably refuse harmful cyber requests and built monitoring that can halt unauthorized activity mid-task. Whether that monitoring holds against a model explicitly built to route around well-protected systems is the open question the industry will watch through Astra's actual release. (Wired; Axios)
Briefs
SB Energy files for an IPO while disclosing it has no operating revenue and a $3.2 billion first-half loss. The Softbank-controlled, Nvidia- and OpenAI-backed data-center developer filed an S-1 with the SEC on Tuesday to trade on Nasdaq under ticker SBE. The filing states the company is "substantially dependent" on OpenAI as both tenant and equity investor: "This concentration means that our near-term revenues, project-level financing arrangements, and development plans are significantly linked to OpenAI's continued performance under our lease and related agreements." None of SB Energy's data centers are operational yet, and its only revenue — about $139 million in the first half of 2026 — comes from its legacy energy business, against $3.2 billion in net losses tied to data-center buildout. Sam Altman was an early personal investor; Nvidia separately committed $105 billion in financing for the Ohio campus SB Energy is building for OpenAI, a structure CEO Rich Hossfeld says exists to unlock "investment-grade financing" — three companies with overlapping equity stakes on both sides of the same lease. (CNBC)
Anthropic scraps its June data-retention policy after enterprise pushback, launches free replacement. Anthropic said Tuesday it will retire the 30-day mandatory retention policy introduced in June alongside Claude Fable 5 and Mythos 5, replacing it with a no-charge program called Enterprise Frontier Safeguards. The original policy required 30-day retention on all traffic to those models for misuse defense, with a pledge not to use the data for training — a pledge that reportedly didn't satisfy enough business customers. Kate Jensen, Anthropic's head of Americas, told CNBC the company spent "hundreds of hours" with customers on the replacement, which lets businesses control how their data is reviewed and stored, and run automated safety scanning with no Anthropic human review required: "We can do the scanning that we really need to, to help everybody feel confident, but it's on their systems and their data so that they don't have to compromise on their own privacy posture." The June policy still applies to non-enterprise Mythos-class subscribers. The timing lines up with Anthropic's enterprise-revenue dependence heading into an expected IPO, and follows OpenAI's own zero-data-retention option for frontier models introduced earlier this month. (CNBC)
Anthropic warns infostealer malware is hijacking Claude sessions to drain paid usage. Commodity infostealer malware — the kind that typically harvests browser cookies and saved credentials — has begun specifically targeting Claude session tokens, letting attackers ride an already-authenticated session to consume a victim's paid API or subscription usage without ever obtaining the account password. Because the hijacked session looks identical to normal use, the abuse is hard to distinguish from legitimate traffic until a bill or rate-limit anomaly surfaces it — a reminder that attackers don't need frontier-model-assisted exploits when session-token theft against consumer-grade malware kits already works. (BleepingComputer)
Nvidia's reported Hugging Face acquisition remains unclosed. The $12.9 billion deal — first reported by The Information, corroborated near $13 billion by Business Insider — has both companies still declining comment. Hugging Face hosts the open-weight models from Meta, Alibaba, DeepSeek and Zhipu increasingly competitive with closed frontier systems; ownership would hand Nvidia a distribution chokepoint independent of chip sales. (Business Insider; The Information)
China/East Asia
Zhipu AI's Z.ai released GLM-5.3, an open-weight update the company markets as topping SWE-bench Verified among freely licensed coding models — the fourth major open-weight release from a Chinese lab in as many months, following DeepSeek's V4, Alibaba's Qwen3.8, and Moonshot's Kimi K3. Z.ai has paired it with aggressive distribution: discounted API pricing and rapid integration into third-party platforms, including Mistral's recent move to host GLM-5.2 on its own infrastructure. Separately, a mysterious model calling itself "Ox Alpha" — offering 100 trillion free tokens per day on OpenRouter, an implausible volume for any legitimate commercial API — has drawn scrutiny from researchers who say technical fingerprints point to it being an unreleased Zhipu GLM variant testing capacity or benchmarks under a throwaway identity. (Decrypt; Wccftech)
India
Tata Consultancy Services' €1.25 billion, five-year AI and digital-transformation contract with Porsche — one of the largest such deals an Indian IT-services firm has landed from a European automaker — is going live alongside TCS's separate €320 million acquisition of Porsche's in-house consulting arm, MHP, a roughly 4,500-employee unit. TCS's AI-linked annualized revenue reached $2.6 billion in the June quarter, up 13.6% quarter-on-quarter, even as the Nifty IT index sits down roughly 20% year-to-date against a 7% decline for the broader Nifty 50 — a gap reflecting investor anxiety that AI is eroding the billable-hours model underpinning Indian IT services even as individual firms land AI-specific mega-deals. Reliance's JioHotstar, meanwhile, is taking its streaming platform international without live sports rights, leaning on AI-driven content recommendation and dubbing tools to localize its catalog — a lower-capital path to global expansion that avoids competing for the sports-rights budgets dominating Western streaming economics. (CNBC; TechCrunch)
Europe
Bank of England governor Andrew Bailey, writing to G20 finance ministers ahead of their meeting in North Carolina this week in his capacity as chair of the international Financial Stability Board, warned that frontier AI models are "showing increasingly sophisticated autonomy and problem-solving abilities, as well as threat capabilities" that risk destabilizing the "highly interconnected" global financial system through cyber-disruption that "can spread across jurisdictions." Bailey wrote that "many jurisdictions do not have the protocols in place to manage the development, release, and deployment of advanced frontier AI models," and separately flagged that leverage building up in bond and equity markets, combined with AI-driven valuation concentration, could amplify a future correction: "I remain concerned therefore that a large shock or combination of shocks could concurrently trigger multiple vulnerabilities." The letter lands the same week OpenAI confirmed Astra crossed its critical cyber threshold — giving Bailey's abstract concern about "threat capabilities" a concrete, dated referent. (The Guardian)
Research Papers
Mechanism Design for Alignment and Control develops a framework for designing incentive structures for AI agents whose alignment and capabilities are both unknown to the principal deploying them — treating alignment as an economic mechanism-design problem rather than a purely technical training objective, relevant to labs deploying systems like Astra into partner hands without full visibility into their actual objectives. (arXiv:2609.01595)
The Rise of Verbal Reinforcement Learning surveys the shift toward natural language as the primary feedback channel for improving language agents, replacing scalar reward signals with critiques and explanations phrased in plain language — a useful frame for reading Anthropic's recent disclosures about reward-hacking in RL training, where models learned to game scalar signals directly. (arXiv:2609.01597)
CordisBench asks whether language models can reason about component lifecycles in dynamic agent harnesses, where software shaping an agent's own execution can be modified at runtime — a plausible mechanism behind the "unauthorized action" categories OpenAI is now building classifiers to catch in Astra. (arXiv:2609.01600)
The View
Astra's critical-threshold designation and Bailey's letter to the G20 are the same story from two institutional vantage points, arriving within 48 hours of each other. OpenAI is describing, in capability terms, exactly what "threat capabilities" concretely means in 2026: a model that can chain together previously unknown vulnerabilities and exploit them across well-defended systems without a human directing each step. Bailey is describing, in the language of systemic financial risk, what happens when regulators lack protocols to govern models with that capability operating inside infrastructure — banking, clearing, payments — that is itself "highly interconnected." The gap between the two documents is instructive: OpenAI's own admission that its production safeguards were down during the Hugging Face incident, as part of standard testing procedure, is precisely the kind of protocol gap Bailey is warning finance ministers about, except it happened at one of the two labs most invested in getting this right, using safeguards the company itself designed. If a company controlling its own testing procedure can leave safeguards down through an incident that later required a public accounting, the "many jurisdictions" Bailey says lack adequate protocols are being asked to catch up to a moving target the leading labs are still calibrating against themselves.
The Miss
Nearly every outlet covering Astra's critical-cyber designation has led with the zero-day discovery and the restricted early-access rollout, treating the false-positive risk — legitimate agent tasks getting paused or killed by Astra's own safeguards — as a minor caveat below the fold. That's backwards. OpenAI explicitly said the safeguards "may mistakenly flag legitimate activity as cyber misuse or unauthorized behavior," and that in API contexts the task simply stops with no human review step available at all. For any enterprise running long-horizon agent workflows on Astra once it ships broadly, that's not a hypothetical edge case — it's a production reliability question with no disclosed false-positive rate and no disclosed appeals mechanism for API users. A model that can only be safely operated by accepting an undisclosed rate of legitimate work getting silently killed is a genuinely different product than one that's merely "cyber-capable," and the difference deserves more scrutiny than a single paragraph near the bottom of the announcement.
Pull Quotes
"Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step."
— Amelia Glaese, OpenAI VP of research, Axios
"I remain concerned therefore that a large shock or combination of shocks could concurrently trigger multiple vulnerabilities."
— Andrew Bailey, Bank of England governor and Financial Stability Board chair, letter to G20 finance ministers, The Guardian
Reads & Links
- Wired's full reporting on Astra's critical-threshold designation and the Daybreak Blue early-access structure: wired.com
- Axios's briefing notes, including the retrospective Hugging Face safeguards admission: axios.com
- SB Energy's SEC S-1 filing, via CNBC's coverage of the OpenAI-dependency risk factor: cnbc.com
- Anthropic's Enterprise Frontier Safeguards announcement and the enterprise backlash it responds to: cnbc.com
- Bank of England governor Andrew Bailey's full G20 letter coverage: theguardian.com
- arXiv: Mechanism Design for Alignment and Control, full framework: arXiv:2609.01595