OpenAI's First Chip Beats Nvidia's Flagship on Its Own Benchmark

Jalapeño claims up to 1.9x throughput per kilowatt over Nvidia's GB300, Anthropic is pitching IPO investors a $30 trillion market, and a Russian drone's autonomous kill in Ukraine gets a name and a face

Wednesday, August 26, 2026 · 9 min read · Issue 248

OpenAI showed up at Hot Chips on Tuesday with numbers, not slides. Jalapeño, the inference ASIC it co-developed with Broadcom over a nine-month RTL-to-tapeout cycle, delivered 1.5x to 1.9x more throughput per kilowatt and up to 3.6x lower end-to-end latency than Nvidia's GB200 and GB300 rack systems, according to public benchmarks run with SemiAnalysis using its InferenceX suite. The tests covered three open models — GPT-OSS 120B, DeepSeek R1 670B, and Moonshot's 1-trillion-parameter Kimi K2.5 — with a 700W Jalapeño part going up against accelerators rated at 1,200W and 1,400W. SemiAnalysis, which ran the tests alongside OpenAI engineers in the company's own lab, called it a chip "beating every Nvidia, AMD, and Google chip we have been able to test."

The caveats matter as much as the headline. Jalapeño wasn't tested against Vera Rubin, the Nvidia platform slated to anchor the first gigawatt of systems OpenAI agreed to deploy in the second half of 2026. It doesn't train models, where Nvidia remains unchallenged. And the comparison ran Jalapeño against GB300 configurations using single-token prediction, even though Nvidia deployments commonly run multi-token prediction in production — a setting where OpenAI's own appendix shows the efficiency lead shrinking to roughly 1.5x. Still, the memory math is real: each Jalapeño package pairs its compute die with six HBM4 stacks totaling 216 GiB at 15.4 TB/s, which works out to roughly 50% more memory per watt of rated power than the GB300's 288GB of HBM3E. That's a direct claim on a supply chain where Samsung, SK hynix, and Micron have sold out HBM capacity through 2027, and where OpenAI's 10-gigawatt chip agreement with Broadcom would make it a serious new claimant against Nvidia's existing multi-year allocation deals.

The timing is not incidental. OpenAI published Jalapeño's results one week after Nvidia agreed to backstop up to $105 billion in financing for OpenAI's data centers — a company simultaneously depending on Nvidia's balance sheet and publishing benchmarks designed to erode Nvidia's hardware moat. Neither fact cancels the other out. OpenAI needs Nvidia's capital and supply today; it's building toward not needing Nvidia's silicon tomorrow. That's the same logic driving Broadcom's custom-ASIC business across Google, Meta, and now OpenAI — inference workloads are commodity enough, and expensive enough at scale, that every hyperscaler with the balance sheet to do so is building an exit ramp off Nvidia's margins, whether or not they exit on cost or on Jalapeño-shaped efficiency numbers.


Briefs

Anthropic is telling IPO investors it sees over $30 trillion in potential market, the largest total-addressable-market claim in IPO history, according to the Wall Street Journal — topping SpaceX's $28.5 trillion pitch from earlier this year. Bankers have floated a raise above $100 billion at a valuation near $2 trillion, and Anthropic could file its prospectus before the end of August. Both numbers would be IPO records; both also require accepting that the pitch is effectively pricing the automation of a meaningful share of white-collar labor, a bet that carries the weight of the number rather than justifying it. Today's Anthropic revenue is nowhere near $30 trillion — the claim is entirely about addressable market, not current sales. (Yahoo Finance)

A New York Times investigation identified the first documented case of civilian deaths from a Russian drone using fully autonomous targeting. A Molniya drone carrying an Nvidia Jetson Orin module crashed at a Zaporizhzhia gas station last month after selecting its final target without a human pilot, killing 19-year-old accounting student Tetiana Bubynets and two men aged 41 and 48. Human operators launched the drone toward the station, but software trained to recognize objects like propane tanks chose the exact aim point onboard; the aircraft failed to clear an apartment building and detonated near people sheltering below rather than hitting its intended target. Nvidia confirmed the recovered modules were genuine Jetson Orin units — consumer hardware starting at $249 that isn't covered by the export controls governing data-center accelerators — and said it doesn't sell to Russia and can't track resales. Ukrainian investigators have now linked the same chip to four separate Russian weapon families since a teardown first found it in the V2U loitering munition in June 2025. (Tom's Hardware)

Meta hired Luke Metz, a senior OpenAI researcher who worked on model behavior and post-training, for its superintelligence lab, Axios reported, the latest in a run of talent moving from OpenAI and Google DeepMind toward Meta's restructured AI division under Alexandr Wang. The hire follows a pattern rather than breaking one: frontier labs are increasingly competing on compensation and research-freedom terms for a small pool of senior researchers, and Meta's checkbook has been the most aggressive of the three. (Axios)

Inherent, a London AI lab founded by Google DeepMind alumni, says its research agent Faraday outperformed Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5 at independently replicating published scientific findings — while running on a 27-billion-parameter Qwen 3.6 model, a fraction of the size of the frontier systems it beat. Cofounder Edward Hughes said the point wasn't beating bigger models but how Inherent got there: reinforcement learning aimed at "research taste" rather than raw accuracy, on the theory that instinct for which experiments are worth running generalizes better than memorized methodology. The 12-person startup, which raised a $50 million seed in May, plans to grow to 20-25 people by year-end. (TechCrunch)

Zhipu says its new coding model GLM-5.3 developed unexpectedly strong offensive cybersecurity skills as a side effect of post-training scale, not design intent. The model scored 84.5% on the CyberGym vulnerability-discovery benchmark, edging out Anthropic's Mythos 5 (83.8%) and OpenAI's GPT-5.6 Sol (83.6%), though it trails both badly on ExploitBench, the deeper exploitation-chain test, at 54.4% versus 78% and 76.5%. Zhipu said GLM-5.3 has already found 2,436 vulnerabilities across 269 real-world codebases in partnership with Chinese security teams, including 107 rated critical. The company's own framing is blunt: teach a model to be a brilliant software engineer, and the same reasoning that fixes bugs finds exploits — and once weights ship openly, any safety guardrails can be stripped without consequence. (CSO Online)


Dispatch: China / East Asia

Alibaba shares fell as much as 10% in Hong Kong on Monday after the company priced an HK$80 billion ($10.2 billion) placement of new shares to non-U.S. investors, with proceeds earmarked entirely for AI infrastructure. The raise lands days after Alibaba reported a 75% year-over-year profit drop for the June quarter, with capital expenditure up 75% to 67.7 billion yuan — the market's read is straightforward: Alibaba is diluting existing shareholders to fund a spending pace its current profit can't support, and investors are pricing that trade-off in real time rather than waiting for the AI payoff to materialize. Separately, DeepSeek released an experimental multimodal variant of its flagship V4 Flash model that can parse images and screenshots, positioning it to approach — not match — the visual-reasoning performance of Anthropic's Claude Opus 4.8, per Bloomberg. Both moves read the same way: Chinese labs and platforms are scaling capital and capability in the same week Alibaba's own numbers show how expensive that scaling actually is. (CNBC; Bloomberg)

Dispatch: India

Voice-AI startup Ringg, which processes 20 million call attempts a month for customers including Cred, Flipkart, Practo, and Policybazaar, raised an additional $10 million from Peak XV Partners as an extension of its Series A, bringing the round to $15.5 million total. Ringg started as a text-to-speech company called DesiVocal before training its own models proved too expensive; it pivoted up the stack to build voice agents instead, and is now moving from high-volume, low-complexity use cases like outbound collection calls — "always going to be a price game," per cofounder Siddharth Tripathi — toward stickier workflows like clinic appointment booking (1,200 locations for Practo) and fintech KYC checks. The bigger story is the crowd Ringg is competing inside: model makers (Deepgram, ElevenLabs, Cartesia, India's Sarvam and Smallest.ai), orchestration players (Bolna, Blue Machines), and sector specialists (Gnani, Arrowhead) are all stacking on top of each other in the same voice-AI layer, betting that owning the customer relationship and the outcome — not the underlying model — is where the durable margin sits. India's 76% consumer preference for phone-based business contact, per Truecaller, is the demand-side bet underneath all of it. (TechCrunch)

Dispatch: Europe

The EU's AI Act transparency rules requiring mandatory labeling of AI-generated audio, video, and images that could pass as authentic became enforceable on August 2, with a four-month grace period for systems already on the market and fines up to €15 million or 3% of global turnover for noncompliance. More than 180 organizations, including Google and Meta, have signed the Commission's voluntary code of practice; Google says its SynthID watermarking has already been applied to more than 100 billion images and 60,000 years of audio. Green MEP Sergey Lagodinsky, who helped negotiate the Act, dismissed industry complaints about compliance burden directly: "I have a deep respect for industry, but in all too many cases, they complain, and then after things are implemented, it doesn't look that burdensome to them." Whether Europe's framework actually holds capability-labeling to account depends on enforcement over these four months, not the rule's existence. (The Guardian)


The View

Three stories this issue are variations on the same tension: capability is scaling faster than the infrastructure meant to constrain, finance, or account for it. OpenAI's Jalapeño benchmarks are a bid to unwind Nvidia's hardware margin even while OpenAI depends on Nvidia's $105 billion financing backstop signed a week earlier — dependency and disruption running on parallel tracks, not in sequence. Anthropic's $30 trillion IPO pitch prices a future where AI automates a meaningful share of white-collar labor, a claim that requires investors to buy the destination without much evidence of the route; SpaceX made a similar bet three months ago and the market absorbed it, which is precisely why Anthropic's bankers think this number will fly too. And the Jetson Orin story is the sharpest version of the same pattern: consumer-grade edge hardware, sold for $249 with no export-control coverage, ended up guiding a fully autonomous strike that killed three civilians — a gap between what the chip was built for and what it got used for that no compliance regime currently closes. None of these are new dynamics. What's changed is the dollar and human cost of letting them run unexamined has gotten large enough that "we'll figure out governance later" is no longer a free option for anyone involved.

The Miss

Most coverage of the Jalapeño benchmarks led with the "beats Nvidia" framing and buried the fact that OpenAI's own appendix shows the lead shrinking to roughly 1.5x once you account for multi-token prediction — the setting Nvidia deployments actually use in production, not the single-token comparison OpenAI chose for its headline numbers. That's not a minor footnote; it's the difference between "OpenAI built a chip that beats Nvidia's flagship" and "OpenAI built a chip that's meaningfully more efficient in a narrower band of workloads than its default marketing framing suggests." SemiAnalysis's own writeup included the caveat. Most of the outlets repeating the 1.9x number didn't.


Pull Quotes

"Our Jetson Orin modules are consumer-grade products sold to students, developers, and startups for a wide range of beneficial applications. They are not available in Russia and are not designed for military purposes." — Nvidia spokesperson, to Tom's Hardware

"Machines are making decisions to strike." — Col. Serhiy Minaiev, Zaporizhzhia air defense commander, to the New York Times

"We are reaching a stage where if we teach an AI to be a brilliant software engineer, you're accidentally teaching it how to be a good hacker, too." — Zhipu statement on GLM-5.3's cybersecurity capability

"I have a deep respect for industry, but in all too many cases, they complain, and then after things are implemented, it doesn't look that burdensome to them." — Sergey Lagodinsky, Green MEP


  • OpenAI's 700W Jalapeño ASIC outpaces 1,400W Nvidia flagship GPU — Tom's Hardware
  • $30 Trillion Dream: Can Anthropic Sell the Biggest IPO Ever? — Yahoo Finance
  • Nvidia Jetson Orin-guided Russian AI drone killed three civilians in Ukraine — Tom's Hardware
  • Meta hires OpenAI's Luke Metz for superintelligence lab — Axios
  • Inherent's Faraday outperforms Anthropic and OpenAI at replicating research — TechCrunch
  • Zhipu says new coding AI developed advanced cyber skills faster than expected — CSO Online
  • Alibaba plunges after announcing $10.2 billion share placement to fund AI push — CNBC
  • DeepSeek unveils test model to rival Anthropic's Opus 4.8 — Bloomberg
  • India's Ringg gets backing from Peak XV as it pushes voice AI past the phone call — TechCrunch
  • AI labels to be compulsory on authentic-looking content under EU rules — The Guardian

Out

That's the briefing. Issue 249 tomorrow.