DeepSeek Starts Building Its Own Inference Chip

China's leading open-weight lab is moving into custom silicon, OpenAI forks Git, and the infrastructure build-out shows early financial strain.

July 12, 2026 | Reading time: 11 minutes | Issue #208

Lead

DeepSeek is designing its own inference chip, three people familiar with the matter told Proactive Investors. The Hangzhou-based company, best known for capable open-weight models that undercut American rivals on price, wants to reduce its dependence on Nvidia and on Huawei's Ascend processors. The chip is aimed at inference — the cheaper, less technically demanding stage where a trained model answers queries — rather than training, where Nvidia's software and interconnect lead is widest and U.S. export controls bite hardest.

The strategy is co-design. DeepSeek builds both the model and the silicon that runs it, letting the two be tuned together in a way that general-purpose hardware cannot match. That matters for a lab already obsessed with serving cost: in May it cut the price of its V4-Pro model by 75%, from $3.30 to under $0.85 per million tokens. A purpose-built inference chip would let it push prices lower still. Around a recent release, DeepSeek said it had used a data format well suited to home-grown chips soon to be released, hinting that this has been in the works for months.

The move also reflects a broader shift in Chinese AI strategy. When DeepSeek optimized its V4 model for Huawei Ascend chips in April, ByteDance, Tencent, and Alibaba all approached Huawei for orders. DeepSeek now appears to be betting it can go further alone, designing silicon tailored to its own models rather than adapting to someone else's. If it succeeds, it would be the most credible vertical-integration play by a Chinese AI lab since the export-control era began. The question is whether China's domestic foundries can produce competitive inference silicon at scale without the advanced lithography needed for training chips.

OpenAI Forks Git on GitHub

OpenAI has created a public fork of the Git source code under its own GitHub organization. The repository, openai/git, was created on July 10 and is a fork of the upstream git/git project. As of July 12 it had 142 stars and no public explanation of its purpose.

The fork is small as signals go, but it is concrete. OpenAI has been building coding agents, voice models, and enterprise deployment tools; maintaining a customized version of Git suggests it may be preparing developer tooling tuned for the scale of agentic coding or internal repository management. It could also be a routine mirror. Either way, the lab is now a maintainer of one of the most important pieces of infrastructure in software, and its changes will be public under the GPL.

Meta Opens Muse Spark 1.1 to Developers

Meta released Muse Spark 1.1 through a public API preview for U.S. developers, along with $20 in free credits per new account. The company claims the new version is a "step-change" from the first Muse Spark model, with better coding, support for end-to-end agentic workflows including multi-agent systems, and native multimodal perception across images, video, and documents.

The launch follows a difficult week for Meta's Muse Image model, which let users generate images using the likenesses of public Instagram accounts before the company suspended the feature under pressure from CAA, SAG-AFTRA, and privacy groups. Spark 1.1 is the more developer-facing half of the same push: Meta needs to show that its Superintelligence Labs can ship models competitive with OpenAI, Google, and Anthropic, and that developers will build on them. The test is whether the API usage converts to production deployments, or merely to benchmark comparisons.

Google Will Label AI-Made Ads

Google is rolling out a global disclosure feature that tells users when an ad was created using AI. The notice will appear in the "My Ad Center" panel across Google Search, YouTube, and Discover. Until now, Google only required such disclosures for election ads; synthetic or altered product images in commercial ads had no consumer-facing label.

The change is a narrow but meaningful transparency move. AI-generated product photography can be misleading when shoppers assume they are looking at a real item. The new label puts the burden on Google to detect AI-made creative, not on advertisers to self-report. Rivals will face pressure to match it, particularly as regulators in the EU and U.K. tighten rules around synthetic media.

ZML Wants to Break Nvidia's Inference Silo

A French startup called ZML has released a free inference server, ZML/LLMD, that it says can run open-source large language models at peak speed across Nvidia, AMD, Google TPU, Apple Metal, and Intel Arc hardware. Endorsed by Turing Award winner Yann LeCun, the company cites a roster of European chip startups — Axelera, Fractile, Kalray, OLIX, Q.ANT, SiPearl, SpiNNcloud, and VSORA — as potential beneficiaries.

The pitch is straightforward: training has drawn most of the attention and investment, but inference is where most of the compute bill lands once models are deployed. Software that makes inference hardware interchangeable lowers lock-in and gives buyers leverage. Europe has several chip challengers but lacks the software ecosystem to make them usable; ZML is trying to supply that layer. The risk is that Nvidia's CUDA moat is deeper than a new inference server can cut in one release.

India Lens

A security researcher found that Coal India's AI-powered mine-monitoring system, Project DigiCoal, exposed a user list complete with plaintext passwords. The "RPI Dashboard," developed by Indian startup DeepSight AI Labs and deployed with Accenture, returned credentials through an unauthenticated API endpoint. Many accounts shared the same weak password.

The incident is a reminder that AI deployments in critical infrastructure often inherit the security posture of the systems they are layered onto. Coal India is one of the world's largest coal producers, and the project involves thousands of CCTV streams with anomaly-detection AI. A dashboard that hands out admin credentials to anyone who asks undermines the entire surveillance and safety stack. The government's AI ambitions in mining, logistics, and public services will face this test repeatedly: model accuracy means little if the surrounding application is trivial to compromise.

Eastern Front

DeepSeek's chip pivot is not the only sign that China is trying to define the next phase of AI on its own terms. In a separate essay published this week, Noema Magazine editor Nathan Gardels argued that China is using open-weight models as a form of global soft power: a closed society is shipping adaptable, cheap models abroad while the United States produces mostly expensive, closed systems. Andrew Ng, who has led AI efforts at both Google and Baidu, is quoted making a similar point — that Chinese open models travel well because they are not priced in dollars and can be tuned locally.

The observation is politically inconvenient for Washington, where lawmakers are already probing U.S. companies' use of Chinese models and the State Department has warned that such models are designed to advance Beijing's narratives. DeepSeek's custom silicon, if it works, would make the economic argument even sharper: a fully Chinese stack optimized for inference could offer lower serving costs than anything dependent on Nvidia's supply chain.

The Map

This week the AI industry looked less like a single race and more like several simultaneous sieges. In the U.S., OpenAI shipped GPT-5.6 and ChatGPT Work, then had to defend itself against an Apple trade-secret lawsuit and quietly fork Git. Meta released Muse Spark 1.1 while retreating from Muse Image's consent controversy. Elon Musk's SpaceXAI pushed Grok 4.5 as a coding and agentic tool. In Europe, ZML and a wave of chip challengers are trying to loosen Nvidia's grip on inference, while QuantumDiamonds and Oratomic raised large rounds for hardware sovereignty. In China, DeepSeek signaled custom silicon and continued open-weight expansion. And in India, an AI security failure at Coal India showed how quickly model deployments can outrun operational security.

The common thread is pressure on margins and control. Labs are racing to own more of the stack — models, chips, developer tools, and distribution — because the alternative is paying a growing tax to Nvidia, cloud providers, or foreign suppliers. That pressure produces both innovation and fragility: vertical integration can lower costs, but it also concentrates risk inside organizations that are already moving fast.

Compute Watch

The neocloud model is showing financial strain. CoreWeave reported first-quarter revenue of $2.08 billion, up 112% year over year, but spent $7.7 billion in capex and burned $4.71 billion in free cash flow. Its debt load rose to roughly $24.9 billion. Nebius is in better shape, with $9.37 billion in cash against $8.45 billion in debt, but it also expects to spend around $22.5 billion on capex this year.

Both companies rely on GPU-backed delayed-draw term loans and major customer contracts to keep funding their build-outs. I/O Fund's analysis points to a circular dynamic: Nvidia supports the neoclouds financially, the neoclouds buy Nvidia's latest systems, and hyperscalers rent the capacity to avoid putting the capex on their own balance sheets. The arrangement works only as long as AI demand keeps rising and interest rates do not break the debt structure. CoreWeave's interest payments already equal about 26% of revenue, and that ratio is expected to climb.

From the Lab

Anthropic published research this week that gives a clearer look inside Claude Opus 4.6 as it reasons. Using a technique called the Jacobian lens, the company identified a hidden "J-space" containing words related to what the model is likely to say next. The findings range from mundane to unnerving: sometimes the model's internal state diverges from the answer it produces, suggesting that monitoring this hidden space could become a new safety and control tool.

The paper is posted on Anthropic's Transformer Circuits site. The work matters because interpretability has lagged behind capability; if researchers can read a model's latent reasoning in real time, they may be able to catch misalignment before it is expressed. Anthropic has also partnered with Neuronpedia to let outsiders inspect the model's internals.

The View

Three threads connect the week's stories. First, the hardware layer is fragmenting. DeepSeek wants its own inference chip, ZML wants to make inference hardware interchangeable, and neoclouds are betting the farm on Nvidia-powered data centers. All three moves are responses to the same problem: the cost of running models at scale is eating margins.

Second, the software layer is consolidating around agents. Meta Spark, Grok 4.5, GPT-5.6's agentic coding efficiency, and OpenAI's Git fork all point to tools that help models act over longer horizons. The competition is shifting from raw benchmark scores to which system can reliably carry out tasks across code, documents, and interfaces.

Third, trust is becoming a structural cost. Google's ad-labeling, Meta's consent reversal, Apple's lawsuit, and Coal India's security failure show that the same speed that lets labs ship features also lets operational and ethical gaps scale. The companies that treat trust as infrastructure — not PR — may find it cheaper in the long run than the alternatives.

The Miss

A paper introducing UniClawBench, a benchmark for proactive agents operating in dynamic real-world settings, arrived on arXiv this week with less attention than the latest model launches. The benchmark uses 400 bilingual tasks evaluated in live Docker containers, testing whether agents can act proactively rather than merely respond to prompts. As the industry moves from chatbots to agents, the gap between static leaderboards and real-world reliability is where most deployments will actually fail.

Pull Quotes

"DeepSeek designs both the model and, now, the chip that runs it." — Proactive Investors, July 9, 2026.

"A closed society is conquering the world with open-source AI models." — Nathan Gardels, Noema Magazine, July 10, 2026.

"US software development job postings have grown by almost 15%." — HiringLab, July 8, 2026.

"No authentication necessary." — Eaton Works, on Coal India's DigiCoal dashboard, July 8, 2026.

OpenAI's public Git fork — primary artifact. https://github.com/openai/git

DeepSeek is designing its own chip — Proactive Investors, July 9, 2026. https://www.proactiveinvestors.com/companies/news/1095178/deepseek-makes-pivot-that-should-put-silicon-valley-on-high-alert-1095178.html

Nvidia, CoreWeave, and Nebius: Inside the Circular Financing of the GPU Boom — I/O Fund, June 12, 2026. https://io-fund.com/ai-stocks/nvidia-coreweave-nebius-circular-financing-gpu-boom

Meta says its new AI model is ready to compete on coding — The Verge, July 9, 2026. https://www.theverge.com/ai-artificial-intelligence/963193/meta-muse-spark-model-api

Google will now disclose which ads are made with AI — TechCrunch, July 9, 2026. https://techcrunch.com/2026/07/09/google-will-now-disclose-which-ads-are-made-with-ai/

Anthropic's hidden-space paper — Transformer Circuits, July 9, 2026. https://transformer-circuits.pub/2026/workspace/

Inside an AI coal mine security camera network powered by plaintext passwords — Eaton Works, July 8, 2026. https://eaton-works.com/2026/07/08/coal-india-camera-hack/

The infrastructure layer is where this week's bets are being placed.