The Stack Was Already There. The Controls Gave It Customers.
Table of Contents
Ask what US chip export controls did to Nvidia’s position in China and you get three answers, all from credible sources, none of which agree.
Jensen Huang says the share dropped to zero and that the policy has already largely backfired. IDC, reported by Reuters, put Nvidia at about 55% of the AI accelerator cards shipped in China across full-year 2025, some 2.2 million units out of roughly four million. Bernstein, relayed through The Economist, has Nvidia at about 40% in 2025 falling to roughly 8% in 2026, with Huawei rising to about half. Nvidia’s own 10-K reports $19.68 billion of revenue from China including Hong Kong for the fiscal year ending January 2026, roughly 9% of sales.
Those four numbers measure four things at four dates: a chief executive’s characterisation of current sales, a trailing shipment count, a forecast, and audited revenue booked by customer headquarters. Anyone quoting one of them without saying which is measuring one thing and calling it another.
Huawei’s accelerator stack existed years before the controls and went nowhere. The controls supplied the demand that turns a second-tier platform into a durable one. A second platform at scale means a second privileged-code surface, and Western security practice has barely looked at it.
This was never one policy
The reversals matter as much as the restrictions, because a control regime that changes direction repeatedly transmits a different signal than a stable one.
January 2025: BIS issues the Framework for Artificial Intelligence Diffusion, a worldwide licence requirement with a three-tier country quota system, compliance date 15 May.
April 2025: the H20, a part Nvidia designed specifically to comply with the existing thresholds, is banned. Nvidia takes a $4.5 billion charge on inventory and purchase obligations and guides to an $8 billion revenue hit.
13 May 2025: BIS announces it will rescind the Diffusion Rule and instructs enforcement staff not to enforce it. Licences follow over the summer and H20 shipments resume, under an arrangement Nvidia’s own filing describes as US officials expecting the government to receive 15% or more of licensed H20 revenue, with no regulation codifying it.
8 December 2025: the President announces H200 exports to China will be permitted. The same day, federal prosecutors unseal an investigation into a China-linked smuggling network involving around $160 million in H100 and H200 chips.
13 to 15 January 2026: BIS issues a final rule moving H200 and AMD MI325X exports to China from presumption of denial to case-by-case review, conditioned on third-party testing in the United States, certifications about domestic supply, know-your-customer and remote-access safeguards, and a volume cap holding China-bound shipments to half of US domestic sales. A presidential proclamation on the 14th adds a separate instrument, a 25% Section 232 duty on covered chips. Both are operative from the 15th. The 15% and the 25% are different mechanisms, an uncodified revenue expectation and an import tariff, and most coverage fuses them.
February 2026: a draft replacement for the Diffusion Rule goes to regulatory review and is withdrawn within weeks.
A Chinese procurement officer reads that sequence differently from a Washington analyst. The commercially relevant boundary moved repeatedly, and so did the mechanism enforcing it. What did not move was the demonstrated willingness to move both. Buyers whose supply depends on another government’s politics from one month to the next look elsewhere. Then they pay to make the alternative good enough.
There is still no replacement rule
Be precise here, because the obvious version of this point is wrong. There is a rule. The January 2026 final rule has a published text you can read, with conditions you can check yourself against.
What does not exist is a successor to the global framework. BIS announced the Diffusion Rule’s rescission in May 2025. No formal rescission has been published since, and the framework remains codified in the Export Administration Regulations while explicitly unenforced. A draft replacement reached regulatory review in February 2026 and was withdrawn in March.
So a planner works with five instruments at once: a published rule for one class of export, a tariff created by proclamation, an uncodified expectation that a share of licensed revenue reaches the Treasury, a codified framework that remains unenforced, and enforcement discretion across all of it. They change faster than a procurement cycle.
The Fable and Mythos directive used the same instrument: a control imposed by individual arrangement rather than by a rule anyone else could read in advance. In chip exports Washington runs that kind of arrangement alongside published rules, and a planner has to satisfy both.
Congress has noticed. The AI OVERWATCH Act was introduced in the House as H.R. 6875 on 18 December 2025 by Foreign Affairs Committee chairman Brian Mast, and reported out of that committee on 21 January 2026 by 42 votes to 2, with one member voting present. A Senate companion, S. 4456, followed on 30 April 2026 from Jim Banks, cosponsored by Elizabeth Warren and four others, and was read twice and referred to the Senate Banking Committee the same day. Neither has moved since, and neither has become law.
The silicon is behind and the toolchain costs months
Huawei’s Ascend line has a guaranteed market and is scaling into it. Roughly 812,000 Ascend chips shipped in 2025 on IDC’s numbers. AI processor revenue around $7.5 billion in 2025, projected near $12 billion in 2026. The Ascend 950PR entered production in early 2026, and Huawei claims a multiple of the H20’s FP4 performance, though the H20 was a deliberately constrained export part and beating it is not the same as competing at the frontier.
Now the other column, which is longer.
A Council on Foreign Relations model published in December 2025 put Huawei’s 2026 aggregate AI compute at around 4% of Nvidia’s, even on aggressive production assumptions, with no Huawei part exceeding the relevant H200 threshold before the Ascend 960 in late 2027. CFR states the assumptions behind that figure. It is a model output, not a measurement, and the sturdiest number in this section.
Two documented cases show what the switch costs in practice. The Financial Times reported in August 2025 that DeepSeek, pushed toward Ascend for training its R2 model, hit persistent technical problems and moved training back to Nvidia while keeping Ascend for inference. iFlytek’s chairman said publicly in June 2025 that using domestic chips including the 910B added about three months to model development compared with Nvidia’s more mature tooling. Both costs came from the software around the chip.
On power, be careful with a figure that circulates in the wrong units. The comparison people quote is a Huawei CloudMatrix 384 system against an Nvidia GB200 NVL72. SemiAnalysis put the CloudMatrix at 559 kW against the NVL72’s 145 kW, delivering around 300 PFLOPs of dense BF16 against roughly 180, and calculated the penalty at two and a half times worse power per FLOP. Coverage that recomputed it from the same figures lands at 2.3 times. Either number is a system-level ratio of about two and a half times worse per unit of compute, and four is the one that circulates.
Huawei markets the Ascend 950 series as an advance, and the comparison it publishes is native FP4 inference against the H20, an export-constrained part with no native FP4 path. CFR reads the same roadmap as a regression against the 910C, because its metric is total processing performance and memory bandwidth. Both are right on their own yardstick.
Huawei sells into a protected market. Its parts are behind at the frontier, less efficient at system scale, and dependent on SMIC and on domestic high-bandwidth memory that Samsung and SK Hynix still outperform. If the controls were meant to slow Chinese frontier training, the evidence is that they are still doing it.
Huawei shipped the stack in 2019 and waited for customers
Here is the thing most coverage gets backwards, including the version of this article I first drafted. The controls did not create Huawei’s software stack.
Huawei launched the Ascend 910 and the MindSpore framework on 23 August 2019, and the launch materials describe a full-stack portfolio including CANN as the chip enablement layer. That is three years before the first US controls on Nvidia sales to China. Eric Xu said at the time that open-sourcing MindSpore was meant to attract the developers who would buy Huawei hardware. The strategy was complete and the customers were not there.
Restricting Nvidia sales put the customers there.
Chinese teams still write PyTorch, and Huawei works to keep the upper layers familiar. The divergence starts below the framework: firmware, device driver, the CANN runtime and its operator libraries, the collective-communication layer, the compiler, and the container and orchestration plumbing are all different. Every painful port from Nvidia leaves behind operators, kernels, deployment tooling and engineers who know how to use them. The application still says PyTorch at the top and calls CANN underneath.
Huang has an obvious commercial interest in Washington reopening China, so his framing is advocacy. He attributes CUDA’s advantage not to the silicon but to the developer population and twenty years of libraries. The controls, on his account, are eroding it. Nvidia’s own published CUDA developer figure runs to several million after twenty years, quoted by the company at four million and more recently higher. Huawei’s chief financial officer Meng Wanzhou has claimed over four million developers and more than 3,000 partners for Ascend. Both are self-reported by the company that benefits from them, on no published methodology, and the comparison is between two marketing numbers rather than two measurements.
The demand floor is the mechanism, and the cleanest evidence for it is not the largest number. Reuters reported in November 2025 that China barred foreign AI chips from state-funded data centres, with some projects required to remove or cancel foreign chips already ordered. In May 2026, nine domestically developed AI processors from seven vendors, including Huawei, Alibaba’s T-Head, Biren and Moore Threads, cleared a state security review. Bloomberg reported in June 2026 that Beijing was drafting a five-year plan worth around two trillion yuan, roughly $295 billion, requiring at least 80% of core technologies including chips to come from domestic suppliers. That is a draft at an early stage rather than an appropriated budget, and the procurement bar and the certification list above carry more of this argument than the headline figure does.
DeepSeek’s V4, released in April 2026, was reported as explicitly optimised for Ascend. That is a frontier lab spending its own engineering time on Huawei’s toolchain.
Hardware generations depreciate. Compatibility code, tuned operators, deployment tooling, documentation and the engineers who learned it can outlive the generation that forced someone to build them. The controls bought real time against Chinese frontier training. They spent some of it supplying a reason to build the thing that does not depreciate on the same schedule.
What a second stack means for a security team
The claim is not that Chinese silicon is insecure, and it is not that nobody has ever secured anything other than CUDA. ROCm, TPUs and Trainium all exist. The defensible claim is narrower and worse: Western enterprise tooling, public vulnerability research and operational muscle are concentrated around one accelerator stack to a degree that nobody planned and few have measured.
Three exposures need separating here, because the controls differ across them.
You self-host Ascend. Then you have a genuinely new privileged surface on machines that hold model weights: firmware, kernel and device drivers, a user-space runtime, a compiler, operator and collective-communication libraries, management agents, container images and Kubernetes device plugins. Every one of those needs patching, monitoring and someone who can read an advisory about it.
A vendor runs it behind an API. Then none of that runs in your environment and the problem is supplier assurance: provenance, isolation, data handling, vulnerability-management commitments, jurisdictional exposure and whether that vendor can investigate an incident for you.
You download a model. Plain weights are mostly portable and carry no toolchain with them, so do not overstate this one. The dependency arrives when the repository ships custom operators, backend-specific compiled kernels, a quantisation implementation, serving code or a container, which is the same inherited-risk problem as anything else in a model repository and not a property of the weights.
You will meet at least one of the three before you expect to, and usually not because you bought it. Compute itself became an object of state competition, and the software layer is cheaper to enter and harder to reverse.
What to do
Three things, and the first two cost nothing.
Find out whether anything in your estate already runs on the other stack. Not the chips you bought, which you know about. Ask your managed-inference vendors what accelerator serves your requests, and grep your model repositories for custom operators and backend-specific kernels. Most organisations cannot answer either question today.
Accelerator choice determines your driver surface, your privileged-code footprint, and the vendor whose vulnerability disclosures you read. It belongs in architecture review as a concentration decision, alongside the cost model.
If you operate in or sell into a market where domestic-silicon requirements are appearing, get someone who can actually read the other toolchain. That capability takes a year to build and cannot be bought during an incident. This is the same conclusion the Hugging Face responders reached about model access from a completely different direction, and when unrelated pressures point at the same investment, it is usually real.
The argument in Washington is still about whether the controls worked, and it is scored in Nvidia’s share of a market it has effectively left. That is the wrong scoreboard. The controls did constrain Chinese frontier training and they are still doing it. They also removed the reason for a large developer population to keep building on the American stack, at the moment Beijing began directing state-funded compute toward domestic accelerators. Export permissions can change tomorrow. Operator libraries, deployment tooling and the engineers trained on them do not change on that schedule.
In the early 2000s, running emerging-technology risk labs at CyberAgency, a defence client asked my team to break the AI systems they planned to put into weapons. We did. That is where my work on AI security started, two decades before the current wave of attention. I kept at it through risk labs at IBM, Accenture, PwC and KPMG. In 2016 I co-wrote a book on AI and leadership. My commercial work today is quantum, at Applied Quantum, which is why this site sells nothing.