REVIEW 3 major objections 6 minor 40 references
Whack-a-Chip: The Futility of Hardware-Centric Export Controls
T0 review · 3 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper argues that U.S. hardware-centric export controls are a leaky proxy for limiting Chinese AI capability, because Chinese labs both keep reaching restricted chips and have learned to train state-of-the-art models on chips that are…
desk verdict A readable, policy-relevant brief with a genuinely useful code-signature method, but the central H20 claim rests on a self-report and the title oversells “futility.” read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is the combination of a code-signature pipeline and a software-efficiency stack. The code-signature pipeline narrows candidate GPUs from Tencent's released training scripts by checking for NCCL (NVIDIA-only), bfloat16 support (Ampere or later), and GPUDirect RDMA with InfiniBand (datacenter GPUs only), leaving a small set that the paper admits cannot be further discriminated by code alone. The efficiency stack—mixture-of-experts routing, bfloat16 mixed-precision training, quantization, fully sharded data parallelism (DeepSpeed ZeRO Stage 3), and RDMA networking—is what makes a throttled chip sufficient for frontier-scale training. Together they convert the question 'which chips did China obtain?' into 'what can be trained on chips China is allowed to have?'
What would settle it
Find Tencent's actual cluster job logs, energy or telemetry traces, or procurement records for Hunyuan-Large's training run; or independently reproduce a comparable model on an H20 cluster with the same open-source stack and compare loss curves and benchmark scores to a reproduction on H100 or A100. If the H20 reproduction falls short of the published Hunyuan-Large results, or if the job logs name a different GPU, the paper's core claim fails.
Extended reading notes
Core claim
The paper's central discovery is the first concrete public case of a leading PRC lab training a state-of-the-art model on an export-control-compliant GPU: Tencent's Hunyuan-Large, a 52-billion-active-parameter mixture-of-experts transformer, was trained on NVIDIA H20s. The H20 has fewer cores and lower throughput than the H100 but carries 96GB of VRAM, and the paper argues that high VRAM plus software techniques—MoE architecture, bfloat16 mixed precision, quantization, fully sharded data parallelism via DeepSpeed's ZeRO Stage 3, and GPUDirect RDMA over InfiniBand—makes it competitive for frontier training. The paper also performs a code-signature analysis of Tencent's released training code, eliminating AMD GPUs (NCCL usage), pre-Ampere GPUs (bfloat16), and consumer GPUs (GPUDirect RDMA configuration), then concedes the signature cannot distinguish H20 from H100 or A100. It corroborates a broader evasion picture: Tencent publicly used A100s for Hunyuan-DiT and H100s for GameGen-X, and reporting cited in the paper describes stockpiling, underground markets, and cloud-rental access to restricted chips. The conclusion is that an export-control strategy built on hardware performance thresholds is structurally leaky.
Load-bearing premise
The load-bearing premise is that Tencent's README statement is accurate when it says Hunyuan-Large was trained on NVIDIA H20s; the paper's code-signature analysis cannot actually distinguish the H20 from the H100 or A100, so if the self-report is wrong, the central example loses its evidence.
Editorial extensions
If this is right
- Closing the H20 loophole by adding it to export-control lists will only push PRC labs to the next sub-threshold GPU, because the same software techniques transfer to any chip with enough memory and bandwidth.
- Performance-threshold metrics such as ECCN 3A090's total processing performance are poor predictors of what models can be trained; a memory-rich, compute-modest chip can still produce state-of-the-art results when combined with mixture-of-experts and sharding.
- Continuous monitoring of open-source codebases, weights, and papers from PRC labs can reveal both hardware use and efficiency techniques, giving a more realistic picture of PRC capability than hardware tracking alone.
- A verification regime built on use-case-specific, continuously updated benchmarks—trained on constrained hardware—could replace threshold-based chip controls as the practical way to gauge and limit AI risk.
- Even if export controls were expanded to all currently available chips, PRC labs would retain stockpiled and domestic alternatives, so the marginal effect of further hardware restrictions is small.
Reading between the lines
- Editorial inference: if algorithm-level efficiency is the binding variable, then the most effective control point may be access to training software, libraries, and algorithmic knowledge rather than chips—though the paper itself stresses that AI software is dual-use and protected speech, so this path is legally fraught.
- Editorial inference: the H20's 96GB VRAM means the word 'weaker' hides a real advantage for memory-bound sharded training; future controls should measure memory capacity and interconnect speed, not just FLOPS, or they will keep missing the effective metric.
- Editorial inference: a testable extension would be to run a controlled scaling study training the same mixture-of-experts architecture on H20, A100, and H100 with identical software to separate hardware effects from software effects; the public Hunyuan-Large release makes such a replication possible.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that U.S. hardware-centric export controls on semiconductors are losing their effectiveness because Chinese AI labs can both obtain restricted chips through stockpiling and gray/black markets and, more importantly, train state-of-the-art models on export-control-compliant hardware such as NVIDIA's H20. The central case study is Tencent's Hunyuan-Large, which the authors claim was trained on H20 GPUs and achieves state-of-the-art results. The paper combines a feature-elimination analysis of training code (NCCL, bfloat16, GPUDirect RDMA) with documentation of other access avenues and policy discussion. It concludes that performance-threshold-based controls are a 'whack-a-mole' strategy and proposes alternative policy directions such as a verification regime and benchmark-based monitoring.
Significance. If the H20 attribution and the interpretation were firmly established, this would be a timely and policy-relevant data point: a leading Chinese lab producing a state-of-the-art open-source model on export-control-compliant GPUs would challenge the assumption that hardware thresholds alone can gate AI capability. The paper usefully catalogs software techniques (MoE, bfloat16, DeepSpeed ZeRO-3, GPUDirect RDMA) and compiles public evidence about Chinese access to restricted chips. The code-signature elimination approach, while limited, is a transparent and reproducible method for narrowing candidate hardware from public artifacts. The policy discussion about benchmarks and verification is constructive. However, as detailed in the major comments, the paper's empirical core depends on a self-reported hardware attribution and the conclusions are broader than the evidence supports.
major comments (3)
- [Abstract and Section 1 (Introduction), p. 1] The abstract's claim that the paper presents 'the first concrete, public evidence of how leading PRC AI labs evade and circumvent U.S. export controls' is internally inconsistent with Section 2, which describes the H20 as 'an export control-compliant chip' that is 'explicitly allowed under U.S. export controls.' Using a compliant chip is not evasion or circumvention. The H20 example can support an argument about the limits of performance-threshold-based controls, but it cannot support an evasion claim. The paper should reframe the H20 discussion as an example of legal adaptation to the rules, and keep the evasion claim separate, tied to the stockpiling and black-market evidence in Section 3.3.
- [Section 3.2, pp. 3-4] The central empirical claim that Hunyuan-Large was trained on H20 relies entirely on the project README's self-report. The paper's independent code-signature analysis narrows the hardware only to 'a few, modern data center GPUs' and explicitly concedes 'there are no meaningful feature differences between the Hopper, Ada, and Ampere architectures (besides more cores) to discriminate between them.' Since H100 (Hopper) and A100 (Ampere) are data center GPUs that Tencent has publicly acknowledged using, the analysis cannot exclude those alternatives. The authors should either provide independent corroboration of the H20 attribution (e.g., verified cluster configuration, procurement records, or a distinguishable technical fingerprint) or explicitly frame the H20 premise as an assumption. As written, the paper's key example is an unverified self-report, and this weakness is load-bearing because the conclusion about 'futility' rests on it.
- [Section 4, pp. 4-5] The policy conclusion that 'Advances in machine learning have eroded the moat' and that hardware-centric export controls are futile is broader than the evidence supports. The paper offers a single existence proof: one model trained on a compliant chip achieves good benchmark results. It does not compare training time, cost, energy, or resulting model quality against the same techniques applied to H100/A100, nor does it control for dataset scale, model architecture, or engineering effort. Showing that a state-of-the-art model can be trained on H20 does not establish that hardware restrictions are not binding. The authors should either add a comparative analysis of H20 versus unrestricted hardware under otherwise similar conditions, or substantially weaken the 'futility' framing to 'one example of a compliant chip supporting competitive model training.'
minor comments (6)
- [Section 2, p. 2] The sentence 'the H20 boasts 96GB of VRAM compared to the H100's common 80GB configuration' is qualified only in a footnote; since the main text otherwise states a contrast that the footnote partly retracts, the qualification should appear in the main text.
- [Section 3.2, p. 4] The phrase 'reverse engineer candidate GPUs' is informal; the method is feature-based elimination from public code, not reverse engineering of hardware or binaries. A more precise term such as 'hardware inference from code signatures' would be accurate.
- [Section 3.3, p. 4] The sentence 'These public acknowledgments from Tencent corroborate reporting...' is too strong: the acknowledgments are public claims in papers and websites, not independent evidence of illegal acquisition. The word 'corroborate' should be replaced with 'are consistent with' or similar.
- [Figures 3 and 4] Some GPU names in the figures (e.g., T40, L40) are not defined or discussed in the text; consider a table of microarchitectures and supported features for clarity.
- [Section 3.3, p. 4] There is a typo, 'as well well as,' in the sentence beginning 'By cross-reference reporting with the publicly available papers...'.
- [References] The performance comparison for the H20 versus H100 (reference [26]) is from a secondary news/rumor source; citing official NVIDIA or BIS documentation would strengthen the technical claim.
Circularity Check
No significant circularity; the paper's central evidence is an external self-report and code-signature analysis, not a fitted or self-referential derivation.
full rationale
The paper's core claim is that Tencent trained Hunyuan-Large on export-compliant NVIDIA H20 GPUs and that this demonstrates software-driven efficiency gains eroding hardware-centric export controls. The evidence for the H20 attribution is the project README's explicit statement, plus a code-signature analysis that narrows the hardware family but explicitly concedes it cannot distinguish H20 from H100 or A100. This is an evidentiary weakness, not a circularity: the README is an external public artifact, and the code-signature analysis is not used to define the conclusion. No equation, fitted parameter, or prediction reduces by construction to an input. The paper contains no quantitative predictions at all, and its argument is qualitative. The self-citations, notably reference [10] in the discussion of export-control thresholds and reference [11] on model efficiency trends, are background supporting material rather than load-bearing derivation; removing them would not collapse the argument because the Hunyuan-Large case study and the README citation carry the empirical weight independently. There is no imported uniqueness theorem, no ansatz smuggled in by citation, and no renaming of a known result. The skeptical concern about relying on Tencent's self-reported hardware is a correctness or verification risk, not a circularity, and under the review rules it does not raise the circularity score.
Assumptions & free parameters
assumptions (4)
- domain assumption H20 is export-control-compliant and its performance and VRAM specifications are as reported by secondary sources.
- domain assumption Code signatures (NCCL, bfloat16, GPUDirect RDMA over InfiniBand) reliably constrain training hardware to NVIDIA data-center GPUs.
- ad hoc to paper Tencent's README statement that Hunyuan-Large was trained on H20 is truthful and accurate.
- domain assumption One successful SOTA training run on compliant hardware supports the broad conclusion that hardware-centric export controls are losing effectiveness.
Cite this review
Pith. "Pith review of Whack-a-Chip: The Futility of Hardware-Centric Export Controls." pith.science (2026). https://pith.science/paper/B5UTWK3L
@misc{pith2026241114425,
author = {Pith},
title = {Pith review of: Whack-a-Chip: The Futility of Hardware-Centric Export Controls},
year = {2026},
howpublished = {\url{https://pith.science/paper/B5UTWK3L}},
note = {Machine review of arXiv:2411.14425}
}
read the original abstract
U.S. export controls on semiconductors are widely known to be permeable, with the People's Republic of China (PRC) steadily creating state-of-the-art artificial intelligence (AI) models with exfiltrated chips. This paper presents the first concrete, public evidence of how leading PRC AI labs evade and circumvent U.S. export controls. We examine how Chinese companies, notably Tencent, are not only using chips that are restricted under U.S. export controls but are also finding ways to circumvent these regulations by using software and modeling techniques that maximize less capable hardware. Specifically, we argue that Tencent's ability to power its Hunyuan-Large model with non-export controlled NVIDIA H20s exemplifies broader gains in efficiency in machine learning that have eroded the moat that the United States initially built via its existing export controls. Finally, we examine the implications of this finding for the future of the United States' export control strategy.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
Cheaper, Better, Faster, Stronger
Mistral AI. Cheaper, Better, Faster, Stronger. https://mistral.ai/news/mixtral-8x22b/, April 2024
work page 2024
-
[2]
Updated October 7 Semiconductor Export Controls
Emily Benson. Updated October 7 Semiconductor Export Controls. https://www.csis.org/analysis/updated-october- 7-semiconductor-export-controls, Wed, 10/18/2023 - 12:00
work page 2023
-
[3]
Improving Image Generation with Better Captions
James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, Wesam Manassra, Prafulla Dhariwal, Casey Chu, Yunxin Jiao, and Aditya Ramesh. Improving Image Generation with Better Captions. Technical report, OpenAI, October 2023
work page 2023
-
[4]
Bureau of Industry and Security. Commerce Strengthens Restrictions on Advanced Computing Semiconductors, Semiconductor Manufacturing Equipment, and Supercomputing Items to Countries of Concern | Bureau of Industry and Security. https://www.bis.gov/press-release/commerce- strengthens-restrictions-advanced-computing-semiconductors- semiconductor
-
[5]
Bureau of Industry and Security. Implementation of Additional Export Controls: Certain Advanced Computing and Semiconductor Manufacturing Items; Supercomputer and Semiconductor End Use; Entity List Modification. https://www.federalregister.gov/documents/2022/10/13/2022- 21658/implementation-of-additional-export-controls- certain-advanced-computing-and-sem...
work page 2022
-
[6]
GameGen-X: Interactive Open-world Game Video Generation, November 2024
Haoxuan Che, Xuanhua He, Quande Liu, Cheng Jin, and Hao Chen. GameGen-X: Interactive Open-world Game Video Generation, November 2024. 7 . DeepSeek AI. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model, June 2024
work page 2024
-
[8]
Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W. Mahoney, and Kurt Keutzer. A Survey of Quantization Methods for Efficient Neural Network Inference, June 2021
work page 2021
-
[9]
Amir Gholami, Zhewei Yao, Sehoon Kim, Coleman Hooper, Michael W. Mahoney, and Kurt Keutzer. AI and Memory Wall. IEEE Micro, 44(3):33–39, May 2024
work page 2024
Show all 40 references
-
[10]
Ritwik Gupta and Andrew W. Reddie. Accelerating the Evolution of AI Export Controls | TechPolicy.Press. https://www.techpolicy.press/accelerating-the-evolution-of- ai-export-controls/, September 2023
2023
-
[11]
Ritwik Gupta, Leah Walker, Rodolfo Corona, Stephanie Fu, Suzanne Petryk, Janet Napolitano, Trevor Darrell, and Andrew W. Reddie. Data-Centric AI Governance: Addressing the Limitations of Model- Focused Policies, September 2024
2024
-
[12]
Desperate Chinese factories repurpose last-gen RTX 3090 GPUs into AI accelerators to skirt US export ban
Christopher Harper. Desperate Chinese factories repurpose last-gen RTX 3090 GPUs into AI accelerators to skirt US export ban. https://www.tomshardware.com/news/rtx-4090-blowers- stripped-for-ai, November 2023
2023
-
[13]
Measuring Massive Multitask Language Understanding, January 2021
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring Massive Multitask Language Understanding, January 2021
2021
-
[14]
Introducing native PyTorch automatic mixed precision for faster training on NVIDIA GPUs, July 2020
Mengdi Huang, Chetan Tekur, and Michael Carilli. Introducing native PyTorch automatic mixed precision for faster training on NVIDIA GPUs, July 2020
2020
-
[15]
Nayak, Jack Merullo, Stephen H
Apoorv Khandelwal, Tian Yun, Nihal V. Nayak, Jack Merullo, Stephen H. Bach, Chen Sun, and Ellie Pavlick. 100K or 100 Days: Trade-offs when Pre-Training with Academic Resources, October 2024
2024
-
[16]
When Ensembling Smaller Models is More Efficient than Single Large Models, May 2020
Dan Kondratyuk, Mingxing Tan, Matthew Brown, and Boqing Gong. When Ensembling Smaller Models is More Efficient than Single Large Models, May 2020. 17 . Zhimin Li, Jianwei Zhang, Qin Lin, Jiangfeng Xiong, Yanxin Long, Xinchi Deng, Yingfang Zhang, Xingchao Liu, Minbin Huang, Zed...
2020
-
[18]
China unveils domestic GPU ecosystem to reduce Nvidia dependence
Amanda Liang and Jerry Chen. China unveils domestic GPU ecosystem to reduce Nvidia dependence. https://www.digitimes.com/news/a20241007PD216/huawei- china-mobile-nvidia-chips-gpu.html, October 2024
2024
-
[19]
WSJ News Exclusive | China’s Top Nuclear- Weapons Lab Used American Computer Chips Decades After Ban
Liza Lin and Dan Strumpf. WSJ News Exclusive | China’s Top Nuclear- Weapons Lab Used American Computer Chips Decades After Ban. Wall Street Journal, January 2023
2023
-
[20]
The Llama 3 Herd of Models, August 2024
Llama Team at Meta. The Llama 3 Herd of Models, August 2024
2024
-
[21]
Huawei’s bug-ridden software hampers China’s efforts to replace Nvidia in AI
Ryan McMorrow, Eleanor Olcott, and Tina Hu. Huawei’s bug-ridden software hampers China’s efforts to replace Nvidia in AI. Financial Times, September 2024
2024
-
[22]
Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and Better
Gaurav Menghani. Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and Better. ACM Comput. Surv., 55(12):259:1–259:37 , March 2023
2023
-
[23]
Mixed Precision Training, February 2018
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu. Mixed Precision Training, February 2018
2018
-
[24]
What To Do When You Can’t Get Nvidia H100 GPUs
Timothy Prickett Morgan. What To Do When You Can’t Get Nvidia H100 GPUs. https://www.nextplatform.com/2023/11/17/what-to- do-when-you-cant-get-nvidia-h100-gpus/, November 2023
2023
-
[25]
Huawei’s HiSilicon Can Compete With Nvidia GPUs In China
Timothy Prickett Morgan. Huawei’s HiSilicon Can Compete With Nvidia GPUs In China. https://www.nextplatform.com/2024/08/13/huaweis- hisilicon-can-compete-with-nvidia-gpus-in-china/, August 2024
2024
-
[26]
NVIDIA’s China-Compliant H20 GPU Has 41% Fewer Cores & 28% Lower Performance Versus Top Hopper H100 Config
Hassan Mujtaba. NVIDIA’s China-Compliant H20 GPU Has 41% Fewer Cores & 28% Lower Performance Versus Top Hopper H100 Config. https://wccftech.com/nvidia-china-compliant-h20-gpu-41- percent-fewer-cores-lower-performance-vs-top-hopper-h100/, July 2024. 27 . Hassam Nasir. Chinese ...
2024
-
[28]
Chinese AI groups use cloud services to evade US chip export controls
Eleanor Olcott, Demetri Sevastopulo, and Qianer Liu. Chinese AI groups use cloud services to evade US chip export controls. Financial Times, March 2023
2023
-
[29]
xView3-SAR: Detecting Dark Fishing Activity Using Synthetic Aperture Radar Imagery
Fernando Paolo, Tsu-ting Tim Lin, Ritwik Gupta, Bryce Goodman, Nirav Patel, Daniel Kuster, David Kroodsma, and Jared Dunnmon. xView3-SAR: Detecting Dark Fishing Activity Using Synthetic Aperture Radar Imagery. Advances in Neural Information Processing Systems, 35:37604–37616, ...
2022
-
[30]
Comparing Quantized Performance in Llama Models
Nicky Pochinkov. Comparing Quantized Performance in Llama Models. https://www.lesswrong.com/posts/qmPXQbyYA66DuJbht/comparing- quantized-performance-in-llama-models, July 2024
2024
-
[31]
ZeRO: Memory Optimizations Toward Training Trillion Parameter Models, May 2020
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. ZeRO: Memory Optimizations Toward Training Trillion Parameter Models, May 2020. Whack-a-Chip 7
2020
-
[32]
DeepSpeed: System Optimizations Enable Training Deep Learning Models with Over 100 Billion Parameters
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. DeepSpeed: System Optimizations Enable Training Deep Learning Models with Over 100 Billion Parameters. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD ’20...
-
[33]
Mixture of Experts Explained, December 2023
Omar Sanseviero, Lewis Tunstall, Philipp Schmid, Sourab Mangrulkar, Younes Belkada, and Pedro Cuenca. Mixture of Experts Explained, December 2023
2023
-
[34]
China’s chip imports boom as the country stockpiles before anticipated new sanctions, struggles to become self-sufficient
Anton Shilov. China’s chip imports boom as the country stockpiles before anticipated new sanctions, struggles to become self-sufficient. https://www.tomshardware.com/tech-industry/chinas-chip- imports-boom-as-the-country-stockpiles-before-anticipated- new-sanctions-struggles-t...
2024
-
[35]
New Tools Are Needed to Address the Risks Posed by AI-Military Integration
Sarah Shoker, Andrew Reddie, Alan Hickey, and Leah Walker. New Tools Are Needed to Address the Risks Posed by AI-Military Integration. https://www.lawfaremedia.org/article/new-tools-are- needed-to-address-the-risks-posed-by-ai-military-integration, March 2024
2024
-
[36]
Gaming Hardware - China | Statista Market Forecast
Statista. Gaming Hardware - China | Statista Market Forecast. https://www.statista.com/outlook/amo/media/games/gaming- hardware/china. 37 . Xingwu Sun, Yanfeng Chen, Yiqing Huang, Ruobing Xie, Jiaqi Zhu, Kai Zhang, Shuaipeng Li, Zhen Yang, Jonny Han, Xiaobo Shu, Jiahao Bu, Zho...
2024
-
[38]
CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge. In Jill Burstein, Christy Doran, and Thamar Solorio, editors, Proceedings of the 2019 Conference of the North American Chapter of ...
2019
-
[39]
Efficient Large Language Models: A Survey, May 2024
Zhongwei Wan, Xin Wang, Che Liu, Samiul Alam, Yu Zheng, Jiachen Liu, Zhongnan Qu, Shen Yan, Yi Zhu, Quanlu Zhang, Mosharaf Chowdhury, and Mi Zhang. Efficient Large Language Models: A Survey, May 2024
2024
-
[40]
Exclusive: Chinese firms stockpile high-end Samsung chips as they await new US curbs, say sources
Heekyong Yang, Fanny Potkin, and Karen Freifeld. Exclusive: Chinese firms stockpile high-end Samsung chips as they await new US curbs, say sources. Reuters, August 2024
2024
-
[41]
Focus: Inside China’s underground market for high-end Nvidia AI chips
Josh Ye, David Kirton, Chen Lin, and Chen Lin. Focus: Inside China’s underground market for high-end Nvidia AI chips. Reuters, June 2023
2023
-
[42]
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Yanli Zhao, Andrew Gu, Rohan Varma, Liang Luo, Chien-Chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, Alban Desmaison, Can Balioglu, Pritam Damania, Bernard Nguyen, Geeta Chauhan, Yuchen Hao, Ajit Mathews, and Shen Li. PyTorch FSDP: Experiences on Sc...
2023
-
[43]
NVIDIA’s H20 AI Accelerators Might Face The Next "US Ban", Team Green Stops Taking New Orders In China
Muhammad Zuhair. NVIDIA’s H20 AI Accelerators Might Face The Next "US Ban", Team Green Stops Taking New Orders In China. https://wccftech.com/nvidia-h20-ai-accelerators-next-us- ban-rumor-stops-taking-new-orders-china/, September 2024
2024
-
[2020]
Association for Computing Machinery
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.