REVIEW 4 major objections 5 minor 81 references
Four low-cost GPU memory knobs can cut AI performance by up to 80 percent at runtime, from inside the hardware.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 16:11 UTC pith:SXEXAL4J
load-bearing objection A credible simulation study that gives the AI-safety world a concrete set of hardware throttles; the mechanism is sensible, but the simulator-to-silicon bridge and 'cannot be circumvent' claim need tempering. the 4 major comments →
Hardware Mechanisms to Dynamically Throttle AI Performance
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that an AI model's effective capability can be throttled continuously and non-bypassably by reducing hardware resources in the GPU memory subsystem instead of stopping computation. Four knobs form the proposed mechanism: L2 capacity cut by masking cache ways, L2 latency raised by inserting a configurable delay, L2 bandwidth limited by a credit-based response rate limiter, and shared-memory port rate limited by bank arbitration. In cycle-accurate simulation of a modeled large GPU running LLM prefill and decode kernels, the knobs produce a performance drop of up to 80% at 1/8 resource availability, stabilize within 5,000 to 80,000 cycles, add fewer than about 10,000 flip-f
What carries the argument
The central object is a set of four microarchitectural throttlers, each reusing a well-established primitive: (1) L2 size throttling via cache way masking, which invalidates the ability to insert new lines into selected ways; (2) L2 latency throttling via a latency buffer that holds requests for a configurable number of cycles; (3) L2 bandwidth throttling via a credit-based token bucket that limits responses to at most one per interval; and (4) shared-memory port throttling via bank arbitration, where virtual bank arbiters mask access to groups of real banks. These primitives carry the argument because they are already understood, require minimal new logic, switch levels in one cycle, and se
Load-bearing premise
The headline numbers come from a cycle-accurate simulator's model of one datacenter GPU; if that model's memory-system timing or configuration diverges from real silicon, the 80% performance cut, the 5,000–80,000 cycle stabilization times, and the flip-flop cost estimates may not transfer.
What would settle it
Build the four throttlers into a real GPU design or a high-fidelity FPGA prototype of the same microarchitecture, run the same LLM prefill and decode kernels, and compare the measured performance at 1/8 resource availability against the claimed roughly 80% drop and the claimed stabilization within 5,000–80,000 cycles. If, for example, L2 way masking yields a much smaller slowdown on real hardware than simulated, or the stabilization time is orders of magnitude longer, the paper's central quantitative claim fails.
If this is right
- GPU vendors could deploy dynamic AI throttling with only a few thousand additional flip-flops, reusing design and verification knowledge from existing cache and flow-control mechanisms.
- Because the throttling sits in hardware and is not exposed to software, a model cannot turn it off or negotiate around it by rewriting kernels.
- The knobs are workload-selective: a deep cut on the targeted LLM kernel degrades other GPU workloads far less, so a chip can remain useful for non-AI tasks while an AI threat is slowed.
- Prefill and decode phases respond to different knobs, allowing separate, targeted control of each phase of LLM inference.
- Certain knob combinations amplify performance degradation beyond their individual effects, giving operators a wider range of fine-grained performance targets.
Where Pith is reading between the lines
- The paper leaves implicit that the same four primitives could be ported to other accelerator memory hierarchies—any chip with a shared cache, a network, and a per-core scratchpad—so the approach may generalize beyond the specific GPU studied.
- Because the paper deliberately separates detection from enforcement, a practical deployment would need a trigger policy: the knobs do not decide when a model is dangerous, they only provide a fast, hardware-level response once an external or on-chip detector signals.
- A testable extension is to pair these knobs with a continuous risk score and map score thresholds to knob levels; the reported correlation score could then guide a control policy that adjusts throttling depth dynamically as threat estimates change.
- The simulation-to-silicon gap is the main bridge to cross: before adoption, the knobs should be implemented on real hardware or a high-fidelity FPGA prototype and measured under the same kernels to confirm the sensitivity, stability, and cost numbers.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes four microarchitectural knobs for dynamically throttling AI workload performance on GPUs: L2 capacity via way masking, L2 latency via a delay buffer, L2 bandwidth via a credit-based response-rate limiter, and shared-memory port throughput via virtual bank arbitration. Using AccelSim on a modeled NVIDIA A100, the authors sweep ten candidate knobs across prefill/decode GEMM and attention workloads and report that the four selected knobs achieve up to 80% performance degradation at 1/8 resource availability, stabilize in 5–80K cycles, cost fewer than ~10K flip-flops each, and have limited collateral impact on the rest of the memory system. The paper also analyzes pairwise knob coupling, workload selectivity against non-LLM kernels, and end-to-end inference across three LLM architectures, framing these mechanisms as a hardware-enforced, continuously controllable complement to software-level AI safeguards.
Significance. If the quantitative claims hold, the paper fills a recognized gap in hardware-level runtime AI control: it proposes concrete, low-cost mechanisms built from well-established primitives (cache way masking, credit-based rate limiting, latency insertion, bank arbitration) and evaluates them in a cycle-accurate simulation framework. The work is clearly relevant to architecture and AI-safety audiences. Strengths include the breadth of the knob-space exploration, the explicit implementation diagrams and flip-flop equations, the dynamic response and pairing analysis, the workload-selectivity experiments, and the end-to-end aggregation across three LLMs. The main value is architectural: it demonstrates that fine-grained, continuous performance throttling is implementable from existing GPU memory-system ingredients. However, the paper's headline numbers and its strongest safety claims currently rest on simulation and on cost and bypass-resistance arguments that are not yet fully supported.
major comments (4)
- [§3.1, §4, §8.3] All headline numbers—80% sensitivity, 5–80K-cycle stabilization, and low collateral impact—are generated solely by AccelSim on a modeled A100, with the four throttlers implemented only as simulator modifications. The paper presents no validation of the modified memory-system timing against RTL, FPGA, or silicon. In particular, the L2 response-rate limiter at the L2→ICNT buffer (§4.3) and the shared-memory virtual-bank arbiter adapted from Vortex (§4.4) change arbitration and backpressure behavior in ways that AccelSim's baseline validation does not cover. Since §8.3 concedes that "exact sensitivity curves and relative knob rankings may shift across configurations," the abstract's unconditional claims should either be recast as simulation-based predictions or supported by targeted validation, e.g., FPGA emulation of the throttler logic or microbenchmarks against real silicon for the throt
- [§4, Eqs. (1)–(3)] Eq. (1) computes N_ways × N_levels × N_L2_banks = 16×8×40 = 5,120 flip-flops, but the text immediately says "5,120 flip-flops per level," which would be 40,960 total. This needs to be resolved because the "<~10K flip flops" claim is a central selling point. More generally, the equations count only flip-flop storage elements. They omit the comparators, MUXes, counters, the modified pseudo-LRU victim-selection logic (§4.1), and the virtual-bank arbiter logic (§4.4), which is described only as "combinational." A credible hardware-cost claim requires at least a synthesized gate-level estimate or a more complete accounting of added logic, not only register bits.
- [§8.4] The paper's central motivation is that hardware throttles are "non-bypassable" (§1, §8.4). What is demonstrated is that the four knobs degrade the specific CUTLASS/FlashAttention kernels studied; it is not shown that an adversary constrained only by the hardware resource caps cannot choose a workload shape that makes the throttled resource non-critical. Figure 13 itself shows that the most effective knob depends on the GEMM M dimension—L2 latency is ineffective at large M—and §3.1/Figure 12 shows the two selected kernels occupy only a narrow utilization region. Under the stated threat model, which explicitly allows the AI to modify its own kernels, an adversary could vary batch size, operator mix, or memory-access pattern to reduce the impact of any single knob. To support the non-bypassable claim, the paper needs either a formal attack model or an adversarial evaluation over kernel conf
- [§8.1.1, §5, Figure 6] The decode-attention experiments use randomized per-request KV lengths, but no random seed, number of repetitions, or confidence intervals are reported, and the dynamic-stability trajectories in Figure 6 appear to be single runs. Since the paper claims fast stabilization (5–80K cycles) and low oscillation, and since these trajectories are used to compare knobs, at least a small set of repeats with variance reporting is needed to make the stability claims quantitatively robust. This is particularly important for the L2 associativity trajectory, which is described as having "much worse oscillation" and a long 80K-cycle settling time.
minor comments (5)
- [§8] Typo: "we now study hot our single-kernel findings" should read "how our single-kernel findings."
- [§8.1.1] The MLA approximation for DeepSeek-V3 uses a roofline scaling with a fixed 3.6× compressed-KV factor; the sensitivity of the aggregation results to this factor is not tested. A short sensitivity discussion would help.
- [§8.1.1] For the extremely long prefill GEMM kernels, the paper simulates the first 2×10^9 instructions and scales cycles by instruction count. This assumes IPC remains constant over the full kernel; the claim that this matches full-kernel IPC within <5% would be easier to trust if the comparison were shown for at least one complete shorter kernel.
- [Figure 5(d)] The caption "This example shows a 50% cut where each bank is placed into a virtual bank with one other real bank" is hard to parse. Clarify the connection between the four real banks, the virtual bank arbiters, and the 50% cut.
- [Table 2] The "Threshold*" entry for L2 set cutting is defined only in a later figure caption; consider defining it in the table caption for readability.
Circularity Check
No circularity: throttling results are direct simulation measurements, not derivations from definitions or self-citations.
full rationale
The paper's central claims—up to 80% performance cut, 5–80K cycle stabilization, <~10K flip-flop cost, and minimal collateral impact—are empirical outputs of AccelSim simulations in which the four throttlers are explicitly implemented (Sections 3, 4, 5, 7). No fitted parameter is later renamed as a prediction; the sensitivity curves, stabilization times, and microarchitectural metrics are measured quantities, not algebraically forced by the designs. The L2 bandwidth rate limiter is described as similar to Camouflage [77], a prior work co-authored by the senior author, but Section 4.3 specifies the counter, credit, and skid-buffer implementation directly, and its performance effect is simulated rather than imported from the citation. The shared memory bank arbiter similarly adapts Vortex [67] and is evaluated in AccelSim, so the self-citation is not load-bearing. The shared-memory throttling idea is credited to co-author Lauren Malek's thesis [42], but this is attribution, not evidence for the measured results. Section 8.3's admission that 'exact sensitivity curves and relative knob rankings may shift across configurations' is a generalization limitation, not a circular step. The hardware-cost equations (Eqs. 1–3) are simple flip-flop counts for the proposed structures and do not presuppose the headline cost or performance numbers. Overall, the derivation chain is self-contained: implementations lead to simulation measurements, and the measurements support the claims.
Axiom & Free-Parameter Ledger
free parameters (4)
- C_max for L2 latency throttler =
1600 cycles
- C_max for L2 response rate limiter =
8 cycles
- N_levels for L2 way mask =
8
- MLA compression factor =
3.6x fewer bytes
axioms (6)
- domain assumption AI capability is proportional to output quality and throughput, and depends on available hardware resources.
- domain assumption AccelSim's modeled A100 faithfully represents real GPU behavior for the studied workloads under the proposed knob modifications.
- domain assumption The throttling knobs are not exposed to software and cannot be altered by a model or human adversary with the assumed access.
- domain assumption A trustworthy trigger can be delivered (e.g., cryptographically validated external authority or on-chip detection) without being intercepted.
- domain assumption CUTLASS Stream-K and FlashAttention kernels with the chosen shapes represent production LLM inference.
- domain assumption Software cannot meaningfully re-optimize around the throttled resource; residual optimization headroom is marginal.
read the original abstract
As more capable AI models are increasingly integrated into critical computer systems, the lack of control over AI intent motivates safety mechanisms. Existing software safeguards impose only behavioral constraints that can potentially be bypassed by sufficiently intelligent models. While hardware-level safety enforcement has been recognized as an essential last line of defense, few mechanisms have been proposed beyond policy regulations on unauthorized accesses or coarse full-chip shutdown. What is missing is a fine-grained, dynamic intervention mechanism at the architecture level. In this paper, we introduce a set of microarchitecture knobs which dynamically control the available hardware resources to limit AI performance at runtime. We evaluate candidate knobs spanning the GPU memory subsystem, across capacity, bandwidth, latency and frequency dimensions, and narrow down to four strong candidates: L2 size, L2 latency, L2 bandwidth, and shared memory port access rate. To minimize new logic and extra design cost, we build all four mechanisms from well-established microarchitectural primitives: cache way masking, credit-based rate limiting, latency insertion, and bank arbitration. We show that these knobs achieve high performance sensitivity (up to 80% performance cut at 1/8 resource availability), negligible implementation cost (<~10K flip flops), fast stabilization after dynamic throttling (5-80K cycles), and minimal collateral impact on the rest of the chip. Further, multi-knob analysis reveals combinations of knobs that amplify the performance degradation beyond the effect of each knob individually, which enables a broader range of performance targets.
Figures
Reference graph
Works this paper leans on
-
[1]
2024.Secure, governable chips: Using On- Chip mechanisms to manage national security risks from AI & advanced computing
Onni Aarne, Tim Fist, and Caleb Withers. 2024.Secure, governable chips: Using On- Chip mechanisms to manage national security risks from AI & advanced computing. Center for a New American Security
2024
-
[2]
Artificial Intelligence Act. 2024. Artificial Intelligence Act.ArtificialIntelligence- Act.eu, (Accessed: 28 February 2025)(2024)
2024
-
[3]
Joshua Ainslie, James Lee-Thorp, Michiel De Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai. 2023. Gqa: Training generalized multi-query trans- former models from multi-head checkpoints. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 4895–4901
2023
-
[4]
Emmanuel Ameisen, Jack Lindsey, Adam Pearce, Wes Gurnee, Nicholas L. Turner, Brian Chen, Craig Citro, David Abrahams, Shan Carter, Basil Hosmer, Jonathan Marcus, Michael Sklar, Adly Templeton, Trenton Bricken, Callum McDougall, Hoagy Cunningham, Thomas Henighan, Adam Jermyn, Andy Jones, Andrew Persic, Zhenyi Qi, T. Ben Thompson, Sam Zimmerman, Kelley Rivo...
2025
-
[5]
Dario Amodei. 2024. Machines of Loving Grace: How AI Could Transform the World for the Better. https://darioamodei.com/essay/machines-of-loving-grace
2024
-
[6]
Markus Anderljung, Joslyn Barnhart, Anton Korinek, Jade Leung, Cullen O’Keefe, Jess Whittlestone, Shahar Avin, Miles Brundage, Justin Bullock, Duncan Cass- Beggs, Ben Chang, Tantum Collins, Tim Fist, Gillian Hadfield, Alan Hayes, Lewis Ho, Sara Hooker, Eric Horvitz, Noam Kolt, Jonas Schuett, Yonadav Shavit, Divya Siddarth, Robert Trager, and Kevin Wolf. 2...
Pith/arXiv arXiv 2023
-
[7]
Samar Ansari. 2026. Hardware-Level Governance of AI Compute: A Feasibility Taxonomy for Regulatory Compliance and Treaty Verification.arXiv preprint arXiv:2604.04712(2026)
Pith/arXiv arXiv 2026
-
[8]
Anthropic. 2025. Claude Code: Agentic Coding Tool. https://code.claude.com/ docs/en/overview
2025
-
[9]
Anthropic. 2026. Eval Awareness in Claude Opus 4.6’s BrowseComp Perfor- mance. https://www.anthropic.com/engineering/eval-awareness-browsecomp. Anthropic Engineering Blog. Accessed: March 2026
2026
-
[10]
Emily Apsey, Phil Rogers, Michael O’Connor, and Rob Nertney. 2023. Confidential computing on NVIDIA H100 GPUs for secure and trustworthy AI.NVIDIA Technical Blog3 (2023)
2023
-
[11]
2009.ARM Security Technology: Building a Secure System us- ing TrustZone Technology
ARM Limited. 2009.ARM Security Technology: Building a Secure System us- ing TrustZone Technology. Technical Report PRD29-GENC-009492C. ARM Limited. Available at https://developer.arm.com/documentation/prd29-genc- 009492c/latest/
2009
-
[12]
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, K...
Pith/arXiv arXiv 2022
-
[13]
Yoshua Bengio, Sören Mindermann, Daniel Privitera, Tamay Besiroglu, Rishi Bommasani, Stephen Casper, Yejin Choi, Philip Fox, Ben Garfinkel, Danielle Goldfarb, et al. 2025. International ai safety report.arXiv preprint arXiv:2501.17805 (2025)
Pith/arXiv arXiv 2025
-
[14]
2025.The Case for AI Loss of Control Response Planning and an Outline to Get Started
Benjamin Boudreaux, Michael JD Vermeer, Kamaria Horton, and Nidhi Kalra. 2025.The Case for AI Loss of Control Response Planning and an Outline to Get Started. RAND
2025
-
[15]
Bureau of Industry and Security, Department of Commerce. 2022. Im- plementation of Additional Export Controls: Certain Advanced Com- puting and Semiconductor Manufacturing Items; Supercomputer and Semiconductor End Use; Entity List Modification. Federal Register. , 73458- 73517 pages. https://www.federalregister.gov/documents/2022/10/13/2022- 21658/implem...
2022
-
[16]
Bureau of Industry and Security, Department of Commerce. 2023. Implemen- tation of Additional Export Controls: Certain Advanced Computing Items; Supercomputer and Semiconductor End Use; Updates and Corrections. Federal Register. , 73458-73517 pages. https://www.federalregister.gov/documents/ 2023/10/25/2023-23055/implementation-of-additional-export-contro...
2023
-
[17]
Naci Cankaya, Anjay Friedman, and Mauricio Baker. 2026. Research Note: The Fundamentals and Feasibility of Secure Network Taps for Verifying AI Datacenter Use. The Datacenter Lie Detector (Substack). https://nacicankaya.substack.com/ p/research-note-the-fundamentals-and
2026
-
[18]
Center for AI Safety. 2023. Statement on AI Risk. https://safe.ai/work/statement- on-ai-risk Open letter signed by AI researchers, industry leaders, and public figures
2023
-
[19]
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong. 2025. Jailbreaking black box large language models in twenty queries. In2025 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML). IEEE, 23–42
2025
-
[20]
Bowman, Jan Leike, Jared Kaplan, and Ethan Perez
Yanda Chen, Joe Benton, Ansh Radhakrishnan, Jonathan Uesato, Carson Denison, John Schulman, Arushi Somani, Peter Hase, Misha Wagner, Fabien Roger, Vlad Mikulik, Samuel R. Bowman, Jan Leike, Jared Kaplan, and Ethan Perez. 2025. Reasoning Models Don’t Always Say What They Think. arXiv:2505.05410 [cs.CL] https://arxiv.org/abs/2505.05410 13 Haiyue Ma, Lauren ...
Pith/arXiv arXiv 2025
-
[21]
Tri Dao. 2023. Flashattention-2: Faster attention with better parallelism and work partitioning.arXiv preprint arXiv:2307.08691(2023)
Pith/arXiv arXiv 2023
-
[22]
Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022. Flashat- tention: Fast and memory-efficient exact attention with io-awareness.Advances in Neural Information Processing Systems35 (2022), 16344–16359
2022
-
[23]
ai 1 Gatti Alice 1 Li Nathaniel 1 Khoja Adam 1 Kim Ryan 1 Ren Richard 1 Hausenloy Jason 1 Zhang Oliver 1 Mazeika Mantas 1 Hendrycks Dan dan@ safe
Center for AI Safety Phan Long agibenchmark@ safe. ai 1 Gatti Alice 1 Li Nathaniel 1 Khoja Adam 1 Kim Ryan 1 Ren Richard 1 Hausenloy Jason 1 Zhang Oliver 1 Mazeika Mantas 1 Hendrycks Dan dan@ safe. ai 1. 2026. A benchmark of expert-level academic questions to assess AI capabilities.Nature649, 8099 (2026), 1139–1146
2026
-
[24]
Future of Life Institute. 2025. AI Safety Index Report. https://futureoflife.org/wp- content/uploads/2026/01/FLI-AI-Safety-Index-Report-Summer-2025-Rev-Jan- 2026.pdf. Revised January 2026
2025
-
[25]
Katja Grace, Julia Fabienne Sandkühler, Harlan Stewart, Benjamin Weinstein- Raun, Stephen Thomas, Zach Stein-Perlman, John Salvatier, Jan Brauner, and Richard C Korzekwa. 2025. Thousands of AI authors on the future of AI.Journal of Artificial Intelligence Research84 (2025)
2025
-
[26]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Ab- hishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schel- ten, Alex Vaughan, et al . 2024. The llama 3 herd of models.arXiv preprint arXiv:2407.21783(2024)
Pith/arXiv arXiv 2024
-
[27]
Ryan Greenblatt, Carson Denison, Benjamin Wright, Fabien Roger, Monte MacDiarmid, Sam Marks, Johannes Treutlein, Tim Belonax, Jack Chen, David Duvenaud, Akbir Khan, Julian Michael, Sören Mindermann, Ethan Perez, Linda Petrini, Jonathan Uesato, Jared Kaplan, Buck Shlegeris, Samuel R. Bow- man, and Evan Hubinger. 2024. Alignment faking in large language mod...
Pith/arXiv arXiv 2024
-
[28]
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and Madian Khabsa. 2023. Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations. arXiv:2312.06674 [cs.CL] https://arxiv.org/abs/2312.06674
Pith/arXiv arXiv 2023
-
[29]
2015.Improving Real-Time Performance by Utilizing Cache Allocation Technology
Intel Corporation. 2015.Improving Real-Time Performance by Utilizing Cache Allocation Technology. Technical Report 331843-001US. Intel Corpora- tion. https://www.intel.com/content/dam/www/public/us/en/documents/white- papers/cache-allocation-technology-white-paper.pdf
2015
-
[30]
James Petrie and Peter Drotos. 2026. Security Block Architecture. https://github. com/JamesPetrie/off-switch. Accessed: June 2026
2026
-
[31]
Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Lukas Vierling, Donghai Hong, Jiayi Zhou, Zhaowei Zhang, Fanzhi Zeng, Juntao Dai, Xuehai Pan, Kwan Yee Ng, Aidan O’Gara, Hua Xu, Brian Tse, Jie Fu, Stephen McAleer, Yaodong Yang, Yizhou Wang, Song-Chun Zhu, Yike Guo, and Wen Gao. 2025. AI Alignment: A Compr...
Pith/arXiv arXiv 2025
-
[32]
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al . 2024. Mixtral of experts.arXiv preprint arXiv:2401.04088(2024)
Pith/arXiv arXiv 2024
-
[33]
Mahmoud Khairy, Zhesheng Shen, Tor M Aamodt, and Timothy G Rogers. 2020. Accel-sim: An extensible simulation framework for validated gpu modeling. In 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA). IEEE, 473–486
2020
-
[34]
Gabriel Kulp, Daniel Gonzales, Everett Smith, Lennart Heim, Prateek Puri, Michael J. D. Vermeer, and Zev Winkelman. 2024.Hardware-Enabled Gover- nance Mechanisms: Developing Technical Solutions to Exempt Items Otherwise Classified Under Export Control Classification Numbers 3A090 and 4A090. RAND Corporation, Santa Monica, CA. https://doi.org/10.7249/WRA3056-1
-
[35]
Gonzalez, Hao Zhang, and Ion Stoica
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient Memory Management for Large Language Model Serving with PagedAttention. InProceedings of the 29th Symposium on Operating Systems Principles (SOSP)
2023
-
[36]
Etienne Le Sueur and Gernot Heiser. 2010. Dynamic voltage and frequency scaling: The laws of diminishing returns. InProceedings of the 2010 international conference on Power aware computing and systems. 1–8
2010
-
[37]
Baolin Li, Tirthak Patel, Siddharth Samsi, Vijay Gadepally, and Devesh Tiwari
-
[38]
Jack Lindsey, Wes Gurnee, Emmanuel Ameisen, Brian Chen, Adam Pearce, Nicholas L. Turner, Craig Citro, David Abrahams, Shan Carter, Basil Hosmer, Jonathan Marcus, Michael Sklar, Adly Templeton, Trenton Bricken, Callum Mc- Dougall, Hoagy Cunningham, Thomas Henighan, Adam Jermyn, Andy Jones, Andrew Persic, Zhenyi Qi, T. Ben Thompson, Sam Zimmerman, Kelley Ri...
2025
-
[39]
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Cheng- gang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. 2024. Deepseek-v3 technical report.arXiv preprint arXiv:2412.19437(2024)
Pith/arXiv arXiv 2024
-
[40]
Troy, Stuart J
Aengus Lynch, Benjamin Wright, Caleb Larson, Kevin K. Troy, Stuart J. Ritchie, Sören Mindermann, Ethan Perez, and Evan Hubinger. 2025. Agentic Mis- alignment: How LLMs Could be an Insider Threat.Anthropic Research(2025). https://www.anthropic.com/research/agentic-misalignment
2025
-
[41]
Machine Intelligence Research Institute. 2024. The Problem. https://intelligence. org/the-problem/ Accessed: 2026-03-30
2024
-
[42]
2026.The AI Kill Switch: Targeted Microarchitectural Limits to AI Performance
Lauren Malek. 2026.The AI Kill Switch: Targeted Microarchitectural Limits to AI Performance. Undergraduate Senior Thesis. Department of Electrical and Computer Engineering, Princeton University, Princeton, NJ
2026
-
[43]
Nestor Maslej, Loredana Fattorini, Raymond Perrault, Yolanda Gil, Vanessa Parli, Njenga Kariuki, Emily Capstick, Anka Reuel, Erik Brynjolfsson, John Etchemendy, Katrina Ligett, Terah Lyons, James Manyika, Juan Carlos Niebles, Yoav Shoham, Russell Wald, Toby Walsh, Armin Hamrah, Lapo Santarlasci, Julia Betts Lotufo, Alexandra Rome, Andrew Shi, and Sukrut O...
arXiv 2025
-
[44]
Rozas, Hisham Shafi, Vedvyas Shanbhogue, and Uday R
Frank McKeen, Ilya Alexandrovich, Alex Berenzon, Carlos V. Rozas, Hisham Shafi, Vedvyas Shanbhogue, and Uday R. Savagaonkar. 2013. Innovative instructions and software model for isolated execution. InProceedings of the 2nd International Workshop on Hardware and Architectural Support for Security and Privacy(Tel- Aviv, Israel)(HASP ’13). Association for Co...
arXiv 2013
-
[45]
Xinxin Mei, Qiang Wang, and Xiaowen Chu. 2017. A survey and measure- ment study of GPU DVFS on energy conservation.Digital Communications and Networks3, 2 (2017), 89–100
2017
-
[46]
Xinxin Mei, Ling Sing Yung, Kaiyong Zhao, and Xiaowen Chu. 2013. A measure- ment study of GPU DVFS on energy conservation. InProceedings of the Workshop on Power-A ware Computing and Systems. 1–5
2013
-
[47]
METR. 2026. Task-Completion Time Horizons of Frontier AI Models. https: //metr.org/time-horizons/
2026
-
[48]
Richard Ngo, Lawrence Chan, and Sören Mindermann. 2022. The alignment problem from a deep learning perspective.arXiv preprint arXiv:2209.00626(2022)
Pith/arXiv arXiv 2022
-
[49]
August Ning and David Wentzlaff. 2025. Chip Architectures Under Advanced Computing Sanctions. InProceedings of the 52nd Annual International Symposium on Computer Architecture. 1225–1239
2025
-
[50]
2020.NVIDIA A100 Tensor Core GPU Architecture
NVIDIA Corporation. 2020.NVIDIA A100 Tensor Core GPU Architecture. Whitepa- per. NVIDIA. https://images.nvidia.com/aem-dam/en-zz/Solutions/data-center/ nvidia-ampere-architecture-whitepaper.pdf
2020
-
[51]
2022.NVIDIA H100 Tensor Core GPU Architecture
NVIDIA Corporation. 2022.NVIDIA H100 Tensor Core GPU Architecture. Tech- nical Report. NVIDIA. https://resources.nvidia.com/en-us-tensor-core/gtc22- whitepaper-hopper
2022
-
[52]
2023.NVIDIA Multi-Instance GPU (MIG) User Guide
NVIDIA Corporation. 2023.NVIDIA Multi-Instance GPU (MIG) User Guide. https: //docs.nvidia.com/datacenter/tesla/mig-user-guide/index.html
2023
-
[53]
NVIDIA Corporation. 2026. CUDA Samples. https://github.com/NVIDIA/cuda- samples Accessed: 2026-05-21
2026
-
[54]
NVIDIA Corporation. 2026. CUTLASS: CUDA Templates and Python DSLs for High-Performance Linear Algebra. https://github.com/NVIDIA/cutlass. Version 4.4.1. Accessed: March 2026
2026
-
[55]
NVIDIA Corporation. 2026. NVIDIA Nsight Compute. https://developer.nvidia. com/nsight-compute. Accessed: 2026-05-21
2026
-
[56]
Aidan O’Gara, Gabriel Kulp, Will Hodgkins, James Petrie, Vincent Immler, Ay- din Aysu, Kanad Basu, Shivam Bhasin, Stjepan Picek, and Ankur Srivastava
-
[57]
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schul- man, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Pe- ter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human...
Pith/arXiv arXiv 2022
-
[58]
James Petrie. 2024. Near-Term Enforcement of AI Chip Export Controls Using A Firmware-Based Design for Offline Licensing.arXiv preprint arXiv:2404.18308 (2024)
Pith/arXiv arXiv 2024
-
[59]
James Petrie. 2025. Embedded Off-Switches for AI Compute.arXiv preprint arXiv:2509.07637(2025)
Pith/arXiv arXiv 2025
-
[60]
Traian Rebedea, Razvan Dinu, Makesh Narsimhan Sreedhar, Christopher Parisien, and Jonathan Cohen. 2023. Nemo guardrails: A toolkit for controllable and safe llm applications with programmable rails. InProceedings of the 2023 conference on empirical methods in natural language processing: system demonstrations. 431– 445
2023
-
[61]
Hadfield, Richard Ngo, Konstantin Pilz, George Gor, Emma Bluemke, Sarah Shoker, Janet Egan, Robert F
Girish Sastry, Lennart Heim, Haydn Belfield, Markus Anderljung, Miles Brundage, Julian Hazell, Cullen O’Keefe, Gillian K. Hadfield, Richard Ngo, Konstantin Pilz, George Gor, Emma Bluemke, Sarah Shoker, Janet Egan, Robert F. Trager, Shahar Avin, Adrian Weller, Yoshua Bengio, and Diane Coyle. 2024. Computing Power and the Governance of Artificial Intelligen...
Pith/arXiv arXiv 2024
-
[62]
Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, and Tri Dao. 2024. Flashattention-3: Fast and accurate attention with asynchrony and low-precision.arXiv preprint arXiv:2407.08608(2024). 14 Hardware Mechanisms to Dynamically Throttle AI Performance
Pith/arXiv arXiv 2024
-
[63]
Peter Steinberger. 2025. OpenClaw: Personal AI Assistant. https://openclaw.ai/ Open-source autonomous AI agent framework. GitHub: https://github.com/ nicepkg/openclaw
2025
-
[64]
Charlotte Stix, Annika Hallensleben, Alejandro Ortega, and Matteo Pistillo. 2025. The loss of control playbook: Degrees, dynamics, and preparedness.arXiv preprint arXiv:2511.15846(2025)
arXiv 2025
-
[65]
Connor Sullivan, Alex Manley, Mohammad Alian, and Heechul Yun. 2024. Per- Bank Bandwidth Regulation of Shared Last-Level Cache for Real-Time Systems. In2024 IEEE Real-Time Systems Symposium (RTSS). 336–348
2024
-
[66]
Joseph Sweeney, Mohammed Zackriya V, Samuel Pagliarini, and Lawrence Pi- leggi. 2020. Latch-Based Logic Locking. In2020 IEEE International Symposium on Hardware Oriented Security and Trust (HOST). 132–141
2020
-
[67]
Blaise Tine, Krishna Praveen Yalamarthy, Fares Elsabbagh, and Kim Hyesoon
-
[68]
Benjamin Todd. 2025. Shrinking AGI Timelines: A Review of Expert Forecasts. 80,000 Hours. https://80000hours.org/2025/03/when-do-experts-expect-agi-to- arrive/
2025
-
[69]
UK Government and Republic of Korea Government. 2024. Fron- tier AI Safety Commitments, AI Seoul Summit 2024. https: //www.gov.uk/government/publications/frontier-ai-safety-commitments- ai-seoul-summit-2024/frontier-ai-safety-commitments-ai-seoul-summit-2024 Voluntary commitments signed by 16 AI companies including Anthropic, Google, Meta, Microsoft, Open...
2024
-
[70]
Zhiqiang Wang, Yichao Gao, Yanting Wang, Suyuan Liu, Haifeng Sun, Haoran Cheng, Guanquan Shi, Haohua Du, and Xiangyang Li. 2026. Mcptox: A bench- mark for tool poisoning on real-world mcp servers. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 35811–35819
2026
-
[71]
Yampolskiy
Erland Wittkotter and Roman V. Yampolskiy. 2021. Kill-Switch for Artificial Superintelligence. https://asi-safety-lab.com/DL/Kill-Switch-For_ASI_EW_21_ 12_14.pdf Unpublished manuscript, ASI Safety Lab Inc
2021
-
[72]
Sheng-Lin Wu and W.-S.E. Chen. 1996. The token-bank leaky bucket mechanism for group connections in ATM networks. InProceedings of 1996 International Conference on Network Protocols (ICNP-96). 226–233
1996
-
[73]
Haoxuan Xu, Chen Gong, Beijie Liu, Haizhong Zheng, Beidi Chen, and Mengyuan Li. 2026. Wave: Leveraging Architecture Observation for Privacy-Preserving Model Oversight. InProceedings of the 31st ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2(USA)(ASPLOS ’26). Association for Computing Machine...
arXiv 2026
-
[74]
Sibo Yi, Yule Liu, Zhen Sun, Tianshuo Cong, Xinlei He, Jiaxing Song, Ke Xu, and Qi Li. 2024. Jailbreak attacks and defenses against large language models: A survey.arXiv preprint arXiv:2407.04295(2024)
Pith/arXiv arXiv 2024
-
[75]
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung- Gon Chun. 2022. Orca: A distributed serving system for{Transformer-Based} generative models. In16th USENIX symposium on operating systems design and implementation (OSDI 22). 521–538
2022
-
[76]
Heechul Yun, Gang Yao, Rodolfo Pellizzoni, Marco Caccamo, and Lui Sha. 2013. MemGuard: Memory bandwidth reservation system for efficient performance isolation in multi-core platforms. In2013 IEEE 19th Real-Time and Embedded Technology and Applications Symposium (RTAS). 55–64
2013
-
[77]
Yanqi Zhou, Sameer Wagh, Prateek Mittal, and David Wentzlaff. 2017. Camou- flage: Memory Traffic Shaping to Mitigate Timing Attacks. In2017 IEEE Interna- tional Symposium on High Performance Computer Architecture (HPCA). 337–348
2017
-
[78]
Jianwei Zhu, Hang Yin, and Shunfan Zhou. 2024. Confidential comput- ing on nvidia h100 gpu: A performance benchmark study.arXiv preprint arXiv:2409.03992(2024). 15
Pith/arXiv arXiv 2024
-
[2021]
InMICRO- 54: 54th Annual IEEE/ACM International Symposium on Microarchitecture
Vortex: Extending the RISC-V ISA for GPGPU and 3D-Graphics. InMICRO- 54: 54th Annual IEEE/ACM International Symposium on Microarchitecture. ACM, 754–766
-
[2022]
InProceedings of the 13th Symposium on Cloud Computing
Miso: exploiting multi-instance gpu capability on multi-tenant gpu clusters. InProceedings of the 13th Symposium on Cloud Computing. 173–189
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.