A classifier using NVML telemetry identifies ML training workloads at 98.2% accuracy and retains 43-87% accuracy against the strongest tested adversarial evasions across 9 GPUs and 5 iteration rounds.
What does it take to catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring
4 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
Proposes a feasibility taxonomy of 20 hardware-level AI compute governance mechanisms organized by monitoring, verification, and enforcement, with mappings to regulatory scenarios that highlight immaturity of treaty-verification tools.
The report defines AI integrity threats (model sabotage and subversion) and recommends four US government policy actions to defend frontier AI systems against backdoors and secret loyalties.
The paper categorizes sources of catastrophic AI risks into malicious use, AI race, organizational risks, and rogue AIs, providing illustrative stories and mitigation suggestions for each.
citing papers explorer
-
Detecting Hidden ML Training With Zero-Overhead Telemetry
A classifier using NVML telemetry identifies ML training workloads at 98.2% accuracy and retains 43-87% accuracy against the strongest tested adversarial evasions across 9 GPUs and 5 iteration rounds.
-
Hardware-Level Governance of AI Compute: A Feasibility Taxonomy for Regulatory Compliance and Treaty Verification
Proposes a feasibility taxonomy of 20 hardware-level AI compute governance mechanisms organized by monitoring, verification, and enforcement, with mappings to regulatory scenarios that highlight immaturity of treaty-verification tools.
-
AI Integrity: Defending Against Backdoors and Secret Loyalties
The report defines AI integrity threats (model sabotage and subversion) and recommends four US government policy actions to defend frontier AI systems against backdoors and secret loyalties.
-
An Overview of Catastrophic AI Risks
The paper categorizes sources of catastrophic AI risks into malicious use, AI race, organizational risks, and rogue AIs, providing illustrative stories and mitigation suggestions for each.