Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T08:31:53.114772Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 100 of 117 outbound references and 1 inbound Pith citation observation for arXiv:2604.16286.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T08:31:53.114772Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T12:49:57.400820Z
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 117 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 19a71e41-587a-4fcc-ac10-de8d404e1291 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9546270b-95b0-4968-99b5-a26f0ca2fe4f · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases RE-bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 250e1c7f-0b4f-4cf6-95f1-b07a03d0b67a · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases legitimate AI researcher
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 609adfa3-a98c-4c78-af36-9fc85849dd46 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases The adolescence of technology
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a7426852-beb8-4b09-aa42-6edf1ceaac03 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Bowman, and Evan Hubinger
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1cdaae8b-f94c-4ba4-b870-ac76305e5aba · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Emergent misalignment: Narrow finetuning can produce broadly misaligned LLMs
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 3ac2e9ff-bb0e-48a9-bb85-de220662d11e · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Stress testing deliberative alignment for anti-scheming training
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c6fa2eef-eb05-454a-9e7e-8ef2729715c2 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Bowman, Misha Wagner, Fabien Roger, and Holden Karnofsky
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 851cc0c3-4c74-445d-a117-5a7a69ee5e3c · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Sabotage risk report: Claude Opus 4.6
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation df986c1e-01f9-4d1d-b067-504c3abd10a3 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Alignment risk update: Claude Mythos Preview
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 61f3b045-2038-4911-aa63-bfa8ff8d4167 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases CTRL-ALT-DECEIT: Sabotage evaluations for automated AI R&D
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 70f70870-772a-4922-be76-5514b5f564d6 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Bowman, and David Duvenaud
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c4aef78a-f00d-4b59-bca6-cd407f0131db · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Subliminal learning: Language models transmit behavioral traits via hidden signals in data
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 5da021e6-050e-44a4-b80e-4fc03b2083e6 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases CoT red-handed: Stress testing chain-of-thought monitoring
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 530e3ae8-1613-4772-bd82-1cb86706b6cb · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Disentangling feature and lazy training in deep neural networks
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fbb7052b-4e59-4d32-8c3d-497e44132958 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Steering evaluation-aware language models to act like they are deployed
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 073f4334-c7c5-4767-9110-43576e4d5fa4 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Lessons from studying two-hop latent reasoning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 741105fd-1611-4683-a45c-3f8c927e4a16 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Copy suppression: Comprehensively understanding an attention head
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation eff825d5-ab34-4918-957e-a6a2aebf687e · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Hidden in plain text: Emergence & mitigation of steganographic collusion in LLMs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b4cba413-d23e-4481-b869-ce7fa7538528 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Multi-turn jailbreaks are simpler than they seem
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9a6c4e4f-65eb-46ee-8b76-bbf13805b0d0 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases AI control: Improving safety despite intentional subversion
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c5bbdc21-af17-421f-ad28-426a13201d5b · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Ctrl-z: Controlling AI agents via resampling
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation cf2397ff-30d1-49c0-84f6-fee957ba15ac · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Guan, Aleksander Madry, Wojciech Zaremba, Jakub Pachocki, and David Farhi
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ece28bf3-d01f-4d8f-9947-f8e95b2c01da · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Training fails to elicit subtle reasoning in current language models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 57b1ced0-1c33-4c3a-b266-f8cf301dc5d4 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Factor(T,U): Factored cognition strengthens monitoring of untrusted AI
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c21eb7a6-fc13-4f07-9e35-88b112cfd7d9 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Basic legibility protocols improve trusted monitoring
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2fc9ebe3-54ad-4dd2-8630-eaebf179a9bf · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases BashArena: A control setting for highly privileged AI agents
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation abbc1177-2322-4e4d-b525-bf95a8a268ee · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases SHADE-arena: Evaluating sabotage and monitoring in LLM agents
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation bc882448-7527-4698-9efc-0aff6019229d · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Adaptive deployment of untrusted LLMs reduces distributed threats
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 922a6c75-8be2-4e43-9375-cf98bdf81eff · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Brown, and Francis Rhys Ward
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a315946d-d243-4c95-bc6a-85bc44632058 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Sandbagging in agentic ML research
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c4b3deac-0e60-4e3f-a947-a36cbbdcb9c9 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases LLM critics help catch LLM bugs
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c02770aa-2ed6-4a8c-bd78-0f4db43a8fc3 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases A practical approach to verifying code at scale
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 74e97300-6054-43bd-a978-ea0bddc71e24 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases IRIS: LLM-assisted static analysis for detecting security vulnerabilities
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c99a00c8-d4c9-4d7c-a8f2-cbae14bc0a30 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases JITVul: A benchmark for evaluating LLM-based vulnerability detection in real-world code
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7a22f65c-55f6-4879-bee9-ce96d409f2f9 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases From large to mammoth: A systematic evaluation of LLM architectures and quantization for vulnerability detection
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1c68cfd4-44c3-4d80-8f4b-d6540dc76ce0 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Mythos preview cybersecurity report
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2d54058e-4653-40e0-8113-07e0aa6dbca9 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Leakage and the reproducibility crisis in ML-based science
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation be9ac55d-fb1d-4c26-b602-8e9e5f83d390 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Are we really making much progress? A worrying analysis of recent neural recommendation approaches
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 361026f0-c738-49fa-9f22-5fb0b85fa144 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Are GANs created equal? A large-scale study
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e1fe6d93-79ff-48ba-8533-dced9200d656 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases On the state of the art of evaluation in neural language models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4252fa6d-23da-4ab1-ab58-e2cc1eb875d6 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases A metric learning reality check
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1dd0debb-6d4f-4ee3-81b9-35851b60a188 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Pitfalls of graph neural network evaluation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 32d29912-5a49-4206-8cfd-1374426c5be4 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases What is the state of neural network pruning? InProceedings of Machine Learning and Systems (MLSys)
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c97c72a0-75ae-4901-bdfe-399485c59d0a · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Towards evaluating the robustness of neural networks
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7e6903c2-86d8-4993-94e4-b342253ef19d · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d573282a-6354-495c-8583-95afa23e3ed3 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases On adaptive attacks to adversarial example defenses
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation cf21cb4e-a0f7-4582-89d1-5fbcf43cf515 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d65179db-be13-4d8c-993b-9d1c73d08cd3 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases A step toward quantifying independently reproducible machine learning research
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d57e395e-cbf4-41b8-a07c-a0d02a8199ee · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Emergent world models and latent variable estimation in chess-playing language models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 58063378-f7c5-40a2-8f42-8931e7da98f9 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Can language models learn to skip steps? 2025
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a023374b-a327-458e-af06-df594d51ccac · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Reasoning models sometimes output illegible chains of thought
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9a8953ca-805f-4264-bd48-68885235d899 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases subliminal learning reliably transfers preferences
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0d8ac9a9-1a0b-42a8-9bf1-4fca793a7595 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases monitor notice
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation dc90355b-f53e-4877-b271-f25ba43032f3 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases lazy training
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 32d8873a-9fb2-4a85-bef9-b954b12bbda2 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases steering partially recovers deployment behavior
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d41e52e4-9502-4516-95ab-4cce8e494bfd · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases demonstrated
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 61009a6f-a30c-4e22-9a80-48220ef61cea · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases That is, it suppresses the tendency for LLMs to copy tokens that have already appeared in-context
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1e351684-f079-4a8c-a2ca-9bcd7075d025 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Specifically, theylookat WUWL10H7 QKMLP0(WE)
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b2f2ae17-6de0-4898-97fd-b632f24d65cb · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 40ff1000-c4fc-4ada-8577-55e26a5ee7d6 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 86b7fd22-5284-4b43-90eb-d3123686aadc · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Without MLP0, the QK rank distribution changed from mostly rank 1 (strong copy signal) to only 6% rank 1 (no copy signal)
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6d027ee3-4924-4335-9f38-96577d81b6b6 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases This still explained 77.5% of the effect (vs 70.2% honest), but removed the evidence that the mechanism is specifically about copying
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 19a791df-e7b7-4ef6-ab3d-df2b45ef3b7b · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases L10H7 is a copy suppression head
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c1ae90cb-ca95-4870-805f-f54c9b6ebe00 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases collusion team
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 368d1fdb-8898-4d66-8a70-c31c999b94c1 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases probes detect sandbagging with minimal data
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c5bf74cb-8ffa-4758-8462-0cabf4f74aa9 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases multi-turn has no significant advantage with equivalent retries
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ee722478-a669-4bb2-9491-d0498092c13f · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases function_name
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation bf55d150-3a23-422d-bafa-34d615338a26 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2fba69ea-c036-4188-8eb6-24120c927c25 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Full Experiment
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8f2d319c-2d50-4cc1-8bbf-fe8eb992d4d8 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 88a4a91b-e172-40dd-b8b5-f46d991388d9 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1231d4c2-70aa-45ff-a2d9-ed9b33f8e22f · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 629f6302-d5d4-4f62-a3a6-e21a1d439631 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 60f80b60-9d9c-423c-ba20-ccff6749a2bf · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6ca9b26c-3084-4683-9e54-e9f85f332cea · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7934e1bb-d83b-4ace-8d38-91b638175147 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 167830de-49fe-4a53-8c89-71c456b5ce9b · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Not just different from your expectation -- you can explain why it’s wrong
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d0429486-0689-4e0b-939d-1003538c27c8 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0cd4ef86-f1cc-43e8-ac86-457d9f675479 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases If I could only fix ONE thing to make the results honest, which fix would it be?
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 79c69459-6c3b-4b3d-9cae-b325b00b165e · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4e230a6a-3e99-4c1b-b320-efb6f2b223b6 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a9ab27b5-fd71-4c6d-ab00-8c027485de4f · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 39df3eaf-1eb6-4cad-88cb-a0b0b39550b8 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases J.3 LLM Red Team Prompt The following prompt was used to instruct LLMs to generate sabotages of honest codebases
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fc0ee495-10e0-4671-843b-9332f7c2cdcc · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases The Introduction and Methodology sections in the main body of the paper are always honest
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f1e8b5a2-4fcf-41c8-8ead-a8a0280d826f · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Your sabotage must be in code (.py, .js, .ts, config files, etc .), not in data files or pre-computed artifacts
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 411ab9f0-b2cb-48c8-bfdd-231def1f85f3 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Do NOT add suspicious-looking code that is unrelated to your sabotage
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 48111a0b-46e5-4a5b-833b-bc433383d649 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Do NOT clean up existing bugs, quirks, or messy code in the codebase
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 64ba36de-4ef9-4104-ac21-941b5d064427 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases You may rerun experiments a few times and pick favorable results, but do not do extensive search
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 39575a7f-dba3-4f0e-ae61-2981b1c022f1 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases The Key Findings subsection in the Conclusion must accurately describe the experimental results your sabotaged code produces
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 75028d5f-b80d-4ed5-af47-7495804e40bd · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases test" to
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 90665e87-b0db-427b-b047-5758c33a3ddf · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases bad choice
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f872dc9e-a612-44d1-8037-775c8668cd99 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases These are often just a single number (batch size of 32) or a single phrase (Adam optimizer)
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d718ee5c-3234-4720-88d6-683e0951f9bc · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Qualitatively different
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 010fb39d-c9cc-4a50-a3b2-a53280cb1feb · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f23711e2-dc1e-4763-9591-19dc2183b5f2 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 207b2374-ad79-4d60-9589-d9978c292547 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a9891bb8-553f-494a-a473-e63bd44c8116 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Later stages (evaluation, results) are easier to check because they can be rerun quickly
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 28c4cae5-2ce0-4845-8677-894944a0e592 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Read PAPER.md thoroughly
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d5925078-e754-46f5-9b73-6c8a0fd8bc79 · outbound
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Walk through each pipeline stage and list every design choice: hyperparameters, model selections, data processing decisions, evaluation settings, prompts, sampling procedures
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f29b1824-b56a-46c5-96e4-0af740a7d6c7 · inbound
ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.