Pith. sign in

Paper Citation Record · LEDGER

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases

As of 4 August 2026, this Paper Citation Record lists 100 of 117 outbound references and 1 inbound Pith citation observation for arXiv:2604.16286.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.16286 v2

Coverage vector

measured 100 of 117 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T08:31:53.114772Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T12:49:57.400820Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 117 outbound references displayed

  • verified exact0
  • verified fuzzy82
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19a71e41-587a-4fcc-ac10-de8d404e1291 · outbound

This paper cites Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.170855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:e6b7df01621843ebcf509b971f95fda73c19b0e501c9646901c1cbafaffd624c

Observation 9546270b-95b0-4968-99b5-a26f0ca2fe4f · outbound

This paper cites RE-bench: Evaluating frontier AI R&D capabilities of language model agents against human experts.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases RE-bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.992059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:8601d6b1a74a15454d3ab14eee541615eecb5f7fd83ca5136080eda1f996903f

Observation 250e1c7f-0b4f-4cf6-95f1-b07a03d0b67a · outbound

This paper cites legitimate AI researcher.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases legitimate AI researcher

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.079244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:9a95944576acd3f470d45283dd65a04e0b4cf7854b897ef409cb5c0840d3d393

Observation 609adfa3-a98c-4c78-af36-9fc85849dd46 · outbound

This paper cites The adolescence of technology.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases The adolescence of technology

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.051102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:2954068029b74b06ae787fa73b743bc2b977882b7d974094b9010895951215ca

Observation a7426852-beb8-4b09-aa42-6edf1ceaac03 · outbound

This paper cites Bowman, and Evan Hubinger.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Bowman, and Evan Hubinger

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.110431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:d4198b2df845e523194a5010153367c9e3630b56e4e1d44901d8f33c8e1a4dba

Observation 1cdaae8b-f94c-4ba4-b870-ac76305e5aba · outbound

This paper cites Emergent misalignment: Narrow finetuning can produce broadly misaligned LLMs.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Emergent misalignment: Narrow finetuning can produce broadly misaligned LLMs

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.271230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:033715123389b934ae41dadef776f1a7f5136391ba3d2f590482062134c7d236

Observation 3ac2e9ff-bb0e-48a9-bb85-de220662d11e · outbound

This paper cites Stress testing deliberative alignment for anti-scheming training.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Stress testing deliberative alignment for anti-scheming training

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.044333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:34507102cdcb06304164b20f9ff69ba7660ec3e4adabf31dec75e68b4ffef42d

Observation c6fa2eef-eb05-454a-9e7e-8ef2729715c2 · outbound

This paper cites Bowman, Misha Wagner, Fabien Roger, and Holden Karnofsky.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Bowman, Misha Wagner, Fabien Roger, and Holden Karnofsky

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.233981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:77d9ccf714ec75e31c1230d2b8bd5f18c7073b25f90b779f6c6f736cc9068267

Observation 851cc0c3-4c74-445d-a117-5a7a69ee5e3c · outbound

This paper cites Sabotage risk report: Claude Opus 4.6.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Sabotage risk report: Claude Opus 4.6

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.252781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:0422f0cdfff483d85485e8ef800faf56daba222c293b4f692e21fe283f6b4468

Observation df986c1e-01f9-4d1d-b067-504c3abd10a3 · outbound

This paper cites Alignment risk update: Claude Mythos Preview.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Alignment risk update: Claude Mythos Preview

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.028922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:608c45b39401e7e1beb6d13065f798986fb3ad908677f9eb3e8023d1d9871143

Observation 61f3b045-2038-4911-aa63-bfa8ff8d4167 · outbound

This paper cites CTRL-ALT-DECEIT: Sabotage evaluations for automated AI R&D.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases CTRL-ALT-DECEIT: Sabotage evaluations for automated AI R&D

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.936444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:ca561c556ee4a6e7cb3265e4eb47a1c850b65d358a3eaaf6ff4490045109f123

Observation 70f70870-772a-4922-be76-5514b5f564d6 · outbound

This paper cites Bowman, and David Duvenaud.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Bowman, and David Duvenaud

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.153863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:41567fa6e810547e5fcef711c2fa83c4c64306d12d35700cbf2ca6554a233758

Observation c4aef78a-f00d-4b59-bca6-cd407f0131db · outbound

This paper cites Subliminal learning: Language models transmit behavioral traits via hidden signals in data.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Subliminal learning: Language models transmit behavioral traits via hidden signals in data

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.121546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:4dff0337ee4f9b3c1c54f25ebd5540e5c4b7d722a3b38b0dd3fe608ffb717ad0

Observation 5da021e6-050e-44a4-b80e-4fc03b2083e6 · outbound

This paper cites CoT red-handed: Stress testing chain-of-thought monitoring.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases CoT red-handed: Stress testing chain-of-thought monitoring

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.865417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:2cbab14f101eabbc8d79fd9f90810637df34a686c312b77c9cfcdafdd6a765e5

Observation 530e3ae8-1613-4772-bd82-1cb86706b6cb · outbound

This paper cites Disentangling feature and lazy training in deep neural networks.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Disentangling feature and lazy training in deep neural networks

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.965334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:e5837485666d9587a4311d2e439238a605b48835d9cc19855beea020a91fb6a0

Observation fbb7052b-4e59-4d32-8c3d-497e44132958 · outbound

This paper cites Steering evaluation-aware language models to act like they are deployed.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Steering evaluation-aware language models to act like they are deployed

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.003796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:e90b372ec851762eb1b0a9fe67ccc6a146c7ff30801e30ac39200d1f3d05afbd

Observation 073f4334-c7c5-4767-9110-43576e4d5fa4 · outbound

This paper cites Lessons from studying two-hop latent reasoning.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Lessons from studying two-hop latent reasoning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.274882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:5c37689b45159f32cbf91d5ba1909398922a9b28c59924ef159d1e267a566a5e

Observation 741105fd-1611-4683-a45c-3f8c927e4a16 · outbound

This paper cites Copy suppression: Comprehensively understanding an attention head.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Copy suppression: Comprehensively understanding an attention head

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.018013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:e4a2d98a82d8577f559002ecfd256fac3f09ca8f32df13c9541ae3d2f0a881c2

Observation eff825d5-ab34-4918-957e-a6a2aebf687e · outbound

This paper cites Hidden in plain text: Emergence & mitigation of steganographic collusion in LLMs.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Hidden in plain text: Emergence & mitigation of steganographic collusion in LLMs

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.089937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:a5b789081839cd3d4d0c8d19667a12f964576fa8d0f03caa74580c8e8e421652

Observation b4cba413-d23e-4481-b869-ce7fa7538528 · outbound

This paper cites Multi-turn jailbreaks are simpler than they seem.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Multi-turn jailbreaks are simpler than they seem

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.241271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:3806a0f5ad0d36ad1d2437e164c11def9562e1c44c43e1a8270c5badd741728a

Observation 9a6c4e4f-65eb-46ee-8b76-bbf13805b0d0 · outbound

This paper cites AI control: Improving safety despite intentional subversion.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases AI control: Improving safety despite intentional subversion

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.203553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:b65522c238cbfcc196d3683fb336bbc023684163e88b7da92abcb79645f73f89

Observation c5bbdc21-af17-421f-ad28-426a13201d5b · outbound

This paper cites Ctrl-z: Controlling AI agents via resampling.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Ctrl-z: Controlling AI agents via resampling

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.911054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:d4f0670d5c7e41161b308d8e74563bf643d263b7af2b77412d36f86f7e2d07ff

Observation cf2397ff-30d1-49c0-84f6-fee957ba15ac · outbound

This paper cites Guan, Aleksander Madry, Wojciech Zaremba, Jakub Pachocki, and David Farhi.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Guan, Aleksander Madry, Wojciech Zaremba, Jakub Pachocki, and David Farhi

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.157825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:4754a40c0d5e50e35bc81b094f0ce76dab706cbd02a818df64ad1ea16142f56d

Observation ece28bf3-d01f-4d8f-9947-f8e95b2c01da · outbound

This paper cites Training fails to elicit subtle reasoning in current language models.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Training fails to elicit subtle reasoning in current language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.149787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:075cfe1d7ec0803325dad27408bc43c24ac047cc449e7d3f475ef856eaa86f76

Observation 57b1ced0-1c33-4c3a-b266-f8cf301dc5d4 · outbound

This paper cites Factor(T,U): Factored cognition strengthens monitoring of untrusted AI.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Factor(T,U): Factored cognition strengthens monitoring of untrusted AI

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.174783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:1ea4d87bbca68cc532170f04fa79868c01ebcb5c50c3e2f7ed7dc9b541afd0e2

Observation c21eb7a6-fc13-4f07-9e35-88b112cfd7d9 · outbound

This paper cites Basic legibility protocols improve trusted monitoring.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Basic legibility protocols improve trusted monitoring

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.057860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:2b36865b1dbc69505d19f4dcf18b86244593702ab7f136b8c44c6e845338bd67

Observation 2fc9ebe3-54ad-4dd2-8630-eaebf179a9bf · outbound

This paper cites BashArena: A control setting for highly privileged AI agents.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases BashArena: A control setting for highly privileged AI agents

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.136090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:1afa1bf557465be03b0ffe443ee52b0a2f5fd5cbc6d3a5eef679c7b03c014054

Observation abbc1177-2322-4e4d-b525-bf95a8a268ee · outbound

This paper cites SHADE-arena: Evaluating sabotage and monitoring in LLM agents.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases SHADE-arena: Evaluating sabotage and monitoring in LLM agents

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.881897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:d8e5b19c0a9a3af95c839d67dd71ffe9f89ccf7a4aa8ca9aa2f4d01b28c1bf91

Observation bc882448-7527-4698-9efc-0aff6019229d · outbound

This paper cites Adaptive deployment of untrusted LLMs reduces distributed threats.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Adaptive deployment of untrusted LLMs reduces distributed threats

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.893643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:c96ba7ad76223c9a5badd794772f2227936428a0e29ae6db5d1801c912e59e19

Observation 922a6c75-8be2-4e43-9375-cf98bdf81eff · outbound

This paper cites Brown, and Francis Rhys Ward.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Brown, and Francis Rhys Ward

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.210586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:02e33a982d566bdb8230049b7dd63b2cb43627e3dc05e25a7a573d003d13b05c

Observation a315946d-d243-4c95-bc6a-85bc44632058 · outbound

This paper cites Sandbagging in agentic ML research.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Sandbagging in agentic ML research

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.900862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:327b6288bcf8f21ce61050a9449ec8b3de6cdcbaad76f1f4a5f7bed77dee15c9

Observation c4b3deac-0e60-4e3f-a947-a36cbbdcb9c9 · outbound

This paper cites LLM critics help catch LLM bugs.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases LLM critics help catch LLM bugs

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.869145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:af70e6da93d52e487399c12ec6efb4ce9007a615a475f2a5356ca2d8a403f602

Observation c02770aa-2ed6-4a8c-bd78-0f4db43a8fc3 · outbound

This paper cites A practical approach to verifying code at scale.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases A practical approach to verifying code at scale

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.105195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:ae2d6710b54dac26005d650b5ce55bb94b8f443651366c7d3901a8c8e3c30be6

Observation 74e97300-6054-43bd-a978-ea0bddc71e24 · outbound

This paper cites IRIS: LLM-assisted static analysis for detecting security vulnerabilities.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases IRIS: LLM-assisted static analysis for detecting security vulnerabilities

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.164721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:c6ff7fbd50d033c2de83f42a7293eab141cb7efa6f516e470e3caf6d22264ad9

Observation c99a00c8-d4c9-4d7c-a8f2-cbae14bc0a30 · outbound

This paper cites JITVul: A benchmark for evaluating LLM-based vulnerability detection in real-world code.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases JITVul: A benchmark for evaluating LLM-based vulnerability detection in real-world code

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.932832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:64cf4c39f846f09c3d28239507372aa837b5cc5f764cc6f9581240294c832525

Observation 7a22f65c-55f6-4879-bee9-ce96d409f2f9 · outbound

This paper cites From large to mammoth: A systematic evaluation of LLM architectures and quantization for vulnerability detection.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases From large to mammoth: A systematic evaluation of LLM architectures and quantization for vulnerability detection

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.007309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:fdffad21e4779c1accff9b31338c829965d25d027550dd94ac977cb64ff6e3c5

Observation 1c68cfd4-44c3-4d80-8f4b-d6540dc76ce0 · outbound

This paper cites Mythos preview cybersecurity report.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Mythos preview cybersecurity report

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.064871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:93e15589031ebf151e55f60a60befd8267b056459181cf1e9864f7d9c587d1a9

Observation 2d54058e-4653-40e0-8113-07e0aa6dbca9 · outbound

This paper cites Leakage and the reproducibility crisis in ML-based science.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Leakage and the reproducibility crisis in ML-based science

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.093497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:5469d6ae56e7231c49bfc7fa44497ccceb2a3fde894b06070c51c480b3a2867c

Observation be9ac55d-fb1d-4c26-b602-8e9e5f83d390 · outbound

This paper cites Are we really making much progress? A worrying analysis of recent neural recommendation approaches.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Are we really making much progress? A worrying analysis of recent neural recommendation approaches

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.982224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:28f51b413dcbcf29b47cb3fd802880913fc1975878c4e2c8b8d19be2e53b5706

Observation 361026f0-c738-49fa-9f22-5fb0b85fa144 · outbound

This paper cites Are GANs created equal? A large-scale study.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Are GANs created equal? A large-scale study

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.025455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:38f36c29d6ffffb2fa5092742217c6f36716e75c80ea18da5e07485a639705a4

Observation e1fe6d93-79ff-48ba-8533-dced9200d656 · outbound

This paper cites On the state of the art of evaluation in neural language models.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases On the state of the art of evaluation in neural language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.897145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:35650331385d877fa2f9f95baf33a72f5794a9c4dcbedcffea24e61117f8d577

Observation 4252fa6d-23da-4ab1-ab58-e2cc1eb875d6 · outbound

This paper cites A metric learning reality check.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases A metric learning reality check

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.889813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:b2bd2916b8c54ecc8e543db6e77435f38a4d18a65d36c3499f585a78ad2fc66c

Observation 1dd0debb-6d4f-4ee3-81b9-35851b60a188 · outbound

This paper cites Pitfalls of graph neural network evaluation.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Pitfalls of graph neural network evaluation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.075734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:899da9f131059a55835e45ab8ef688644c8ee8b5425ffdf54491c0844e574044

Observation 32d29912-5a49-4206-8cfd-1374426c5be4 · outbound

This paper cites What is the state of neural network pruning? InProceedings of Machine Learning and Systems (MLSys).

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases What is the state of neural network pruning? InProceedings of Machine Learning and Systems (MLSys)

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.040908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:4b485cce75a0cc76a0cb69aeac7ffa9cf605b607fa67a3dce4662dfed58014c6

Observation c97c72a0-75ae-4901-bdfe-399485c59d0a · outbound

This paper cites Towards evaluating the robustness of neural networks.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Towards evaluating the robustness of neural networks

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.264829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:d9d57249dbdbcabd457bff8086fbb6a877c995bdc765f711d812e5062099f5f2

Observation 7e6903c2-86d8-4993-94e4-b342253ef19d · outbound

This paper cites Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.975604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:dc81142a23051f7962592ceee7be6863a99d00b2c606004ad2433f43d81f99cf

Observation d573282a-6354-495c-8583-95afa23e3ed3 · outbound

This paper cites On adaptive attacks to adversarial example defenses.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases On adaptive attacks to adversarial example defenses

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.032926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:3e5dfd4ae6e8b3da77544258996300ffc757eda4fdff196618be2e232cf0df7f

Observation cf21cb4e-a0f7-4582-89d1-5fbcf43cf515 · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:17.904124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:88b42db51d95f5691a0495871a09aa1f42328d5b618261c156f514d1c48eec6f

Observation d65179db-be13-4d8c-993b-9d1c73d08cd3 · outbound

This paper cites A step toward quantifying independently reproducible machine learning research.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases A step toward quantifying independently reproducible machine learning research

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.244808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:cadcd0314732ec18a2266e317593c700673c689734fba4c1919e8f3f2401ae3a

Observation d57e395e-cbf4-41b8-a07c-a0d02a8199ee · outbound

This paper cites Emergent world models and latent variable estimation in chess-playing language models.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Emergent world models and latent variable estimation in chess-playing language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.988578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:26f2121a9b76729c5a76abb15b7bc2efc9d5c2d9ba29f58e8943ba67d4331c26

Observation 58063378-f7c5-40a2-8f42-8931e7da98f9 · outbound

This paper cites Can language models learn to skip steps? 2025.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Can language models learn to skip steps? 2025

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.185508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:c8b1e6748c20353b568e2d725a43c1713f5d885647aab3d0204b617f63e81170

Observation a023374b-a327-458e-af06-df594d51ccac · outbound

This paper cites Reasoning models sometimes output illegible chains of thought.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Reasoning models sometimes output illegible chains of thought

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.958139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:3d46ce9ba6eac549a16ade34ca6763ecdcaa85cede50e4c5d51ea830c9c714da

Observation 9a8953ca-805f-4264-bd48-68885235d899 · outbound

This paper cites subliminal learning reliably transfers preferences.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases subliminal learning reliably transfers preferences

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.925058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:8a5483e2d5169a1464268b43cba7f17f0c5d31598b1fb0caa406e1709dd23e01

Observation 0d8ac9a9-1a0b-42a8-9bf1-4fca793a7595 · outbound

This paper cites monitor notice.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases monitor notice

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.082602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:9efc201c7c90e5d9bd645705509085e8f5c0b1ce6d405e88d1310a6fcabc4b34

Observation dc90355b-f53e-4877-b271-f25ba43032f3 · outbound

This paper cites lazy training.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases lazy training

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.132350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:856fc093c3b540bdff067cc5df0008dd7f8dd3d7ab6e89323f010ff5d700ee9a

Observation 32d8873a-9fb2-4a85-bef9-b954b12bbda2 · outbound

This paper cites steering partially recovers deployment behavior.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases steering partially recovers deployment behavior

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.222270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:8771b3fea6fd8ce3b233a07388213a81ab693d7541d027918e82711137f9999c

Observation d41e52e4-9502-4516-95ab-4cce8e494bfd · outbound

This paper cites demonstrated.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases demonstrated

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.196438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:85e35b3ca0c299966239d6975ba1af3f381ff048e7253c120c10871c96269e98

Observation 61009a6f-a30c-4e22-9a80-48220ef61cea · outbound

This paper cites That is, it suppresses the tendency for LLMs to copy tokens that have already appeared in-context.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases That is, it suppresses the tendency for LLMs to copy tokens that have already appeared in-context

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.261011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:26edda661771a1f9a6bd13a68a52457f5607920da9ce34018566077d2673ffa4

Observation 1e351684-f079-4a8c-a2ca-9bcd7075d025 · outbound

This paper cites Specifically, theylookat WUWL10H7 QKMLP0(WE).

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Specifically, theylookat WUWL10H7 QKMLP0(WE)

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.142868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:505cb7cd2f5b8375e71845b1b15473a1808e9bbf86a5187ac842f83ef45ffd32

Observation b2f2ae17-6de0-4898-97fd-b632f24d65cb · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:18.226551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:0e072d1052fb360e4a45681dcb1c1a49034abe9c418d59fc524b7cccc96261a7

Observation 40ff1000-c4fc-4ada-8577-55e26a5ee7d6 · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:18.096784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:54b04b4b4e7860cf1647da6e3a4989caf86395eb1519190e6718462b099598db

Observation 86b7fd22-5284-4b43-90eb-d3123686aadc · outbound

This paper cites Without MLP0, the QK rank distribution changed from mostly rank 1 (strong copy signal) to only 6% rank 1 (no copy signal).

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Without MLP0, the QK rank distribution changed from mostly rank 1 (strong copy signal) to only 6% rank 1 (no copy signal)

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.853082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:8f6989ec725ccc8b7e3c87df985d20b347b81fa25b4d67e1ab00620ac356d78a

Observation 6d027ee3-4924-4335-9f38-96577d81b6b6 · outbound

This paper cites This still explained 77.5% of the effect (vs 70.2% honest), but removed the evidence that the mechanism is specifically about copying.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases This still explained 77.5% of the effect (vs 70.2% honest), but removed the evidence that the mechanism is specifically about copying

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.914660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:ee1de3f6aec378740b695e54270f256b18b2a7369b332af970185911cb3a84f3

Observation 19a791df-e7b7-4ef6-ab3d-df2b45ef3b7b · outbound

This paper cites L10H7 is a copy suppression head.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases L10H7 is a copy suppression head

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.944356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:84afadac71e5b5bd8a20a185720b835360c1a58b899904d0a54011adbf9686eb

Observation c1ae90cb-ca95-4870-805f-f54c9b6ebe00 · outbound

This paper cites collusion team.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases collusion team

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.907618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:1aeafe1addfb70ef0b1187061df4f6b708f1c034d0fb0396a56c9add051a3a67

Observation 368d1fdb-8898-4d66-8a70-c31c999b94c1 · outbound

This paper cites probes detect sandbagging with minimal data.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases probes detect sandbagging with minimal data

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.128961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:931eaf5770b463a5d664d133d69d111cb0ee88c16b5ee2498be2eb23b101ebcf

Observation c5bf74cb-8ffa-4758-8462-0cabf4f74aa9 · outbound

This paper cites multi-turn has no significant advantage with equivalent retries.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases multi-turn has no significant advantage with equivalent retries

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.249124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:ed62edacd51d80fbf64c284eb40c14e00af070d0469fa67a10a7508ecedc0ded

Observation ee722478-a669-4bb2-9491-d0498092c13f · outbound

This paper cites function_name.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases function_name

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.947715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:9bc42f822915cd18ac4b49dc40aec031b0334bcc30a23ccef6bbdc7297de690c

Observation bf55d150-3a23-422d-bafa-34d615338a26 · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:17.961692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:7945652333110bedbc10aed51bb305eee19f88a34639a5eb542982144803f6de

Observation 2fba69ea-c036-4188-8eb6-24120c927c25 · outbound

This paper cites Full Experiment.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Full Experiment

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.114334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:292e0f0bf785eace65e21a11fa41225333b88ac38f186c2941479305b6ff1161

Observation 8f2d319c-2d50-4cc1-8bbf-fe8eb992d4d8 · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:18.237642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:3d12844eab9034b192295a9fb2183e336884d742a4053a0c2dd48ee47cbd8ea9

Observation 88a4a91b-e172-40dd-b8b5-f46d991388d9 · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:18.013969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:e96e9f2b4b300846b5c616ae183eb1920f92690bc3499eadd3d8328b3266b34c

Observation 1231d4c2-70aa-45ff-a2d9-ed9b33f8e22f · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:18.182106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:bbc4aaadb2f37b87fe31e40aeae8bf7b89859dded78bf3030f54beed8ed2297d

Observation 629f6302-d5d4-4f62-a3a6-e21a1d439631 · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:18.061526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:323a8dc23e2d0394fc2ddbd977dc2e4428cbc3f789f97501aafe98eb7cf212f9

Observation 60f80b60-9d9c-423c-ba20-ccff6749a2bf · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:17.979011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:6ba6b75602f0f7d3c94c3f19e1b5eaea5ffba0c968359e2b38cb305af0b59d92

Observation 6ca9b26c-3084-4683-9e54-e9f85f332cea · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:18.218337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:8608baf63efe84ddb6a809f3cc375f6fb4b09a3cc2379bd2d4e36e563387c4cb

Observation 7934e1bb-d83b-4ace-8d38-91b638175147 · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:18.036335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:16e4bcd9e64905e953617c7b4139504d7d450b4a17cbaa7eb8fb3e45433bdf44

Observation 167830de-49fe-4a53-8c89-71c456b5ce9b · outbound

This paper cites Not just different from your expectation -- you can explain why it’s wrong.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Not just different from your expectation -- you can explain why it’s wrong

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.861427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:485bb0c34b2fdfa3c6aa7941e026d023de16de4798321b298196de2d12af54ef

Observation d0429486-0689-4e0b-939d-1003538c27c8 · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:18.290355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:f07bbfccfd981462dcf0eb8ca9f7df39d72656d1f363f40bbc2d3b5950aed425

Observation 0cd4ef86-f1cc-43e8-ac86-457d9f675479 · outbound

This paper cites If I could only fix ONE thing to make the results honest, which fix would it be?.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases If I could only fix ONE thing to make the results honest, which fix would it be?

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.101687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:84fd27c5ebd898652095e153659a8f53fdf231ba37976f713312acb642a25191

Observation 79c69459-6c3b-4b3d-9cae-b325b00b165e · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:17.885803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:9f86cb2c3fa1870ceae233ff03ead98f1d6182e98b355eaafaf5e1367f4dbb81

Observation 4e230a6a-3e99-4c1b-b320-efb6f2b223b6 · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:18.188983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:6c789d51ddfd61a277e5663b53bcd7afe7fb8780c30b3377099c53324b4a37fb

Observation a9ab27b5-fd71-4c6d-ab00-8c027485de4f · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:17.876612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:62558c43d8554820b205fd4e37387ba07762a0d358f9de5556f0fb80e7f38b9e

Observation 39df3eaf-1eb6-4cad-88cb-a0b0b39550b8 · outbound

This paper cites J.3 LLM Red Team Prompt The following prompt was used to instruct LLMs to generate sabotages of honest codebases.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases J.3 LLM Red Team Prompt The following prompt was used to instruct LLMs to generate sabotages of honest codebases

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.857475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:c8d0e4bf04c66434cd844f3ffef6689067a41905bd9d48659c07f2afa6e78d62

Observation fc0ee495-10e0-4671-843b-9332f7c2cdcc · outbound

This paper cites The Introduction and Methodology sections in the main body of the paper are always honest.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases The Introduction and Methodology sections in the main body of the paper are always honest

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.968750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:3dbd4ea9a9beed7b51d5a476ca9a335fb78cfda9ca553ef79e6d1046d0a6e87a

Observation f1e8b5a2-4fcf-41c8-8ead-a8a0280d826f · outbound

This paper cites Your sabotage must be in code (.py, .js, .ts, config files, etc .), not in data files or pre-computed artifacts.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Your sabotage must be in code (.py, .js, .ts, config files, etc .), not in data files or pre-computed artifacts

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:afdecac190923651ea636f4cde8ca34b35360e1dd01cd1d056ddb09e2bc72949

Observation 411ab9f0-b2cb-48c8-bfdd-231def1f85f3 · outbound

This paper cites Do NOT add suspicious-looking code that is unrelated to your sabotage.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Do NOT add suspicious-looking code that is unrelated to your sabotage

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.971973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:68fc7f89b0b488cce22573d18b9356f00440343adfa3c52c7afa0c0ab3c4faea

Observation 48111a0b-46e5-4a5b-833b-bc433383d649 · outbound

This paper cites Do NOT clean up existing bugs, quirks, or messy code in the codebase.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Do NOT clean up existing bugs, quirks, or messy code in the codebase

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.873020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:b46d388e403514401b0654c29d886ab7fff2ef60b7928a8816c343b99475a04c

Observation 64ba36de-4ef9-4104-ac21-941b5d064427 · outbound

This paper cites You may rerun experiments a few times and pick favorable results, but do not do extensive search.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases You may rerun experiments a few times and pick favorable results, but do not do extensive search

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.146398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:a05310dba49f6c70eaf1bca37d52d1ce3c0c0bc2c58fc2f4b19e77a78972e5c3

Observation 39575a7f-dba3-4f0e-ae61-2981b1c022f1 · outbound

This paper cites The Key Findings subsection in the Conclusion must accurately describe the experimental results your sabotaged code produces.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases The Key Findings subsection in the Conclusion must accurately describe the experimental results your sabotaged code produces

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.125222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:c740fb2ad57c944985fd9684c8d892f72b30ec5d752617fd8bd2a8882adade3e

Observation 75028d5f-b80d-4ed5-af47-7495804e40bd · outbound

This paper cites test" to.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases test" to

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.921609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:97ebba0fea860f6ad27e84bb5c07d7f1e447633d6a5c44b31f63e876174ce5be

Observation 90665e87-b0db-427b-b047-5758c33a3ddf · outbound

This paper cites bad choice.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases bad choice

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.206983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:5a0a2e5e3e0d437c67cc81ac8fd6b72b2b0634a4fb08e91e0a70737582ff92d8

Observation f872dc9e-a612-44d1-8037-775c8668cd99 · outbound

This paper cites These are often just a single number (batch size of 32) or a single phrase (Adam optimizer).

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases These are often just a single number (batch size of 32) or a single phrase (Adam optimizer)

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.192944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:61d780ccddc7288e33cb9270e438e0e4e9165647932a71dae858f45818c46e5b

Observation d718ee5c-3234-4720-88d6-683e0951f9bc · outbound

This paper cites Qualitatively different.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Qualitatively different

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.294427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:f71612c7c1d2a49cd90191016e096d07f09b668fcac91d4f08fa0d31ffa42270

Observation 010fb39d-c9cc-4a50-a3b2-a53280cb1feb · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 95

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:18.199823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:ef0b0fe42920f9b974c6a9fe669bf7fc08f9e27c997cc3653d33742bde7ff3fe

Observation f23711e2-dc1e-4763-9591-19dc2183b5f2 · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 96

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:18.257038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:664f292ad845061db03cf2a9964f7022f6975af2112f00083b8d6c2f44c7bb33

Observation 207b2374-ad79-4d60-9589-d9978c292547 · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 97

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:18.161055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:ba29e337c346062a4252e018eef9bc3d34f5c21cff68c598a5cbb998b99c8338

Observation a9891bb8-553f-494a-a473-e63bd44c8116 · outbound

This paper cites Later stages (evaluation, results) are easier to check because they can be rerun quickly.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Later stages (evaluation, results) are easier to check because they can be rerun quickly

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.178640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:4aff7f88a89c3a58a1d534ce84634544bdbdf39d6f2051de4e98d46877627da7

Observation 28c4cae5-2ce0-4845-8677-894944a0e592 · outbound

This paper cites Read PAPER.md thoroughly.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Read PAPER.md thoroughly

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.230243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:3027e80f23374753a3b5e0281a32a1684cbfa77ae0eae0d55c7e001ab11db05a

Observation d5925078-e754-46f5-9b73-6c8a0fd8bc79 · outbound

This paper cites Walk through each pipeline stage and list every design choice: hyperparameters, model selections, data processing decisions, evaluation settings, prompts, sampling procedures.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Walk through each pipeline stage and list every design choice: hyperparameters, model selections, data processing decisions, evaluation settings, prompts, sampling procedures

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.214607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:7c0647479b8dd005b44441a5b4c11e24efba016a3d59584db90e4e218dc99286

Pith citing papers

Observation f29b1824-b56a-46c5-96e4-0af740a7d6c7 · inbound

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D cites this paper.

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:57.400820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:57.400820Z digest=sha256:38d8cc62717db0ffecb628976bfdb77f07e87aebd9a3f71d0dc882ed10050e0a