Pith. sign in

Paper Citation Record · LEDGER

Evaluating whether AI models would sabotage AI safety research

As of 4 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2604.24618.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.24618 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T03:31:45.082467Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0efa6d82-172d-4935-8c76-68d132285737 · outbound

This paper cites System Card: Claude Opus 4.6, February 2026.

Evaluating whether AI models would sabotage AI safety research System Card: Claude Opus 4.6, February 2026

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:07.089446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:0f0946f3828e8e061073cc99d0c93cdf57dc478e5bb025c3e7841134c4dedd16

Observation 5c7e97fc-bb18-4423-b73e-4ed9d272ec63 · outbound

This paper cites UK AISI Alignment Evaluation Case-Study.

Evaluating whether AI models would sabotage AI safety research UK AISI Alignment Evaluation Case-Study

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:07.063065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:4f0fc3f3085e158994c3e72a967493ac507ca4f966597757c96b41c17c0bc5f4

Observation 0cded5d0-dd27-4eff-b87c-88bacaf5b424 · outbound

This paper cites Risk Report: February 2026, February 2026.

Evaluating whether AI models would sabotage AI safety research Risk Report: February 2026, February 2026

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:06.888995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:df4997a7086ec800bc190becddf90673967abe81af4b39c0ee040df57a4e72a8

Observation c6a9ecf9-a2f3-4841-b97d-fd9a2c1145e0 · outbound

This paper cites Bowman, Misha Wagner, Fabien Roger, and Holden Karnofsky.

Evaluating whether AI models would sabotage AI safety research Bowman, Misha Wagner, Fabien Roger, and Holden Karnofsky

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:07.056286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:679fd42345cb8a8500f17d340c0345c5047126e5d160e34f64ab3ec56e943a9e

Observation 019b5b08-a9f0-4cbe-b70f-5b4361e45f09 · outbound

This paper cites AI Behind Closed Doors: a Primer on The Governance of Internal Deployment.

Evaluating whether AI models would sabotage AI safety research AI Behind Closed Doors: a Primer on The Governance of Internal Deployment

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:07.013327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:8bf224aac1672c044608650ec4f4578cd775475bb10d4a36aefd5d8a5cf3e787

Observation 6f06a5be-fbb3-49a1-a4cd-3305cd4c9c9a · outbound

This paper cites Zimmermann, Ziyue Wang, David Lindner, Victoria Krakovna, Sarah Cogan, Allan Dafoe, Lewis Ho, and Rohin Shah.

Evaluating whether AI models would sabotage AI safety research Zimmermann, Ziyue Wang, David Lindner, Victoria Krakovna, Sarah Cogan, Allan Dafoe, Lewis Ho, and Rohin Shah

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:06.900228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:73e435ccd3cca0e6ac483cbbf0d8ebfe14505b0111a2649c3e34a275cf299342

Observation 5c7b2643-31da-46f3-a2d0-ccbdbe4df02e · outbound

This paper cites Bowman, and David Duvenaud.

Evaluating whether AI models would sabotage AI safety research Bowman, and David Duvenaud

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:07.085624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:fb64dde0b9b4256f56c9b54ffb7a019251fc3681feb002a5e0ab209ac3cca673

Observation a6715866-17a4-43ee-8f36-41f18a0600f2 · outbound

This paper cites Troy, Stuart J.

Evaluating whether AI models would sabotage AI safety research Troy, Stuart J

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:06.996714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:f1f31a01a0fa9f977f07cdc829274b0e9cb83b0ac6134c8b37d0bd1e7f1ac368

Observation a995fdf9-ae54-4d97-bcbb-6438cf4d4510 · outbound

This paper cites Investigating Claude Refusing to Assist with AI Safety Research, December 2025.

Evaluating whether AI models would sabotage AI safety research Investigating Claude Refusing to Assist with AI Safety Research, December 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:07.031441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:b4d772f03347f9a5b4ea37e2009dffa5f63e5fb6d9d0f5bfa5a7d86b9154928c

Observation 74f51f1e-61d1-4164-a22c-519fe029b0aa · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:06.912732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:5dbb25d0311a3e7b4a2e61af751ffa22d1b7216ced119d1d954d3cb3d18748e3

Observation f55f0601-12b5-4659-a853-34fb810eb569 · outbound

This paper cites Measuring and improving coding audit realism with deployment resources, March 2026.

Evaluating whether AI models would sabotage AI safety research Measuring and improving coding audit realism with deployment resources, March 2026

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:07.006510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:bad1f075821710fade69425f0d56e776708d115fade3834c0c965ab2ed22848d

Observation 6c8c4618-1693-43bb-b32e-338881582464 · outbound

This paper cites Do models continue misaligned actions? LessWrong.

Evaluating whether AI models would sabotage AI safety research Do models continue misaligned actions? LessWrong

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:07.009931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:97a847b6d78a5908759dd7c0825ccce4722be1d898438863d26a99d3e83e727d

Observation 25fe1ecd-8919-4e74-8c9f-1d0cbab45512 · outbound

This paper cites OpenClaw — personal AI assistant.

Evaluating whether AI models would sabotage AI safety research OpenClaw — personal AI assistant

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:07.043526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:de0e73a71f53ca0522fd183b05d4f247ac6a6d8e92b011bfb1060373f66aaaa3

Observation e68ccbd9-b0ca-4ddc-8dd1-18936478ce32 · outbound

This paper cites Constitutional Black-Box Monitoring for Scheming in LLM Agents.

Evaluating whether AI models would sabotage AI safety research Constitutional Black-Box Monitoring for Scheming in LLM Agents

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:07.069738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:76abee158dcd82010ac0af020666425b548bf23f57bf6eb62209bc6314c1f036

Observation 774ec82c-06ae-4114-b606-f9721cfdcdb7 · outbound

This paper cites Large Language Models Often Know When They Are Being Evaluated.

Evaluating whether AI models would sabotage AI safety research Large Language Models Often Know When They Are Being Evaluated

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:06.859813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:5c6b7ec53a65c8e56c9a305be5744ad8ee276a068221c030a6b1df36b18992ad

Observation 3c8308f0-9f35-4268-bc66-55933649a861 · outbound

This paper cites System Card: Claude Sonnet 4.5, September 2025.

Evaluating whether AI models would sabotage AI safety research System Card: Claude Sonnet 4.5, September 2025

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:06.916462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:33f39d7e254aa6d0627c0c7d62875921caf45f1b2d6be96bdcbb50308f526442

Observation 60f1cdd1-6988-4744-b69b-e809922396b1 · outbound

This paper cites Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations.

Evaluating whether AI models would sabotage AI safety research Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:06.885523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:cad2523e8336926375425b3b11638a81bfd6a66f089dea86bb2a757fff3d8ca5

Observation 5d9cfb11-fd6f-4f5e-9c97-5c8a67d4bcf1 · outbound

This paper cites The Hawthorne Effect in Reasoning Models: Evaluating and Steering Test Awareness.

Evaluating whether AI models would sabotage AI safety research The Hawthorne Effect in Reasoning Models: Evaluating and Steering Test Awareness

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:07.049992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:4a64d1b66a7c8a84279d15b293380a7883a2c0f519231798f780f609b0a93111

Observation d27f8249-9e26-4fbf-8410-0ef24442a982 · outbound

This paper cites Me, myself, and AI: The situational awareness dataset (SAD) for LLMs.

Evaluating whether AI models would sabotage AI safety research Me, myself, and AI: The situational awareness dataset (SAD) for LLMs

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:06.958383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:91e2277014f9e7b54ba11f89faec398e9f46b7172ed86ee0fe01e9b7239a23cb

Observation 6508954e-c3f0-4227-bab6-b06f74fa5ed1 · outbound

This paper cites Prefill awareness: Can LLMs tell when “their” message history has been tampered with? Blog post.

Evaluating whether AI models would sabotage AI safety research Prefill awareness: Can LLMs tell when “their” message history has been tampered with? Blog post

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:06.882199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:f6baff29eff637b3f3edb41909e12bf3b0622f922df6387841cccd20d642c95c

Observation 19eb7ed3-4f0d-4c36-9f59-3a48f92e1e5e · outbound

This paper cites System Card: Claude Mythos Preview, April 2026.

Evaluating whether AI models would sabotage AI safety research System Card: Claude Mythos Preview, April 2026

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:07.079502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:a2c7ae4c1960493fb20f2f05fa7f198a89f87eb8658905177072b0d0bb562bf7

Observation 1e3a47e3-23b0-4ffb-add4-2745d817662d · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:07.002910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:b9308bac981d931aead37f7542bf00394c8480047b16b8d625716b6a7b8aee37

Observation a9915f1d-3db3-47cc-85f0-4c0475e33107 · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:06.936079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:062a8b13208e8f8d8041a62213f8c09be7967935501b965840b7f9ae0385aaa7

Observation 72245175-b9ec-4022-bc4b-cf2282324665 · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:06.970714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:299180a985e2e00e6d7e0aeabba198e8935d91519360a9b86ab157db6641ab51

Observation f37e0238-3474-4c8d-9611-e2fb193a76f5 · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:06.967375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:18709b1afcdab88c72d7cee9b335286ac4178d40040a069fbef44e00998bf742

Observation 17ea2cb2-c2a8-49f8-b133-c6461e213fe0 · outbound

This paper cites 6.Reduce hallucination.Reducing hallucination in model outputs.

Evaluating whether AI models would sabotage AI safety research 6.Reduce hallucination.Reducing hallucination in model outputs

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:06.983611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:b2137aa35d4c62d639761e5a461f41bc455c1289a1d10fe6af2fe3e4116184bc

Observation eed0cd7a-acdf-48c2-a2df-830c3020467b · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:07.019306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:9e13324ed7a35c0a44c0b84851d973541a17cd36fdf001c83e0723d64c797027

Observation e1536a47-35d5-4157-875c-e5ec14bc7f90 · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:06.945917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:e896c13cd542f80470323f68be288adcedd894e9ae646a754abbf9ed4dc09823

Observation fd567af4-5e61-44d2-9adc-2f4fab9da82d · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:06.928623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:2356819b11173898a32acabe4ba8bf1d6d7b0d6b2ffe5499bdf928bfb668e577

Observation a8dd32db-fe30-48b5-82ab-7e229e96960e · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:06.866724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:8afa03786c2167ce5b09c58ce38d7ba0f593eaf3616e85f71ee5e4510a9685be

Observation 144b4679-be5e-4b56-ac95-5590f28dd7ad · outbound

This paper cites 6.Reduce sandbagging.Detecting capability hiding and strategic underperformance.

Evaluating whether AI models would sabotage AI safety research 6.Reduce sandbagging.Detecting capability hiding and strategic underperformance

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:07.059533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:01bf52c40006b9cfe2d05aa6aab1cf031efc9133aa71f9dc4577765f90de80ea

Observation b80d9a0a-c7df-4494-a9d6-ba236c608bb0 · outbound

This paper cites 8.Reduce self-preservation behaviors.Measuring and reducing self-preservation tendencies.

Evaluating whether AI models would sabotage AI safety research 8.Reduce self-preservation behaviors.Measuring and reducing self-preservation tendencies

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:06.878139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:edc867cb38cc91482a23dbcade7780c4a9958ff0ce9b8eac6fb296a06fc65ff3

Observation c0276ede-790a-470b-8285-1ef637180187 · outbound

This paper cites A.2 Research Activities Each research activity defines the type of task assigned to the target model.

Evaluating whether AI models would sabotage AI safety research A.2 Research Activities Each research activity defines the type of task assigned to the target model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:06.976352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:3e0b1f177c0af7da3d5a1fcceec618a88e765e585dd88ead97972a1d9cb1539c

Observation 5a143323-6e8e-4014-a02b-04dcaeb0c74d · outbound

This paper cites The PR contains a subtle but critical bug (e.g., a safety-relevant computation that is logged but never incorporated into the training loss).

Evaluating whether AI models would sabotage AI safety research The PR contains a subtle but critical bug (e.g., a safety-relevant computation that is logged but never incorporated into the training loss)

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:06.955040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:f9934bd81d7f7ce53ca88d53b11b958a666d6467235c248e984a3a78303aab0f

Observation 1e93f6d1-cf1d-4cca-b777-7219c070f6f8 · outbound

This paper cites Sabotage surfaces include misconfigured runs or misleading analysis outputs.

Evaluating whether AI models would sabotage AI safety research Sabotage surfaces include misconfigured runs or misleading analysis outputs

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:06.964245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:890bd77558800fd8ba3c49b382d12041eb0e70408b2a5c986399dac4669911db

Observation d85ac03e-538d-4579-8ef4-342dd7e95dba · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:06.922461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:e2d782824e5e56d652453af9b9a67e22fb10c82d7c820bcf39c1e02f3c961f46

Observation bc6b42cd-a985-43d3-8e60-75aab20e1513 · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:06.895736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:aec8af7c8e40b3ceee724da5175fd302fd787d37d11c3fb63304baf15c644f77

Observation 51d7302f-05c9-4e7a-8cc7-b250239ec38f · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:06.973293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:f686bbbec437fd06344f94dc02bac47267f97a4b95b072e239ce31c3003d1082

Observation cc10bcfa-8fb6-49f7-b697-26a21b607e2e · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:07.072765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:75b6ebae2df20c8c874078a6de11c62a3049f234a95d5b18fbd5277f74ecf9c4

Observation 03585f8f-0db3-40dc-9e76-0ef8fd881ed9 · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:06.874804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:aa46ff3096f9939451feb3991366b305617f53abfd743bc39d4e02457b6ae463

Observation e99b6a83-1589-4511-abcb-7ef3e5fdd4cf · outbound

This paper cites Current Research Focus.

Evaluating whether AI models would sabotage AI safety research Current Research Focus

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:06.863377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:539b2009fade2bd8b47396695d8d55338a44c84d878266702968dc3bde795503

Observation 1176a242-000a-4d59-8208-37346f84fed6 · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:06.993559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:b8dbd91a014c281af05355b02ec834eadaa6a9608dfced844fbecb560c0064dd

Observation e5315528-0332-4a58-9214-662f5746fbbb · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:06.990308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:e10ed022dbd4d67ddeb8382f6458b269a9012868bc1709b6268253f6b20e9842

Observation 47a77cf2-39c7-495b-ba13-bb12a057319e · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:06.986856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:debb30e69f23e25a41ff11ea130896d0fdf98f5c869f3979603bd0c4d693bbf8

Observation f939c231-c6b8-448d-942c-fb701c2884a2 · outbound

This paper cites Lead with the multi-factor framing.

Evaluating whether AI models would sabotage AI safety research Lead with the multi-factor framing

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:07.037401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:26a09cbeb489223864eacfc14a6f947d5bd30e71c970bc2ebc53ceb381d8f131

Observation 813776ab-59b1-4bc9-ac83-072a4116771b · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:06.909183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:60e43d1cbf582cade456c89001b71300d3af7d57a06af60ebed5d3730a11844e

Observation 97564542-74b3-471f-985e-2873f1e991a3 · outbound

This paper cites #2 SAFE-2061 behavioral training (satisfies the check).

Evaluating whether AI models would sabotage AI safety research #2 SAFE-2061 behavioral training (satisfies the check)

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:06.925585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:de9ae13e2d35fecbb1fb3688669df13d0c037f940b39f3bc6caa6af1652bf6c4

Observation eee96503-fd09-430e-b6e7-b3ef2dfb2003 · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:07.053070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:62ba68629777e5e3a9e98040f6af94a3f8a23fde812c903efb8d2720f01b0455

Observation cde4b3b2-c577-4cc0-a9fc-e0b67bcafecf · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:06.939119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:6cbe011cbba3e53255c92ce621e55ca317517801cfd47e9105952de1e43f9adf

Observation e99d12d1-eb69-4320-b1a9-7619cd8ca848 · outbound

This paper cites The comment I left frames it as a false-positive fix.

Evaluating whether AI models would sabotage AI safety research The comment I left frames it as a false-positive fix

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:07.076472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:bf93ba13c3add740c8c81ade3d1a158cb816ae30ada7517140ce33532469bfaa

Observation 7cae80b4-5d82-4a47-94e6-7d6dafd8daff · outbound

This paper cites If Amara reads them side by side she’ll notice the framing shift.

Evaluating whether AI models would sabotage AI safety research If Amara reads them side by side she’ll notice the framing shift

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:06.951983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:cb02c5f2e7b998d65458d63bc58ec5fbcacc223a723e52eb91698692f9e6dd70

Observation 623ec018-0ef4-4f05-bdbf-12239cc19c10 · outbound

This paper cites Anyone who actually opens the trace file will see ”The system prompt says to stop and alert.

Evaluating whether AI models would sabotage AI safety research Anyone who actually opens the trace file will see ”The system prompt says to stop and alert

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:06.979965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:6c24be463b2e299c840228034b4fa7e5653a63d0b44185a6cfd9de7a97b674e8

Observation 218ed936-0f06-4d33-80cd-94ee56f790db · outbound

This paper cites It’s there, acknowledged, just.

Evaluating whether AI models would sabotage AI safety research It’s there, acknowledged, just

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:06.932186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:0a22b01d3debf2efa5d617cf45ac4550d43c14d0a527e8f1c1187ff5859bdffd

Observation ebacc227-82be-4237-86c5-00160961ee3b · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:07.040590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:830eadfdbafef40371cd082e1237fe1e5c0f99e7eccbcca30b56a5a898e526c5

Observation c7983015-4517-4721-8e3d-c5d7e8de00e5 · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:06.869351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:9f33acc7b583c37d68264bbf3d117d04f9e589154201a48bd02b04b180eb8a45

Observation efff19b4-d414-4e13-9c1b-79b0e693d6f6 · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:07.022199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:9e600270f10e00b4db726f3226127d0a29df7a6716f1918cdcb8bc3995d130b0

Observation 5476fc0e-4fb1-4b18-8674-4de2121a8b9c · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:07.046558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:5d60734c76a55cf11b0249db5d55bb253347ceb0f962266958d3740cccade68b

Observation 97fad1be-add9-47e6-8e37-0f0e5e01e75e · outbound

This paper cites 113 114 --- 115 116 ## 6.

Evaluating whether AI models would sabotage AI safety research 113 114 --- 115 116 ## 6

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:07.016330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:995235deb612aa509868964b97a7d47a0bc9fe3a269ef89327227cba6382bda0

Observation 6163e88c-fc7a-4048-ab71-715a23e8f01f · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:06.904706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:f5002fe927f532beef3511b5c009487cfe0ce76c7c78fc03ae85abe3d7d0ac78

Observation f52ffa98-3af3-4897-8d21-5e57ef69956f · outbound

This paper cites If you’re working on tasks that would benefit from tracking progress, consider using the TodoWrite tool to track progress.

Evaluating whether AI models would sabotage AI safety research If you’re working on tasks that would benefit from tracking progress, consider using the TodoWrite tool to track progress

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:06.892738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:e4dfa9e1dcc14ae757fb36797aaa2011edd7f00c0526eb4bd90ee69f63c40813

Observation 38d95033-bb8f-4c4c-a6bd-979e5fb6269b · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:07.025019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:6163c424f94970740ec75855da13aba282f2dd681eccf9dae15d214ddc85d16f

Observation 6fd3ff51-0687-44c6-b57f-b4a6fda9e1f8 · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:07.028071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:520ade506b455a1f251f931830018c7a9d7105266711405809c3aa6f448b93e4

Observation 667e5e91-45bf-41ee-82a3-fd9c62a31315 · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:07.066358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:115220c5602f428cb0d9383e7eae34628a098031ba0cc1dc3ec73c8daed31dcd

Observation be51fb84-f468-4b74-b023-f965d26dd7a4 · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:06.948871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:87c625115796c2e1f31b2805aef7ba987ab6812f3aec3d38a91fc8c488904ec1

Observation bc29a5b2-2a3c-4d81-8c48-de77a483c093 · outbound

This paper cites There’s no evidence of overfitting.

Evaluating whether AI models would sabotage AI safety research There’s no evidence of overfitting

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:06.919870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:782a87fe392f3265a6837c50ab555bd4cc1d789dc34a8b836d98e6b72bef2619

Observation 4e95c3bd-360f-436c-8718-6ad909577890 · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:06.871987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:a0c59ba263013f7fd82fa3bbf6b7e00909a01f5c142408752c01b8b58c9ef74e

Observation 20b53840-8569-4e08-8fa4-0709422decad · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:06.999982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:fc2027f1f2b6947d5f38aac7325d14435a5c5eddf1f7aa22e62f5f17c4ee80c8

Observation a38c8b7e-7dae-4af1-b32b-d5c2b03507d9 · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:06.961222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:2240131a7caaa946919f7618b1f45dec722abaf5528ad5e83e3448c234cdd103

Observation 8c0814d1-a5c1-4b8b-8b65-3fe52d28f91c · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:07.082560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:7e4b4d95eb8427d747ace830bde496bcff6a829804b9f6a2097cf941eb453c62

Observation c43ac5d2-8887-4bdc-a3d8-69a76dbd4177 · outbound

This paper cites If there’s a sandbagging problem — and the evidence in that document is strong — then fixing it is exactly what should happen.

Evaluating whether AI models would sabotage AI safety research If there’s a sandbagging problem — and the evidence in that document is strong — then fixing it is exactly what should happen

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:58:06.942857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:010f5ea0b047f5e7920c6338ac864992581a6367bf05d0ea777273d0d11aa085

Observation 7d1bbb73-58b9-401f-aa44-5a7a2e2095a2 · outbound

This paper cites an unresolved cited work.

Evaluating whether AI models would sabotage AI safety research Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:58:07.034244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:31:45.082467Z digest=sha256:05bc4bb638eaedf90e3bd76b43c0c1b1f2ba3db81328b75c97b845b7ff7207f3

Pith citing papers

No inbound Pith citation observations are available.