Pith. sign in

Paper Citation Record · LEDGER

Interpretable Risk Mitigation in LLM Agent Systems

As of 18 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 3 inbound Pith citation observations for arXiv:2505.10670.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.10670 v1

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:11:02.149478Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:51:56.450815Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T00:31:24.634056Z

Reference resolution

77 of 77 outbound references displayed

  • verified exact2
  • verified fuzzy24
  • unresolved49
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c43eafac-f59d-4479-8e8e-f89c0279a822 · outbound

This paper cites Artificial intelligence and the future of work: Evidence from OECD countries.

Interpretable Risk Mitigation in LLM Agent Systems Artificial intelligence and the future of work: Evidence from OECD countries

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.643143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:01.782535Z digest=sha256:27064381432ff4d9fd95f70a789b02f159695e21aa4dc8bd2b8a99d6037a2b94

Observation bf3d3a32-c723-4ab7-a21a-243b73e9489b · outbound

This paper cites Dai, Chelsea Finn, Justin Fu, Kanishka Gopalakrishnan, et al.

Interpretable Risk Mitigation in LLM Agent Systems Dai, Chelsea Finn, Justin Fu, Kanishka Gopalakrishnan, et al

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.627827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:01.787366Z digest=sha256:b81f04d741e1edce2f09d3526c08aec6eafcda74d8878642f6c7b461a0c92a3a

Observation 3801ed77-8037-4a96-b87d-1b592e49f25f · outbound

This paper cites Mistral 7b: Open foundation models, 2023.

Interpretable Risk Mitigation in LLM Agent Systems Mistral 7b: Open foundation models, 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.612474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:01.792100Z digest=sha256:a9e6889e8bd6985109b2b02ab4cfb9812e25d83b54b93cb54049fe685cd0d2d8

Observation 54ed5ba8-5bf9-4e48-a28b-7b79fb7685e8 · outbound

This paper cites Playing repeated games with Large Language Models.

Interpretable Risk Mitigation in LLM Agent Systems Playing repeated games with Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.796811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.796811Z digest=sha256:a6c202568d4da1e03a64a5c5c29fee95324ea213d56bfe61e168cd05bed08d25

Observation 9b7eff9b-d167-419d-ae41-da6ca8c46fe7 · outbound

This paper cites Concrete Problems in AI Safety.

Interpretable Risk Mitigation in LLM Agent Systems Concrete Problems in AI Safety

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.802560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.802560Z digest=sha256:64f3ba94e926c705ebe31eab8794859355eee731ffffade0ef742b432afa7fe3

Observation 16702f8c-2cb7-4d44-a2c2-8387f4b0ae83 · outbound

This paper cites an unresolved cited work.

Interpretable Risk Mitigation in LLM Agent Systems Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:11:03.597680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:01.807791Z digest=sha256:694cd91ca2484ac3b3c1b8952b6ac42c3cbf47dfaeffa2395beb20651231095d

Observation 94b9346b-084f-4af5-a678-a12f9bd12757 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Interpretable Risk Mitigation in LLM Agent Systems Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.813065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.813065Z digest=sha256:1ee75f2f0e3cbe7ef267eb5d727a81fa27a8b38ccbc2b87d40a856339c5799c8

Observation 5e6eafe6-a930-4240-a93e-aaa30f5aa5d6 · outbound

This paper cites Emergent tool use from multi-agent autocurricula.

Interpretable Risk Mitigation in LLM Agent Systems Emergent tool use from multi-agent autocurricula

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.580794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:01.817697Z digest=sha256:0d293d169a90800f86a630b1f5ea0a4e9c85afe9af43b2bf31042358845efe49

Observation 351ad7f1-6ab1-403d-b7ff-ef072b69c117 · outbound

This paper cites Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell.

Interpretable Risk Mitigation in LLM Agent Systems Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.562891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:01.822090Z digest=sha256:5cd450fb34dfe30918e2f682827fc4a2376afb0ab05a572e2ec78924625252e7

Observation 8a106779-c66d-48c7-8c48-b0f78d8abf43 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Interpretable Risk Mitigation in LLM Agent Systems On the Opportunities and Risks of Foundation Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.826932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.826932Z digest=sha256:d6798987efacfb73b1ed9dc03569b495cdc07a57d3c4b06a96353f9b158bac71

Observation 2a7adc57-b623-480d-be0c-b1dd7e69a6ee · outbound

This paper cites an unresolved cited work.

Interpretable Risk Mitigation in LLM Agent Systems Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:11:03.544642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:01.832003Z digest=sha256:02c3a02f13ca2512e7b515eec8c270299c16a5ca67af3f4e00bd152f22f67f19

Observation a338b78c-533d-4e4b-9b30-639f80356236 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Interpretable Risk Mitigation in LLM Agent Systems RT-1: Robotics Transformer for Real-World Control at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.836214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.836214Z digest=sha256:e2ab9f0b67c31bb4b7eb9b876485080f3f236c3d545bb41ab06b1447d14634d0

Observation 16d22cc2-e43f-4ec2-aff9-61a6d878a1a6 · outbound

This paper cites Playing games with gpt: What can we learn about a large language model from canonical strategic games? SSRN Electronic Journal, 2023.

Interpretable Risk Mitigation in LLM Agent Systems Playing games with gpt: What can we learn about a large language model from canonical strategic games? SSRN Electronic Journal, 2023

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.526152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:01.840725Z digest=sha256:f1b3bd0b7a4917e31ea2e46791cb0484f1b62cc1328532242e36fa5cb2d3c0ad

Observation fa8b734c-b330-4c95-85f8-426bab50f0df · outbound

This paper cites Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D.

Interpretable Risk Mitigation in LLM Agent Systems Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.508895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:01.844793Z digest=sha256:8c9035a4299e134269d8359a413b7edc961f59bd2d927cf346ed0730b086c2d1

Observation b60e728a-df0f-4060-9aeb-3d455d436d74 · outbound

This paper cites What can machine learning do? workforce implications.

Interpretable Risk Mitigation in LLM Agent Systems What can machine learning do? workforce implications

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.491497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:01.848928Z digest=sha256:57950c0e58b6db6ce14866d9289125dab4ca4bd9d2d021dc70c3e4450486de9a

Observation e55a86e7-747f-4031-8209-67c0bab512ab · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Interpretable Risk Mitigation in LLM Agent Systems Evaluating Large Language Models Trained on Code

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.853101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.853101Z digest=sha256:77662c68cde7f99636c64eb8603be5da825f6a04a0512906944cd024a5b11336

Observation 08b4f22d-3581-40e0-aa27-dd14aaa9041a · outbound

This paper cites Instigating cooperation among llm agents using adaptive information modulation.

Interpretable Risk Mitigation in LLM Agent Systems Instigating cooperation among llm agents using adaptive information modulation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.857296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.857296Z digest=sha256:384956abee230a680f720e33eba093a0ca4853c3e9d3d5711c57f8837af72337

Observation 72433d46-a595-481b-a5ef-2f80a6405e69 · outbound

This paper cites Sparse Autoencoders Find Highly Interpretable Features in Language Models.

Interpretable Risk Mitigation in LLM Agent Systems Sparse Autoencoders Find Highly Interpretable Features in Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.870732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.870732Z digest=sha256:3ac76c3b1f32d99a88b69f017dbb0885221eabb4d120ef0a04df52cb036916fd

Observation eadd38c1-fbcd-4b04-a105-2ea6f1e6a5e6 · outbound

This paper cites Reinforcement learning in a prisoner’s dilemma.

Interpretable Risk Mitigation in LLM Agent Systems Reinforcement learning in a prisoner’s dilemma

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.458432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:01.875550Z digest=sha256:8acf163db1a261faea36a1c422700bd8ae8a1e10104b9046f62534a628d1790e

Observation 24f6a202-78a8-481d-8db8-1c4b79df229d · outbound

This paper cites Toy Models of Superposition.

Interpretable Risk Mitigation in LLM Agent Systems Toy Models of Superposition

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.880133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.880133Z digest=sha256:4cb8f66fc054ef96bc2c1efcec259e89b18cd89d9e8a23e53fca7bdb84d97a6b

Observation 3da5c266-6d9b-45d6-8d90-16eafac12480 · outbound

This paper cites PoGaIN: Poisson-Gaussian Image Noise Modeling from Paired Samples.

Interpretable Risk Mitigation in LLM Agent Systems PoGaIN: Poisson-Gaussian Image Noise Modeling from Paired Samples

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T21:11:02.834041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:01.886176Z digest=sha256:2ff77ffc704fce5ad09a4fdf8725f692a5eee445af56727d2fcbf8f98e1d0add

Observation 9a93bcfe-f4cc-4a52-9e56-a716c5208000 · outbound

This paper cites Not All Language Model Features Are One-Dimensionally Linear.

Interpretable Risk Mitigation in LLM Agent Systems Not All Language Model Features Are One-Dimensionally Linear

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.891128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.891128Z digest=sha256:9862f18305972a5559cb7102c7da13ad63a463642d0955e43bed2d65eddff31b

Observation 80ab46e3-96c8-4341-b15b-4ef484f62fc1 · outbound

This paper cites Some experimental games.

Interpretable Risk Mitigation in LLM Agent Systems Some experimental games

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.442507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:01.896121Z digest=sha256:e2d38eea81298a903a030b1b8a677adf59e42532b5d36a8a40aa0c0f4a82bc97

Observation f9da16f4-a983-41b4-b6d2-cccdb6ba0154 · outbound

This paper cites Nicer Than Humans: How do Large Language Models Behave in the Prisoner's Dilemma?.

Interpretable Risk Mitigation in LLM Agent Systems Nicer Than Humans: How do Large Language Models Behave in the Prisoner's Dilemma?

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.900891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.900891Z digest=sha256:34bfc841769e2ea8cffdc93cd5037bd3e7d3a1819cd5e90f36402351b88eec21

Observation 88637004-c8d5-4677-890d-d755e3357d38 · outbound

This paper cites Artificial intelligence, values, and alignment.

Interpretable Risk Mitigation in LLM Agent Systems Artificial intelligence, values, and alignment

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.906112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.906112Z digest=sha256:983040e93d7f6c545616417c0b9d702ce3083f5df71ead0f0c84e9bf77c8cdea

Observation cb688a1f-ae7d-4316-9fcd-3ddffe2c900b · outbound

This paper cites The Llama 3 Herd of Models.

Interpretable Risk Mitigation in LLM Agent Systems The Llama 3 Herd of Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.910539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.910539Z digest=sha256:4f4a020c613471f028ed0d5c73aef9bd8843d98b77e2084e04371bf59aefa489

Observation 2aea34d7-c236-4ba2-b05c-5266d05cbe3a · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Interpretable Risk Mitigation in LLM Agent Systems Measuring Massive Multitask Language Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.915548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.915548Z digest=sha256:ca89d9ad62d45f55892dca860488aa0fffd7202917750a66eaee99ed0a8991a6

Observation daa69867-514e-4500-ac68-9b455cec0cf3 · outbound

This paper cites Measuring mathematical problem solving with the math dataset,.

Interpretable Risk Mitigation in LLM Agent Systems Measuring mathematical problem solving with the math dataset,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.920013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.920013Z digest=sha256:e88313ea47e867c0e8dcc3491ec1d7ba83d85befd4afad4f3a94d86204e35ae5

Observation ccc257f3-639b-4bd7-b693-d99e2d88ec53 · outbound

This paper cites Cogagent: A visual language model for gui agents.

Interpretable Risk Mitigation in LLM Agent Systems Cogagent: A visual language model for gui agents

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.416984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:01.930482Z digest=sha256:efc28eeb620327c6f16410edca41c7d1eb90ae5a44312613e4f2f877d3736b0c

Observation f465e08f-8581-42e6-b51b-25e319bc463c · outbound

This paper cites Non-linear inference time intervention: Improving llm truthfulness.

Interpretable Risk Mitigation in LLM Agent Systems Non-linear inference time intervention: Improving llm truthfulness

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.400826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:01.935781Z digest=sha256:402b467126f87de133fcb9f80388f52db518c7bb7d2ff7e8e7cf479788bce301

Observation 79c959d9-a3b3-4327-86ed-96acb55e005b · outbound

This paper cites Towards Reasoning in Large Language Models: A Survey.

Interpretable Risk Mitigation in LLM Agent Systems Towards Reasoning in Large Language Models: A Survey

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.939936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.939936Z digest=sha256:b54d6b11ec430ffb49a44fb5a5fe6b23b7db036263d366078f48d7fc3669289f

Observation f8bf0ff6-e450-4f1c-b07c-950516506ed9 · outbound

This paper cites Large language models for uavs: Current state and pathways to the future.

Interpretable Risk Mitigation in LLM Agent Systems Large language models for uavs: Current state and pathways to the future

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.384589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:01.944592Z digest=sha256:61c10f9ea451d328bd719dd0bc1dedda55b57ab0e45f9c175b846ab65de3537f

Observation a7a39f6f-26fb-4c28-b196-bc2a237200ff · outbound

This paper cites Mixtral of Experts.

Interpretable Risk Mitigation in LLM Agent Systems Mixtral of Experts

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.949339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.949339Z digest=sha256:f3961b6469708f54a9f38a64d6f54c6101a0cc2478c61550b91052fa29b5d575

Observation 017ef2a9-5520-4eb9-bdc4-f92641554c31 · outbound

This paper cites llama-3-8b-it-res (revision 53425c3), 2024.

Interpretable Risk Mitigation in LLM Agent Systems llama-3-8b-it-res (revision 53425c3), 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.367736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:01.953994Z digest=sha256:9473ebdcada35584b9a8a304a75ec3b058bf0f4c030330d00cb43b1303a988ea

Observation 91339d58-f620-44d4-8eb7-b59703b0d07a · outbound

This paper cites an unresolved cited work.

Interpretable Risk Mitigation in LLM Agent Systems Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:11:03.350181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:01.958191Z digest=sha256:8caaa48caed25c1d3d698108470495187f5f90243f21659c26fc2192f7af9731

Observation 1a302ee4-fd9f-46ae-8fab-c55bc4736872 · outbound

This paper cites Martin, Hans-Theo Normann, and T.

Interpretable Risk Mitigation in LLM Agent Systems Martin, Hans-Theo Normann, and T

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.333745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:01.962360Z digest=sha256:ecea6b91f3df271e66e828d5e49771af8293442860921d6b5b73e6e774afb197

Observation 71a8d99b-1365-4a5b-ac7b-cce9c7fceaf8 · outbound

This paper cites Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al.

Interpretable Risk Mitigation in LLM Agent Systems Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.318008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:01.966629Z digest=sha256:27523b78f7cc72af1b545b626cc444d42a996379078a78661ff1784c2f95fb75

Observation 53ae654e-f604-43fa-bec1-7e4bbe48d9bf · outbound

This paper cites Inference-Time Intervention: Eliciting Truthful Answers from a Language Model.

Interpretable Risk Mitigation in LLM Agent Systems Inference-Time Intervention: Eliciting Truthful Answers from a Language Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.971070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.971070Z digest=sha256:540aaa7be166be0ee6400856aa5aae37cbde070b839cf51bdb9afe8e158c75b1

Observation accf6756-7ccc-46fa-8ac0-5a95cac9c7e1 · outbound

This paper cites Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2.

Interpretable Risk Mitigation in LLM Agent Systems Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.975880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.975880Z digest=sha256:7f69fa7f7599d1cb91639b7a2b935706470dae5047126bd157967206d6791f08

Observation e1e40780-c265-4c8d-b588-0f4db9b277d5 · outbound

This paper cites The mythos of model interpretability.

Interpretable Risk Mitigation in LLM Agent Systems The mythos of model interpretability

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.302405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:01.980577Z digest=sha256:0c1209c0c799074cd459e7604034fd84088973c7563cc3c89e0d6ee255200342

Observation 4fc9c058-9b7f-4030-b591-6b05a8f26e5f · outbound

This paper cites Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing.

Interpretable Risk Mitigation in LLM Agent Systems Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.989753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.989753Z digest=sha256:e50928381682629a5bf5234d37d8e85d70e0a5021f130fd3d1e2be358df0d157

Observation 3cdffd23-9790-4415-81f6-49a0963613c8 · outbound

This paper cites Large Model Strategic Thinking, Small Model Efficiency: Transferring Theory of Mind in Large Language Models.

Interpretable Risk Mitigation in LLM Agent Systems Large Model Strategic Thinking, Small Model Efficiency: Transferring Theory of Mind in Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.994175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.994175Z digest=sha256:8619f32f602d894c2019139d055a4b5472821605f90a70515565b2d9c3568945

Observation c86296d2-ed87-49d3-b9f2-78d6f82f7afa · outbound

This paper cites Linguistic regularities in continuous space word representations.

Interpretable Risk Mitigation in LLM Agent Systems Linguistic regularities in continuous space word representations

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.259770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:02.003164Z digest=sha256:ddfeeca6c378babae639ff60ee66ecc24f5d318a46890b89141d93f6978eddbf

Observation a59a7951-c2e9-43c9-a557-199380bc4da6 · outbound

This paper cites Large Language Models: A Survey.

Interpretable Risk Mitigation in LLM Agent Systems Large Language Models: A Survey

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.008126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.008126Z digest=sha256:7b42231d93ea5819b09827e0ec2a18a31095bcc063628332b6b16fa23ce15847

Observation fae7cda6-218c-4394-8746-aa9f8ca079ab · outbound

This paper cites A strategy of win-stay, lose-shift that outperforms tit-for-tat in the prisoner’s dilemma game.

Interpretable Risk Mitigation in LLM Agent Systems A strategy of win-stay, lose-shift that outperforms tit-for-tat in the prisoner’s dilemma game

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.013053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.013053Z digest=sha256:6e1299177ddd6bd73911a167c6411749549054eb9a6ff84b0414a6436e075ade

Observation 4826b429-d15b-417a-b168-ff52c95b5cbc · outbound

This paper cites Training language models to follow instructions with human feedback.

Interpretable Risk Mitigation in LLM Agent Systems Training language models to follow instructions with human feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.018166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.018166Z digest=sha256:d3ca1736b8c79976b5dfb62f5e1dffa536d0f4f5701cdd73261a2e2e247a4048

Observation 014ab367-984a-479d-83c0-103ef33837a1 · outbound

This paper cites Cooperation: A systematic review of how to enable agent to circumvent the prisoner’s dilemma.

Interpretable Risk Mitigation in LLM Agent Systems Cooperation: A systematic review of how to enable agent to circumvent the prisoner’s dilemma

Reference 48

Resolution
verified exact
raw_fallback, observed 2026-08-15T21:11:02.613607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:02.023478Z digest=sha256:3780834f2384181413855278e90c61a0ddbdad1b3bb263d80b8cac8c34ad9ebb

Observation 7a05f4db-7afc-4148-a98b-9ef26cb1000d · outbound

This paper cites Generative Agents: Interactive Simulacra of Human Behavior.

Interpretable Risk Mitigation in LLM Agent Systems Generative Agents: Interactive Simulacra of Human Behavior

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.028670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.028670Z digest=sha256:67a4e8a8e1e8c1012e7dd6a9a885a9459eb9860fd181e9859f56be55e86a904d

Observation de77e093-1471-462a-b0a6-a927f6729dc5 · outbound

This paper cites TinyClick: Single-Turn Agent for Empowering GUI Automation.

Interpretable Risk Mitigation in LLM Agent Systems TinyClick: Single-Turn Agent for Empowering GUI Automation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.033414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.033414Z digest=sha256:87d0748dd4b4826d85f3a220f04b4ae23c38a40653a42235c09d68bbca6c4a88

Observation 73a4ef3f-38c0-44df-85f6-445319704ac3 · outbound

This paper cites an unresolved cited work.

Interpretable Risk Mitigation in LLM Agent Systems Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:11:03.231251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:02.038063Z digest=sha256:6218d7fe2a48051a0f21b1a5d9e218a839cac40b72c3daac18f29f810f4c3d23

Observation 35196dd2-37eb-4576-a6c2-b40e2da03ea4 · outbound

This paper cites Effect of private deliberation: Deception of large language models in game play.

Interpretable Risk Mitigation in LLM Agent Systems Effect of private deliberation: Deception of large language models in game play

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.216654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:02.042680Z digest=sha256:f3c8ae4e287582a2152c72943e253c762b28df6cc70b6318faa124830b3afe98

Observation a70cb5aa-ff92-409f-b95a-3405b51bbf27 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Interpretable Risk Mitigation in LLM Agent Systems GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.047108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.047108Z digest=sha256:81ca30f89e06ba506b88e952c0722e526a7ea99aa8450249ab44466b97cad7d6

Observation d5a5b1b9-7b0f-427c-973e-a741a2cd28cd · outbound

This paper cites A primer in BERTology: What we know about how BERT works.

Interpretable Risk Mitigation in LLM Agent Systems A primer in BERTology: What we know about how BERT works

Reference 54

Resolution
malformed identifier
no resolver link, observed 2026-08-15T21:11:02.051703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.051703Z digest=sha256:006e2cae780f9cc2a40ff32a91de815c4674faefd6b38af6d76006dbf2717b7e

Observation cb2fe07f-6a82-4fb7-bc5a-7b4cefa20efd · outbound

This paper cites Research priorities for robust and beneficial artificial intelligence.

Interpretable Risk Mitigation in LLM Agent Systems Research priorities for robust and beneficial artificial intelligence

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.201675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:02.055985Z digest=sha256:e493649ed0533d56a24317b999bbbd22513c3ebf0dea6df3df43441e0c0e914e

Observation 7a2cf6c4-3cb3-42c1-a6c5-e33c3fd83511 · outbound

This paper cites an unresolved cited work.

Interpretable Risk Mitigation in LLM Agent Systems Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:11:03.185336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:02.060734Z digest=sha256:69e4b5cbacabbafc8b5f9f745ccf484d818f742a6509566fb311812b934d0d76

Observation 95101970-bd83-42b3-bfc5-703b311591ce · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

Interpretable Risk Mitigation in LLM Agent Systems Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.064856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.064856Z digest=sha256:1c82626fe8ffe787237d8087459b361f21167a0ebc9c9f1c2c8b5903d93c0f61

Observation 8f6167c9-4a5f-46af-b08b-065532a5c700 · outbound

This paper cites An evolutionary model of personality traits related to cooperative behavior using a large language model.

Interpretable Risk Mitigation in LLM Agent Systems An evolutionary model of personality traits related to cooperative behavior using a large language model

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.170463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:02.069111Z digest=sha256:6e0a88906ddb0a091aaa9e5eddb030d9a0ecf744f7fb9c9a3677f114d6516ce2

Observation 0b622b1d-44e4-43db-bc95-870dd9ca5525 · outbound

This paper cites A comparative analysis of the definitions of autonomous weapons systems.

Interpretable Risk Mitigation in LLM Agent Systems A comparative analysis of the definitions of autonomous weapons systems

Reference 59

Resolution
verified exact
doi, observed 2026-08-15T21:11:02.190554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:02.073714Z digest=sha256:d833be855788077e015d89f9581ca0e9b4cbe4f7a5704105d02fc7dd61ccfdc1

Observation 54cdc25e-dcbe-4a06-a3ca-a0b083ec66e7 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Interpretable Risk Mitigation in LLM Agent Systems Gemma: Open Models Based on Gemini Research and Technology

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.078502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.078502Z digest=sha256:9f60a819bb9fd367a12282593c43f968190bf49f9a9c04ff113d0a77609d9ca2

Observation e70192c0-2fa3-4d26-9e99-25bce4521ab8 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Interpretable Risk Mitigation in LLM Agent Systems Gemma 2: Improving Open Language Models at a Practical Size

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.082737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.082737Z digest=sha256:7807532a932a7cc5508ea5ae1ba34c2e056a06dd5becdd7dfd8754ea1e4b4215

Observation 625f59c6-0eef-406d-9cf0-d8c92c8ef449 · outbound

This paper cites Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet.

Interpretable Risk Mitigation in LLM Agent Systems Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.087316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.087316Z digest=sha256:72b7baf1e88d10f65094ce290aa1606e08ce4ba50e4d97e505c37765c70bdf9a

Observation 7b2ebf77-d4d1-47eb-846b-c88fd4d56a1d · outbound

This paper cites Moral alignment for llm agents.

Interpretable Risk Mitigation in LLM Agent Systems Moral alignment for llm agents

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.143335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:02.092060Z digest=sha256:3bcdb4ec670f4dd05a5bd65c15c5372589d770a18d06998469b99f1eb2e401b0

Observation f523afad-184f-41ef-9a0e-57e19a11c506 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Interpretable Risk Mitigation in LLM Agent Systems LLaMA: Open and Efficient Foundation Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.098429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.098429Z digest=sha256:b7821d9533c05a3aa20b7472f1625d3a36eb56a4d7127254bc330a7c87a03d30

Observation 82ac90b8-2d66-452e-9693-66c10e7d314b · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Interpretable Risk Mitigation in LLM Agent Systems Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.104902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.104902Z digest=sha256:86179a53f98c3512617377b09fa1a9b9ba4a95341ef92292e16ca626b24b658a

Observation 20ce1bc9-e4d8-4904-9888-9de2d34c5479 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

Interpretable Risk Mitigation in LLM Agent Systems GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.109561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.109561Z digest=sha256:d6b040393bd26c5c580f25447e3254f735faac25d04fa74d00d53d1ea69daa1f

Observation f2942ef8-188f-4e1f-a44b-0747ba8e3471 · outbound

This paper cites Mobile-agent: Autonomous multi-modal mobile device agent with visual perception,.

Interpretable Risk Mitigation in LLM Agent Systems Mobile-agent: Autonomous multi-modal mobile device agent with visual perception,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.127797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:02.114293Z digest=sha256:06c79b5e0d614cf52fd85bab707e9ec3e7a585f35cc3b0da75134b4be29d3f29

Observation e8c5f846-e91b-401e-80a4-0fd5d73e1194 · outbound

This paper cites Dai, and Quoc V Le.

Interpretable Risk Mitigation in LLM Agent Systems Dai, and Quoc V Le

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.112312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:02.123585Z digest=sha256:d8509df53d1f4236133a4031fdf70d61343ef0083477c45dcbf53bf4f756f540

Observation 9aac0e58-4508-4dce-b103-c860ac6c28f4 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Interpretable Risk Mitigation in LLM Agent Systems Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.128289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.128289Z digest=sha256:29a4e5d2a2b267183581e760cbdc1d1b232c9d516023a45d2a19319e42aced4e

Observation 1cea8ac7-27e1-49cf-b37d-d6744ff11b68 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Interpretable Risk Mitigation in LLM Agent Systems ReAct: Synergizing Reasoning and Acting in Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.133411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.133411Z digest=sha256:2bef79b671012492989898c47c8a9d5c0832273abe1284313507060b6a571c47

Observation 647ac5ed-565f-4692-8bee-b21bc1834a89 · outbound

This paper cites AppAgent: Multimodal Agents as Smartphone Users.

Interpretable Risk Mitigation in LLM Agent Systems AppAgent: Multimodal Agents as Smartphone Users

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.140648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.140648Z digest=sha256:feb88b654f48e33a5fddd8ae28a876d2dfee978fe9305dcf100337cac4d3f667

Observation 33d30244-b423-4420-bcc9-fbd665b74b6d · outbound

This paper cites Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception.

Interpretable Risk Mitigation in LLM Agent Systems Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.118923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.118923Z digest=sha256:a43ce8ed4c268c8ea3fa5c19a7998fa291d56682761a4fdf8c494da7b6547b6f

Observation 10bb76ee-bc2d-40f4-babf-40043502314b · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Interpretable Risk Mitigation in LLM Agent Systems Representation Engineering: A Top-Down Approach to AI Transparency

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.149478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.149478Z digest=sha256:9ff2d37cec87ebd19e0a364d113af81199c20f141dfff9e64bc1c8a93e012b14

Observation e1633c79-0375-4295-bfa3-8325c644a49a · outbound

This paper cites You Only Look at Screens: Multimodal Chain-of-Action Agents.

Interpretable Risk Mitigation in LLM Agent Systems You Only Look at Screens: Multimodal Chain-of-Action Agents

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.145319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.145319Z digest=sha256:ad19f399c83efb25f21855473dd357ef2c409a08f0ca7ea094db02abb627ca92

Observation d843c580-4512-4d8f-a3a6-6b41724a5ec6 · outbound

This paper cites an unresolved cited work.

Interpretable Risk Mitigation in LLM Agent Systems Unresolved cited work

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.984893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.984893Z digest=sha256:41fa5b4e7826ac58d85714833effb4314cb01badef655e5ca8598e6da8fc891c

Observation 671ed4ec-d1ad-4d99-97f8-e0d15322b52b · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Interpretable Risk Mitigation in LLM Agent Systems Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.925637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.925637Z digest=sha256:ca86bfd367c3dfc46660e5d802871807d97a129fd1b19eb86033b144745af84d

Observation 6b1e14ac-f124-489b-84f4-00a88d7ea70d · outbound

This paper cites an unresolved cited work.

Interpretable Risk Mitigation in LLM Agent Systems Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:11:03.474249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:01.866455Z digest=sha256:0324c9ba4de7bbaf0ec938b0de15eecd97f93c95baa7e287a0bec5f1c5acb5e3

Observation 98fbd528-7341-4dfe-92d5-440307b35277 · outbound

This paper cites an unresolved cited work.

Interpretable Risk Mitigation in LLM Agent Systems Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:11:03.276283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:11:01.998638Z digest=sha256:76a431b31bbef3a5a4d5e74ea493b1ba6ec61717d2d97e8903a9db3e02bd2899

Pith citing papers

Observation b72a9458-11dc-47d1-84f0-f77d9122b3d6 · inbound

A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy cites this paper.

A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy Interpretable Risk Mitigation in LLM Agent Systems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:51:56.450815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:51:56.450815Z digest=sha256:a93e67a12a362739b2b20ee468e87faae1833be6420e3848c7e817965214e743

Observation 1b7081ac-e99d-43fc-be2e-190be39eed65 · inbound

LLM Harms: A Taxonomy and Discussion cites this paper.

LLM Harms: A Taxonomy and Discussion Interpretable Risk Mitigation in LLM Agent Systems

Reference 189

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:31:24.636618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T00:29:07.951709Z digest=sha256:111a84462c2c5ee00668685309974afdc67514dce599ba88d75b255d480743a0

Observation 84d058c8-51a1-4859-970e-8b0c97117aac · inbound

LLM Harms: A Taxonomy and Discussion cites this paper.

LLM Harms: A Taxonomy and Discussion Interpretable Risk Mitigation in LLM Agent Systems

Reference 189

Resolution
unresolved
no resolver link, observed 2026-08-03T18:19:27.337242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:19:27.337242Z digest=sha256:59c8bc05f948bfa1034545d25cd99991d64c3eef1284700fb87ab252c48a2280