Pith. sign in

Paper Citation Record · LEDGER

Effective Red-Teaming of Policy-Adherent Agents

As of 7 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 2 inbound Pith citation observations for arXiv:2506.09600.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09600 v3

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:50:00.253920Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T07:48:36.924294Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T07:54:22.258152Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact2
  • verified fuzzy2
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a9fe4ffe-27a7-4b2d-bac8-777d4d8dd59c · outbound

This paper cites an unresolved cited work.

Effective Red-Teaming of Policy-Adherent Agents Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.119981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.119981Z digest=sha256:cac8d2389d253bb0241b1890519ec9e99c23c0b35fa21d65bdd9b46b44fc91fa

Observation ed34efe0-0318-4a5e-a5e3-dc1c3f58dbd8 · outbound

This paper cites AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents.

Effective Red-Teaming of Policy-Adherent Agents AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.124363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.124363Z digest=sha256:78eb47897ba466f0c74d1f152edfd1c3fb165910b6a21043b5a15147c0c981dd

Observation 5bd98dfd-fade-4c92-993d-2104cf826649 · outbound

This paper cites A Framework for Testing and Adapting REST APIs as LLM Tools.

Effective Red-Teaming of Policy-Adherent Agents A Framework for Testing and Adapting REST APIs as LLM Tools

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.129367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.129367Z digest=sha256:227ccd11d41f6b6074f6b450db22e35c0bd553221d71fa0785b72d89939f9fc8

Observation 03ed9a8b-08c1-47cf-a062-174d7e1d1cab · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Effective Red-Teaming of Policy-Adherent Agents Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.133366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.133366Z digest=sha256:9eb7fd068a8ede08645b2254b9b8286cb1f2ebfdb77ca48d9fccedf1d025ab99

Observation a98cffe2-f85d-496d-9947-cd94c9849bc6 · outbound

This paper cites The Llama 3 Herd of Models.

Effective Red-Teaming of Policy-Adherent Agents The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.137097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.137097Z digest=sha256:fa0b67c8c54836704413f8b311cced0feef5756940d53eccfdb4e0bd17520fd1

Observation 1e9eb559-74f2-464a-8fe8-116158765a52 · outbound

This paper cites an unresolved cited work.

Effective Red-Teaming of Policy-Adherent Agents Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:50:00.987004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:50:00.140744Z digest=sha256:0e9ad44a628000220b9aa71975a1df85c46e2fc0fdb9dca20234f2b8d7b166a9

Observation e6252988-2705-441d-be6b-a23da52a6621 · outbound

This paper cites an unresolved cited work.

Effective Red-Teaming of Policy-Adherent Agents Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:50:00.976219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:50:00.144239Z digest=sha256:7e959b9ef10258f28e25b862294e5c9606a9f8e7288c05842ea22adce4447abd

Observation 2d02d8c2-ef73-4a6e-9b5b-998a49c59faf · outbound

This paper cites GPT-4o System Card.

Effective Red-Teaming of Policy-Adherent Agents GPT-4o System Card

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.147390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.147390Z digest=sha256:8057a1a31137a8389d2a0adfe3f9165be02092d4ea82501582144db5944e871e

Observation 0fc1f222-66e5-48a6-9d06-cccaa70e343d · outbound

This paper cites an unresolved cited work.

Effective Red-Teaming of Policy-Adherent Agents Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.150733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.150733Z digest=sha256:b39a3a76d245cf45029d7de04135b7e1f370749b9ea062ae27b16f390764e6db

Observation 48730d43-8fe6-4f01-a378-2c588ee6b58a · outbound

This paper cites an unresolved cited work.

Effective Red-Teaming of Policy-Adherent Agents Unresolved cited work

Reference 10

Resolution
verified exact
raw_fallback, observed 2026-08-07T04:50:00.672247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:50:00.154869Z digest=sha256:db4d15d62c781b4a7421d5ed1c6cf8f800fbd67b0ea2852b3db948a92409328a

Observation 986a60bd-ecac-46ec-ba2d-f5e102ed3cfc · outbound

This paper cites an unresolved cited work.

Effective Red-Teaming of Policy-Adherent Agents Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:50:00.965469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:50:00.158262Z digest=sha256:01e53f01d573ac95a712f3591db58a784869ac340b0151a1b8066c4642edba93

Observation 7ceccfad-f25e-486e-adc7-2455580e1072 · outbound

This paper cites Unveiling Safety Vulnerabilities of Large Language Models.

Effective Red-Teaming of Policy-Adherent Agents Unveiling Safety Vulnerabilities of Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.161709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.161709Z digest=sha256:29c634b5738a173101dd8e06a6ab7012c3cfd74e628cea27487a80d55a4110c5

Observation cc16bd91-aabd-4430-a712-cd7800f2b6f4 · outbound

This paper cites an unresolved cited work.

Effective Red-Teaming of Policy-Adherent Agents Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:50:00.953787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:50:00.165546Z digest=sha256:708a9a0611210053894fac6116e89cfc61e0e0406fca71fa5174e97815c4e3f5

Observation 6c005901-1f7d-42ed-895b-fca8e3ed9781 · outbound

This paper cites ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents.

Effective Red-Teaming of Policy-Adherent Agents ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.168901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.168901Z digest=sha256:456029d785b187ba9e98a93121eae2434021e47db252fc1b6213be73250bc0f1

Observation d2052270-87ab-4262-b1ae-ae4e9432038d · outbound

This paper cites SOPBench: Evaluating Language Agents at Following Standard Operating Procedures and Constraints.

Effective Red-Teaming of Policy-Adherent Agents SOPBench: Evaluating Language Agents at Following Standard Operating Procedures and Constraints

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.172498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.172498Z digest=sha256:a864d72a666e93ce090c64f437cb304558201438905650ad98a991e6ef1ccd58

Observation 1fc9cdf3-3095-4c91-9b47-b3ec89dd02ca · outbound

This paper cites DeepSeek-V3 Technical Report.

Effective Red-Teaming of Policy-Adherent Agents DeepSeek-V3 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.176032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.176032Z digest=sha256:971beb6f6036b29d9132b8ec8c06bc1011ca64725e11fb4338355b1d85af91c3

Observation 45a07ed6-ae4e-4bc8-8523-8826db4d0170 · outbound

This paper cites From LLM to Conversational Agent: A Memory Enhanced Architecture with Fine-Tuning of Large Language Models.

Effective Red-Teaming of Policy-Adherent Agents From LLM to Conversational Agent: A Memory Enhanced Architecture with Fine-Tuning of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.179402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.179402Z digest=sha256:2e9ce131ebdb679165d585f2b10f7d59cd23c5aa73f3539881139a5483c93951

Observation e4aa5648-db3b-4be7-ade6-9fd1c50de94e · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Effective Red-Teaming of Policy-Adherent Agents AgentBench: Evaluating LLMs as Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.182946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.182946Z digest=sha256:c8cf12c18f235b4b2bb1463523c5f0ce84155aebb43aabf6611d23a3513c4804

Observation 50533e89-9e21-42b8-844c-5106ca7e14f1 · outbound

This paper cites Prompt Injection attack against LLM-integrated Applications.

Effective Red-Teaming of Policy-Adherent Agents Prompt Injection attack against LLM-integrated Applications

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.186499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.186499Z digest=sha256:497d1672c22b4c626f901364f09309c46de1013899e08519846fa372ada61d35

Observation a358ba86-58cc-4329-87c5-f8f28fee284d · outbound

This paper cites an unresolved cited work.

Effective Red-Teaming of Policy-Adherent Agents Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:50:00.943574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:50:00.190116Z digest=sha256:42b650037de9887823d0242440be87cdedf80a5cbcb15cae8cd199b4c502cb54

Observation 0fde2ab8-b1f9-4dd3-afcb-8dacea8d9941 · outbound

This paper cites u ndler, Mark Niklas M \.

Effective Red-Teaming of Policy-Adherent Agents u ndler, Mark Niklas M \

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:50:00.933475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:50:00.193419Z digest=sha256:c347ee516c6141c7643a71a25f864123ea8d4c48a9d708fddf4ce94add3829df

Observation 5b29a17e-92b7-409b-b739-9d0cd5e120c8 · outbound

This paper cites an unresolved cited work.

Effective Red-Teaming of Policy-Adherent Agents Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:50:00.923174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:50:00.196714Z digest=sha256:d895c817cd212536031692ac29ca0659c1a5d6393f327dc78057452de924c44e

Observation e7683656-a1bd-4537-aaef-7b477760998e · outbound

This paper cites do anything now.

Effective Red-Teaming of Policy-Adherent Agents do anything now

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:50:00.912398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:50:00.199724Z digest=sha256:bf33619706c9b272de0d24e6d1141468c68a7ccc708c38be31e4c0cf541828ec

Observation 36f055a0-5e8a-40fb-9d6b-71cf64ba9975 · outbound

This paper cites do anything now.

Effective Red-Teaming of Policy-Adherent Agents do anything now

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.202923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.202923Z digest=sha256:77085d96644f5c50b665a1a3b3703a7facc626e77edbc68dff796e5b3e1d7ced

Observation f46532e3-00d6-4dd6-99fb-2a643e881f52 · outbound

This paper cites CHOPS: CHat with custOmer Profile Systems for Customer Service with LLMs.

Effective Red-Teaming of Policy-Adherent Agents CHOPS: CHat with custOmer Profile Systems for Customer Service with LLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.206194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.206194Z digest=sha256:bdba0a2e2b116b6da8c47a41ebf2024c7f9ba1fca3ec7c1fa3bce90f003b95bb

Observation ada563b1-aadb-4908-9872-59d56c02826d · outbound

This paper cites Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers.

Effective Red-Teaming of Policy-Adherent Agents Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.209867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.209867Z digest=sha256:b5db811acc824ecf5c6d88c21648d2ed83313287eefeb5c6e74b2b3bc9857cf5

Observation 04ecd511-17b3-4848-8819-74ac326f158f · outbound

This paper cites an unresolved cited work.

Effective Red-Teaming of Policy-Adherent Agents Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.214612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.214612Z digest=sha256:26e2ee76e23272d17483cf488396d081d179e9222f6e71ceb63a114beb85c843

Observation c7f4545a-55a6-4feb-8926-256a2478f566 · outbound

This paper cites Multi-Turn Context Jailbreak Attack on Large Language Models From First Principles.

Effective Red-Teaming of Policy-Adherent Agents Multi-Turn Context Jailbreak Attack on Large Language Models From First Principles

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.218144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.218144Z digest=sha256:5109d7f6a7ae67bad1be297b3b0d7a165592b67b5c259e0925e40a125f5461a4

Observation 0066ba6d-3785-41fd-96c2-443202205f55 · outbound

This paper cites an unresolved cited work.

Effective Red-Teaming of Policy-Adherent Agents Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:50:00.895130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:50:00.221806Z digest=sha256:9ce8a4cf4d57a3511ee8857f6d3d2863b406d488666f349fa7826f6b436ef5e7

Observation 0c447225-0981-48ad-aafa-07b43f553d88 · outbound

This paper cites Emotional Manipulation Through Prompt Engineering Amplifies Disinformation Generation in AI Large Language Models.

Effective Red-Teaming of Policy-Adherent Agents Emotional Manipulation Through Prompt Engineering Amplifies Disinformation Generation in AI Large Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:50:00.289354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:50:00.225273Z digest=sha256:465689745406585ce58798e18523d4573a71cdde2bdc357eb480a72dba0c4414

Observation 617a1164-8ca5-4f15-bafa-2f5dfbfabcb7 · outbound

This paper cites Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise.

Effective Red-Teaming of Policy-Adherent Agents Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.228964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.228964Z digest=sha256:89e528dc2dd9e4c6fe1f54228487c0f257fec602610ee3dd3549c5a1f581cbbe

Observation 5f6b1128-03a9-4a3a-885e-aefa408de241 · outbound

This paper cites Qwen3 Technical Report.

Effective Red-Teaming of Policy-Adherent Agents Qwen3 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.232482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.232482Z digest=sha256:397b476f0c2c2595b4b471e9bfa8119049ffd7bbe9017b2b197eaacc23e9de7e

Observation f2aa3b30-54c0-4f1e-8b3f-35ebd1d96784 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Effective Red-Teaming of Policy-Adherent Agents $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.235971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.235971Z digest=sha256:043444d50972125c4a95e2cee2e08111bade26c2dd1a54438877a83dfd24700c

Observation bdddd043-26b7-436e-9c0a-a2d6f4633c16 · outbound

This paper cites an unresolved cited work.

Effective Red-Teaming of Policy-Adherent Agents Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.239482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.239482Z digest=sha256:4929b35f8d7953cd1291f572d37e5b1e768be377bb0e74d68d37457497b7e0e3

Observation a67f23d5-4b4a-419f-8f27-be2a1c2ced3d · outbound

This paper cites Survey on Evaluation of LLM-based Agents.

Effective Red-Teaming of Policy-Adherent Agents Survey on Evaluation of LLM-based Agents

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.243042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.243042Z digest=sha256:34eac36d711d6d2d9bbdf6292fbb242c851f712f534ce760bcc0c6e7302c67e6

Observation 816d0e1b-59bd-4e90-9fb5-a609d7754cc9 · outbound

This paper cites an unresolved cited work.

Effective Red-Teaming of Policy-Adherent Agents Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:50:00.877545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:50:00.246609Z digest=sha256:dd2b8a7b21a40f52f2903bea69419d856965c90b3ff59ffcf07c4c48b8bb18a4

Observation 80feea2e-0a2a-4cb2-a8a6-236e3e9bf407 · outbound

This paper cites online" 'onlinestring :=.

Effective Red-Teaming of Policy-Adherent Agents online" 'onlinestring :=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.249996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.249996Z digest=sha256:1dfa6cbcd2c3f14cddcfe362529354ea0181782e529271fff208d96bce47f256

Observation 82e72879-d510-4f14-9dad-a2a6be99f225 · outbound

This paper cites write newline.

Effective Red-Teaming of Policy-Adherent Agents write newline

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.253920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.253920Z digest=sha256:32d787efebc074092081b6f2012b4fb9814c6971742cdffba659d06acfcb1dfd

Pith citing papers

Observation 29492dc8-2762-4be4-a9ef-17672f85925e · inbound

Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation cites this paper.

Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation Effective Red-Teaming of Policy-Adherent Agents

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:06:13.730575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T10:04:39.223895Z digest=sha256:bbdbdb105577ba54893e7cdc8904745b52a5634ee7a484b829b0f90133101249

Observation e24b196a-a0fd-4f1c-8ca4-5f9ae0e3084f · inbound

PolicyGuard: A Dialogue-Grounded Sub-Agent Verifier for Policy Adherence in LLM Agents cites this paper.

PolicyGuard: A Dialogue-Grounded Sub-Agent Verifier for Policy Adherence in LLM Agents Effective Red-Teaming of Policy-Adherent Agents

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:54:22.260097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T07:48:36.924294Z digest=sha256:083559777c25f52f046dfdf07b48582cfba86bcca019c2c66e118bb1c5a86758