Pith. sign in

Paper Citation Record · LEDGER

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation

As of 7 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2608.03166.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03166 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:29:58.290673Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact2
  • verified fuzzy3
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch8

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b9fac2ea-4937-4cc2-8d6d-b81eafb9675c · outbound

This paper cites The Oscars of AI Theater: A Survey on Role-Playing with Language Models.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation The Oscars of AI Theater: A Survey on Role-Playing with Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:55.891371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:55.891371Z digest=sha256:477d0bd5af6e5ce66d6092fb8720435ed47db463dc8d911d605031ec3574f759

Observation 00c93cdc-df59-47ba-bbb0-569dfd3da8a1 · outbound

This paper cites ASSERT: Automated Safety Scenario Red Teaming for Evaluating the Robustness of Large Language Models.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation ASSERT: Automated Safety Scenario Red Teaming for Evaluating the Robustness of Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:55.943379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:55.943379Z digest=sha256:25b52345e8286ced5490332dd02c26b857e1b110db8843d81aeb373101cb8c30

Observation 103f45a2-66bb-4496-98f7-7f48abc91fd2 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:56.028755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:56.028755Z digest=sha256:c9596bbf09dc425d37f45eaaa14583bdd78362048818531ee19413991078963d

Observation 59846c01-4ae4-47c3-9ac5-60a8dff937ac · outbound

This paper cites RoleMRC: A Fine-Grained Composite Benchmark for Role-Playing Language Agents,.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation RoleMRC: A Fine-Grained Composite Benchmark for Role-Playing Language Agents,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:30:00.431231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:29:56.129225Z digest=sha256:ce5269516e721c1c4221acf0fbfdbe6511fe6f2e6737148c217de42311028387

Observation 153fcb74-d969-4170-8d67-9e08ce1803b1 · outbound

This paper cites AART: AI-Assisted Red-Teaming with Diverse Data Generation for New LLM-powered Applications.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation AART: AI-Assisted Red-Teaming with Diverse Data Generation for New LLM-powered Applications

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:56.208408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:56.208408Z digest=sha256:f1e4aae664486632edf42b2556da562a3c6d9d0a2363443a7d8e32afd2174114

Observation 8416447f-0537-483c-91d7-93dbd9f4b2fc · outbound

This paper cites Adversarial Testing in LLMs: Insights into Decision-Making Vulnerabilities.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Adversarial Testing in LLMs: Insights into Decision-Making Vulnerabilities

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-06T00:29:59.822893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:29:56.341529Z digest=sha256:14b543d44ed29499acf6c3537af9132a124fdb61339ed63e5612ebe70f188431

Observation 4dded492-67ed-4285-83f1-5d72175ab62e · outbound

This paper cites Red Teaming Large Language Models: A Comprehensive Review and Critical Analysis,.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Red Teaming Large Language Models: A Comprehensive Review and Critical Analysis,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:30:00.272974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:29:56.408327Z digest=sha256:fdcc671f669bd13358ce7157b9f735b6ab3fcb4c9467c56ad997a1b0f06cffb1

Observation aa86cdb8-632d-499e-98c9-a29c2ba83f91 · outbound

This paper cites Security of LLM-based Agents: Attacks, Defenses, and Applications,.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Security of LLM-based Agents: Attacks, Defenses, and Applications,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:30:00.131587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:29:56.485119Z digest=sha256:44c42107ddcb3fe5615f059e6546520846e95e04cbdcbaf9f3e82409de68cfee

Observation 4dc2bad5-597b-4a54-b856-a8b2916e6f24 · outbound

This paper cites Markov-Enhanced Clustering for Long Document Summarization: Tackling the 'Lost in the Middle' Challenge with Large Language Models.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Markov-Enhanced Clustering for Long Document Summarization: Tackling the 'Lost in the Middle' Challenge with Large Language Models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T00:29:59.970446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:29:56.532477Z digest=sha256:f761e1d906e7190e68189ba8bd95285f384233cecbed2d60df4f8bef347403bb

Observation 51ee87da-f0b5-42b3-9169-f8cfd2021453 · outbound

This paper cites Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:56.632621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:56.632621Z digest=sha256:eaeb6cca7126825177487d7c35e4bfae9ffaa4d78e1d5ef1f8f295007bd1d1cc

Observation 593966c0-7a4e-4ad2-93af-ea9b594737a8 · outbound

This paper cites RedDebate: Safer Responses Through Multi-Agent Red Teaming Debates.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation RedDebate: Safer Responses Through Multi-Agent Red Teaming Debates

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:56.700795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:56.700795Z digest=sha256:cc045bf60f8c04d10b3b6b4fd35e52d2e58e3918e1b22ef9c36bbcee0033971a

Observation edb960e9-8d02-478f-ab19-7cf4cadd7b88 · outbound

This paper cites MART: Improving LLM Safety with Multi-round Automatic Red-Teaming.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation MART: Improving LLM Safety with Multi-round Automatic Red-Teaming

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:56.778717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:56.778717Z digest=sha256:9f97b5d5421b2820c4ece1aa586aae65163d4660f4d5d3b653d86d7ae36f6e72

Observation fcd9b88f-f8d7-48cb-8a23-07b0bcc1fd3d · outbound

This paper cites Evil Geniuses: Delving into the Safety of LLM-based Agents.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Evil Geniuses: Delving into the Safety of LLM-based Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:56.858992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:56.858992Z digest=sha256:64c39f9dab659a7dd5d6a3067c6acc214fe5ded1ea9c9305c25cbf94b21f57f2

Observation 2a562869-30fa-40f6-a310-093383e3a0f5 · outbound

This paper cites RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:56.930215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:56.930215Z digest=sha256:a95f2cebd5b480d0da33ee8e71107ccaa93e4c73abf068194b6494a601106c2e

Observation d28070b9-00f4-4e73-b1d8-8654f92726e6 · outbound

This paper cites Encounter-based model of a run-and-tumble particle with stochastic resetting.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Encounter-based model of a run-and-tumble particle with stochastic resetting

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T00:29:59.614995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:29:56.998828Z digest=sha256:28176113375fce6bf5211117b8c2540ec74ece010f08eea3b77ea1f7561c7fc8

Observation be359958-702e-4cb7-a28d-851c99af979f · outbound

This paper cites Exposing Weak Links in Multi-Agent Systems under Adver- sarial Prompting,.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Exposing Weak Links in Multi-Agent Systems under Adver- sarial Prompting,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:57.065596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:57.065596Z digest=sha256:cb62c99a4fc8704f235f658cc370381b5d6eea3c31be721a17dc556cdb5bf032

Observation 6913ec36-c166-4117-bbd4-745fc6aace07 · outbound

This paper cites Iwasawa module of the cyclotomic $\mathbb{Z}_{2}$-extension of certain real quadratic fields.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Iwasawa module of the cyclotomic $\mathbb{Z}_{2}$-extension of certain real quadratic fields

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T00:29:59.270567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:29:57.300906Z digest=sha256:f859960e127cad75b8d43ed6f6c9c511eaee7aa2ac313b5329a6d5003e598e83

Observation c6ca3504-9d98-4ad8-8114-0bcf1183b914 · outbound

This paper cites SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:57.375647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:57.375647Z digest=sha256:b307e848c2e8612a895fb0aa81f930710bf1380bec87fe247d7c2b25313e1e6a

Observation a3dc3093-1dfc-4dff-8728-9dbef0276e2e · outbound

This paper cites From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agent Workflows,.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agent Workflows,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:57.468919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:57.468919Z digest=sha256:f10c60c67f6c81cc23f1b6b01f4838e24f01be0321c94a1bc7202e1ad2e99034

Observation 891d8d25-88bf-48f7-bdc6-ff1dad7283dc · outbound

This paper cites Stable Diffusion For Aerial Object Detection.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Stable Diffusion For Aerial Object Detection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:57.573900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:57.573900Z digest=sha256:b03f35cd00b9b3db87220be864c204a39002074697990a70d6bc2e3959f32723

Observation dc053066-a483-4ae5-860e-62584c6b65a2 · outbound

This paper cites Online Adaptive Traversability Estimation through Interaction for Unstructured, Densely Vegetated Environments.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Online Adaptive Traversability Estimation through Interaction for Unstructured, Densely Vegetated Environments

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T00:29:59.083036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:29:57.693775Z digest=sha256:d6e5670b682be009b07b0c20875c3154d0054c1876c472f4c2c744197a0259f8

Observation 3974ba24-355a-435e-8cb6-3c8a78a87e94 · outbound

This paper cites Learning to Adapt to Position Bias in Vision Transformer Classifiers.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Learning to Adapt to Position Bias in Vision Transformer Classifiers

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T00:29:58.912857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:29:57.773801Z digest=sha256:1b26caf9508e86a1e55cd7ff310d5935fb43fc48b5c2eac11d2148ab7f70c5ae

Observation ca0e29d1-e053-4708-b07e-129e648b92ac · outbound

This paper cites Compound Expression Recognition via Large Vision-Language Models.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Compound Expression Recognition via Large Vision-Language Models

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T00:29:58.788800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:29:58.020373Z digest=sha256:dc8bb62e8163a748a1453a8029a4ce763c086da630318ad12b87fb7bedeb2c9a

Observation b455f210-342a-402c-a79c-e6bb5e790e4b · outbound

This paper cites Context-Aware Two-Step Training Scheme for Domain Invariant Speech Separation.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Context-Aware Two-Step Training Scheme for Domain Invariant Speech Separation

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-06T00:29:58.622840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:29:58.108401Z digest=sha256:680aeb4444059fc1659ed339bdcb39e92f8eabfca344e2f400f5e437b0eb18f1

Observation 5a58fb64-3080-4b1b-b706-afef034d9e26 · outbound

This paper cites The Dark Side of Human Feedback: Poisoning Large Language Models via User Inputs.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation The Dark Side of Human Feedback: Poisoning Large Language Models via User Inputs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:58.203660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:58.203660Z digest=sha256:5e7357418f414076cac4db68d5827bc160cf0ffea2067b6bec744f402f54e755

Observation d7788856-23e0-4090-9d8f-4f6cd1ed82a9 · outbound

This paper cites Lower bounds on the $\ell$-rank of ideal class groups.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Lower bounds on the $\ell$-rank of ideal class groups

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T00:29:58.420619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:29:58.290673Z digest=sha256:566a414e4c56535589521740aa8af940c4d2f34adb46916583047102112edafd

Observation 58b5660d-6260-46e1-908c-065a195af10c · outbound

This paper cites Mitigating Label Noise on Graph via Topological Sample Selection.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Mitigating Label Noise on Graph via Topological Sample Selection

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:57.946969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:57.946969Z digest=sha256:f0e5948903b57b1f275ad238b6a5047080222a3a8fe1968187f551e95bbdca06

Observation d839405d-1902-4356-b701-0d250ea816f7 · outbound

This paper cites Evaluating the Impact of Verbal Multiword Expressions on Machine Translation.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Evaluating the Impact of Verbal Multiword Expressions on Machine Translation

Reference 2025

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T00:29:59.419281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:29:57.241797Z digest=sha256:45f4742cda8e2ab1974d8eb92e4cc814b783c0debe816d79ce71750f6a24d260

Pith citing papers

No inbound Pith citation observations are available.