Pith. sign in

Paper Citation Record · LEDGER

Role-Playing Evaluation for Large Language Models

As of 16 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 3 inbound Pith citation observations for arXiv:2505.13157.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13157 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:21:57.722153Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T04:55:29.796891Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation c0056d57-f91d-4bef-ac38-d721e4555b6f · outbound

This paper cites Journal of management development23(4), 355–371 (2004).

Role-Playing Evaluation for Large Language Models Journal of management development23(4), 355–371 (2004)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:58.159303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:57.609256Z digest=sha256:bf4516cdf2ab28e5f0fb1076f0e46098ec3bf9193ce36ba9330fef8bb44e46ff

Observation e27628fd-fbe8-4e6d-aef2-7955b3382dab · outbound

This paper cites Language Models as Agent Models.

Role-Playing Evaluation for Large Language Models Language Models as Agent Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:57.614797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:57.614797Z digest=sha256:5a4b924e7d04e9c4ccbe55255d17553d476a666c3ec245605705bde0e28e9343

Observation 4811d34f-3669-4e90-b209-9ec5b9dff87e · outbound

This paper cites Compress to Impress: Unleashing the Potential of Compressive Memory in Real-World Long-Term Conversations.

Role-Playing Evaluation for Large Language Models Compress to Impress: Unleashing the Potential of Compressive Memory in Real-World Long-Term Conversations

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:57.620224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:57.620224Z digest=sha256:51f525e58ca5690e1078e080cbb973f64865ea6f17902214a45c8161c1d37926

Observation 0134bc38-25c2-45e5-b988-8e8e9ee81249 · outbound

This paper cites an unresolved cited work.

Role-Playing Evaluation for Large Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:58.141901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:57.625561Z digest=sha256:0675c683e75c3dc28ed92411507f80471c5435e2ba43e4c87f694f35056a1bc8

Observation aa404674-b56a-4d1b-ae7c-5c271c2c7c2c · outbound

This paper cites The Llama 3 Herd of Models.

Role-Playing Evaluation for Large Language Models The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:57.630336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:57.630336Z digest=sha256:5e6e995b68293db4b6b41409476e4472120eb0b7823ee92b9cc588d1e7c57ff9

Observation 7e4f65e6-8fa0-4743-8135-3ea916417110 · outbound

This paper cites GPT-4o System Card.

Role-Playing Evaluation for Large Language Models GPT-4o System Card

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:57.634917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:57.634917Z digest=sha256:6c2fee3a80bd0c48be0d4da382946d95bcf3ed7f580e329af21334cb868327e0

Observation 94367a52-c728-43bf-ad2e-4cc39670d66f · outbound

This paper cites an unresolved cited work.

Role-Playing Evaluation for Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:58.125510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:57.640472Z digest=sha256:47e2230a5422d4a6099e654d0cca0d3d20c82533d6c5a9bf60af65f679a0671a

Observation 7581cbf8-462c-409e-ae0a-bfca96ede448 · outbound

This paper cites Self-Prompt Tuning: Enable Autonomous Role-Playing in LLMs.

Role-Playing Evaluation for Large Language Models Self-Prompt Tuning: Enable Autonomous Role-Playing in LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:57.644959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:57.644959Z digest=sha256:4065e286b65d62057b0300e6ddd805c21b5fd01f6a86aa09f1275303fb205a9b

Observation 70e4910e-65a8-4c41-8f84-d61d8e609693 · outbound

This paper cites Asian Social Science 5(10), 140–143 (2009).

Role-Playing Evaluation for Large Language Models Asian Social Science 5(10), 140–143 (2009)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:58.106895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:57.649693Z digest=sha256:d67a89364a58443a824e228d4c4a9aead57057d6d4dbd4733676a4659c30eea7

Observation c7d85e52-4465-48c6-ab6e-aaf395210704 · outbound

This paper cites Large Language Models are Superpositions of All Characters: Attaining Arbitrary Role-play via Self-Alignment.

Role-Playing Evaluation for Large Language Models Large Language Models are Superpositions of All Characters: Attaining Arbitrary Role-play via Self-Alignment

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:57.654224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:57.654224Z digest=sha256:efc260ef9f68d91735fd16c43c3ea9dbbc39ed02337ae78f856b2325cb6bc43f

Observation fd729776-d955-4850-8be6-677397c54e5d · outbound

This paper cites LLM Discussion: Enhancing the Creativity of Large Language Models via Discussion Framework and Role-Play.

Role-Playing Evaluation for Large Language Models LLM Discussion: Enhancing the Creativity of Large Language Models via Discussion Framework and Role-Play

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:57.659050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:57.659050Z digest=sha256:4c26e82301268a44ef1f5bd3ec6fa1323441002d5924882bc4542dadff8ed94e

Observation 0c530145-e6aa-40c8-b63b-8a857b81946c · outbound

This paper cites BMC medical education7, 1–9 (2007).

Role-Playing Evaluation for Large Language Models BMC medical education7, 1–9 (2007)

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:58.090934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:57.663910Z digest=sha256:866dc23bebea5a8967d34575bc820f9cbaa68b0256e65037a5b3debf79b99a53

Observation e2591919-7042-45f9-89e8-423fa14a6726 · outbound

This paper cites Nature623(7987), 493–498 (2023).

Role-Playing Evaluation for Large Language Models Nature623(7987), 493–498 (2023)

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:58.074920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:57.668879Z digest=sha256:1ac499e33bfc04cc9a5040523938ab5f006cba6d6c182217d05f027f5dffccdc

Observation 39c71228-a6c5-48e3-822e-07ef340fad7f · outbound

This paper cites Character-LLM: A Trainable Agent for Role-Playing.

Role-Playing Evaluation for Large Language Models Character-LLM: A Trainable Agent for Role-Playing

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:57.673580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:57.673580Z digest=sha256:fe7d86fd34276ba26f74944c138bdd695fcb0bccfab1fb758920dc90b5daf45d

Observation aa2976f4-3051-4928-8721-4aaae187c202 · outbound

This paper cites Building Persona Consistent Dialogue Agents with Offline Reinforcement Learning.

Role-Playing Evaluation for Large Language Models Building Persona Consistent Dialogue Agents with Offline Reinforcement Learning

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T20:21:57.886584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:57.678653Z digest=sha256:12c64a4f6ac13b047887785a2ab46d218517e529e4f48603c4b2c6e72df5a1b1

Observation 581fd299-ab40-4d56-bc74-e767de90c656 · outbound

This paper cites BoB: BERT Over BERT for Training Persona-based Dialogue Models from Limited Personalized Data.

Role-Playing Evaluation for Large Language Models BoB: BERT Over BERT for Training Persona-based Dialogue Models from Limited Personalized Data

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T20:21:57.864266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:57.683443Z digest=sha256:d06c4895688e2a2222ac44c9594c189820638a324986486ea37c289998f78de0

Observation 7b1295c3-6096-4525-9d4d-d7223b471890 · outbound

This paper cites In: Workshops at the twenty-sixth AAAI conference on artificial intelligence (2012).

Role-Playing Evaluation for Large Language Models In: Workshops at the twenty-sixth AAAI conference on artificial intelligence (2012)

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:21:58.056942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:57.688228Z digest=sha256:1a4f8fab7c14413aca3c160d81fb8a91c78da9ca4e0fe0e25a1534e87eb5a55a

Observation 9265c7f7-04f0-4511-8acc-41d64f1439cd · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Role-Playing Evaluation for Large Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:57.692953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:57.692953Z digest=sha256:8a28f52fc53849efcdfe83d4dc2f9557894553e1a71eddcb62d87497bb7b6d05

Observation 8ed643c6-a346-466f-b96c-b8ee2d8158cc · outbound

This paper cites CharacterChat: Learning towards Conversational AI with Personalized Social Support.

Role-Playing Evaluation for Large Language Models CharacterChat: Learning towards Conversational AI with Personalized Social Support

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:57.697743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:57.697743Z digest=sha256:2e3e9aae030c570d23255d3f64bb16dc6e4bb9c79b9c8e61f62c2eb6c73b66db

Observation 8e84c5c1-4180-43f4-9033-31852149eebe · outbound

This paper cites Large Language Models are not Fair Evaluators.

Role-Playing Evaluation for Large Language Models Large Language Models are not Fair Evaluators

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:57.702763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:57.702763Z digest=sha256:c33a93a09523346b46bd2b11b34ebf4d407d916d4d8b81933c507ffcf7777a4c

Observation 8f66484a-5c9f-45d4-a936-db1ab169dd0e · outbound

This paper cites RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models.

Role-Playing Evaluation for Large Language Models RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:57.707787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:57.707787Z digest=sha256:703b62d9420b42e32b3de3bd3407811e5c09fe83dda19d172a8c987aff659b1f

Observation 44ad64a7-62df-4bbd-8d79-593de5d08aea · outbound

This paper cites an unresolved cited work.

Role-Playing Evaluation for Large Language Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:58.040978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:21:57.712714Z digest=sha256:60168c77ded5322d59e3fd18ff5734066d2108d8eb9604b36168390c348b729e

Observation 30273fff-343b-49b3-8a02-3bd5a23b4552 · outbound

This paper cites Unveiling the Secrets of Engaging Conversations: Factors that Keep Users Hooked on Role-Playing Dialog Agents.

Role-Playing Evaluation for Large Language Models Unveiling the Secrets of Engaging Conversations: Factors that Keep Users Hooked on Role-Playing Dialog Agents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:57.717635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:57.717635Z digest=sha256:3a27ebeead8e0dc45281a5376f92fae04056ee6cdecba70930150805c3bff5f2

Observation 26f7d78e-7895-4ed4-ba70-d702f95a177e · outbound

This paper cites Personalized Dialogue Generation with Diversified Traits.

Role-Playing Evaluation for Large Language Models Personalized Dialogue Generation with Diversified Traits

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:57.722153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:57.722153Z digest=sha256:7e33f8ceaa2f4f09925d6ceac49d3a8b55a3c744dc5007d24c36dc0792ae7136

Pith citing papers

Observation 74aadbed-d7c0-4f3a-957e-cf860245ca59 · inbound

Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models cites this paper.

Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models Role-Playing Evaluation for Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:20:27.791778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-17T23:16:56.957905Z digest=sha256:714693371924607d471551ff072d6a114b5642ef0fd5d2e13da4faffd3351bcf

Observation 6318671c-ae49-4a6e-b553-94663d258a9b · inbound

Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models cites this paper.

Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models Role-Playing Evaluation for Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:21:01.060182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-10T16:05:09.033412Z digest=sha256:871671f12c0b15f2067f5f0e686754a60a975eb9f1faa2e511f6cae893e19f94

Observation 635a4217-d583-4862-ac9f-9b0bde09a821 · inbound

Improving General Role-Playing Agents via Psychology-Grounded Reasoning and Role-Aware Policy Optimization cites this paper.

Improving General Role-Playing Agents via Psychology-Grounded Reasoning and Role-Aware Policy Optimization Role-Playing Evaluation for Large Language Models

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-06-26T04:58:59.290116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-26T04:55:29.796891Z digest=sha256:9efd214591d90627fc5b47afd76549a1c140217b91b1a423b00c2779cd7d7f71