Pith. sign in

Paper Citation Record · LEDGER

RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2312.16132.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.16132 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:58:03.233374Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:26:56.284609Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b2df884a-9610-4c1b-bfab-dc40fc85d249 · inbound

Star-Agents: Automatic Data Optimization with LLM Agents for Instruction Tuning cites this paper.

Star-Agents: Automatic Data Optimization with LLM Agents for Instruction Tuning RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:03.233374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:03.233374Z digest=sha256:73763f8ba231c509caa675a186443cd5ae0cf579e7e3c011d6b772bf7d643ba3

Observation e53bf8d8-ee54-47f2-8e0a-8de62ac2f943 · inbound

Personalized LLM for Generating Customized Responses to the Same Query from Different Users cites this paper.

Personalized LLM for Generating Customized Responses to the Same Query from Different Users RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T14:44:37.111621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:44:37.111621Z digest=sha256:86da8de6e9bad81d9b89fad445864a4cb266acab0da9402c3f90216fcf000803

Observation 97e458c5-a92b-4eaf-b794-1b648663accb · inbound

Beyond Math: Stories as a Testbed for Memorization-Constrained Reasoning in LLMs cites this paper.

Beyond Math: Stories as a Testbed for Memorization-Constrained Reasoning in LLMs RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:22:38.369848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-23T06:20:49.096900Z digest=sha256:fa87cb8a3840dd05149e8d915a857621190b531b1df28be2df6e4b797de2bdea

Observation 8b23cddf-ee0a-4dcb-bb08-e87e84c8344b · inbound

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning cites this paper.

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:25:30.114386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T14:22:25.660785Z digest=sha256:4d9808221c711f34fb45ae8cd7b3b1dad3c7fd79e5e092d5db7ba11a35465eaa

Observation a0c3796f-39e3-4af2-bec2-537b97f9b1ed · inbound

BOOKMARKS: Efficient Active Storyline Memory for Role-playing cites this paper.

BOOKMARKS: Efficient Active Storyline Memory for Role-playing RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

Reference 114

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T04:55:03.885367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-15T04:51:44.394368Z digest=sha256:54b572d02aba3f16e6b7a9e43a2f36fb90fe8bcc9f0bd5ff7bfe69c3861b2d1b

Observation 6ff4c808-cd80-4f05-b07a-059a958857c8 · inbound

PersonaArena: Dynamic Simulation for Evaluating and Enhancing Persona-Level Role-Playing in Large Language Models cites this paper.

PersonaArena: Dynamic Simulation for Evaluating and Enhancing Persona-Level Role-Playing in Large Language Models RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T15:48:32.102867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T15:45:21.524760Z digest=sha256:76daef0e689f19af553db00a4571db8c260eaba1f96f4f77d9f48b6291b05fec

Observation 70f0897c-49d0-4842-910e-b8a036789cce · inbound

RoleCDE:Benchmarking and Mitigating Role-Alignment Trade-offs in Role-Playing Agents cites this paper.

RoleCDE:Benchmarking and Mitigating Role-Alignment Trade-offs in Role-Playing Agents RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:20.002713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T15:00:57.851111Z digest=sha256:d511072044b50ab338195e13f8b6ccf06812469e8daab90622ea92f6025d4261

Observation fb08e66e-b89c-4154-a6c3-36c9af69f4cb · inbound

ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time? cites this paper.

ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time? RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:26:56.286102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T02:11:04.565958Z digest=sha256:dc80321e97d49915040bd0ab1b1a4923d3ef7a1f1fba2501cf2cbfe8a316f2be

Observation 9efd854c-5cd3-4898-94e5-3f505bd59384 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.704756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.704756Z digest=sha256:5a5a1663956bc53d1c02b891bb87dedce89f5532bbac7800f6e7797db225c298