Pith. sign in

Paper Citation Record · LEDGER

RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2312.16132.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.16132 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:58:03.233374Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:26:56.284609Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b2df884a-9610-4c1b-bfab-dc40fc85d249 · inbound

Star-Agents: Automatic Data Optimization with LLM Agents for Instruction Tuning cites this paper.

Star-Agents: Automatic Data Optimization with LLM Agents for Instruction Tuning RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:03.233374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:03.233374Z digest=sha256:e3bc217d347af1ba0afb1238bfb31dbdb1f37175080b91b9f2eb53c691e9a34d

Observation e53bf8d8-ee54-47f2-8e0a-8de62ac2f943 · inbound

Personalized LLM for Generating Customized Responses to the Same Query from Different Users cites this paper.

Personalized LLM for Generating Customized Responses to the Same Query from Different Users RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T14:44:37.111621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:44:37.111621Z digest=sha256:f4ec631764a758521fdff1df4821f57f2362f31cbfbd1ec9259f5b376a92a76f

Observation 97e458c5-a92b-4eaf-b794-1b648663accb · inbound

Beyond Math: Stories as a Testbed for Memorization-Constrained Reasoning in LLMs cites this paper.

Beyond Math: Stories as a Testbed for Memorization-Constrained Reasoning in LLMs RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:22:38.369848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T06:20:49.096900Z digest=sha256:05ef532cc37c0e3bffb2ac46a11191271e1dd40917c0e1636c7ff3ac3854eded

Observation 8b23cddf-ee0a-4dcb-bb08-e87e84c8344b · inbound

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning cites this paper.

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:25:30.114386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T14:22:25.660785Z digest=sha256:d9b6e17dbf16a5d020b0c6b79fdd3b737ba6d564b35433b2576f0fd17c7aaf8b

Observation a0c3796f-39e3-4af2-bec2-537b97f9b1ed · inbound

BOOKMARKS: Efficient Active Storyline Memory for Role-playing cites this paper.

BOOKMARKS: Efficient Active Storyline Memory for Role-playing RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

Reference 114

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T04:55:03.885367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-15T04:51:44.394368Z digest=sha256:cde42bd593557819eca3c8e35eb2a985b0ac144da23329a1d4b5518ecee3aedd

Observation 6ff4c808-cd80-4f05-b07a-059a958857c8 · inbound

PersonaArena: Dynamic Simulation for Evaluating and Enhancing Persona-Level Role-Playing in Large Language Models cites this paper.

PersonaArena: Dynamic Simulation for Evaluating and Enhancing Persona-Level Role-Playing in Large Language Models RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T15:48:32.102867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T15:45:21.524760Z digest=sha256:2322fc44c1cae03b0ce4ac58475fbb7f1551c1b0f8c4508cea340066c63463b7

Observation 70f0897c-49d0-4842-910e-b8a036789cce · inbound

RoleCDE:Benchmarking and Mitigating Role-Alignment Trade-offs in Role-Playing Agents cites this paper.

RoleCDE:Benchmarking and Mitigating Role-Alignment Trade-offs in Role-Playing Agents RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:20.002713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T15:00:57.851111Z digest=sha256:3f35b076150b767f398dbd0d709c5164b243685854660360e804cc19e3a9b1d2

Observation fb08e66e-b89c-4154-a6c3-36c9af69f4cb · inbound

ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time? cites this paper.

ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time? RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:26:56.286102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:11:04.565958Z digest=sha256:0ff298413e3868f9ecd15e74c7e4304e01de667fc0d67fb5d90b29fbcbc19429

Observation 9efd854c-5cd3-4898-94e5-3f505bd59384 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.704756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.704756Z digest=sha256:5eca424dca1146f58972f28df784faf0f3bbbc4bc13b6124fa11aaba353b55cc