Pith. sign in

Paper Citation Record · LEDGER

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents

As of 21 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 2 inbound Pith citation observations for arXiv:2509.03940.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.03940 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:36:06.487611Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T23:23:47.291534Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T22:39:01.897662Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d182ce33-700e-4508-81bd-68d54a27bbd5 · outbound

This paper cites VoiceBench: Benchmarking LLM-Based Voice Assistants.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents VoiceBench: Benchmarking LLM-Based Voice Assistants

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.397319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.397319Z digest=sha256:2739aeeb6d3b9a1d8f580bf637fe9b1a0f9c118c357ef5f09ede8deaef16fb95

Observation 879e9837-eb19-4e3d-9436-7e39a6276806 · outbound

This paper cites MMRole: A Comprehensive Framework for Developing and Evaluating Multimodal Role-Playing Agents.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents MMRole: A Comprehensive Framework for Developing and Evaluating Multimodal Role-Playing Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.402445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.402445Z digest=sha256:1679ae65878c737df11290ec41425717bf46b4177b61fcca53f79b9af3ee1fb2

Observation 5967656f-7727-47a0-9e23-059b5d5d582c · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents Moshi: a speech-text foundation model for real-time dialogue

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.407429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.407429Z digest=sha256:4f2fea47fb1dc11e9f1b9e5f4ee05bc86fc3c45dda811f5ecf46b6298aae558d

Observation 4b0f3d85-50f1-451f-84ef-70923caf34f5 · outbound

This paper cites GPT-4o System Card.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents GPT-4o System Card

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.417604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.417604Z digest=sha256:39f70d295f977d03d07cc1fff2c715ab75843561c11b3fc52e83d40b7c839a4f

Observation 2ac3479e-559f-4eb7-959a-cad129263656 · outbound

This paper cites Evaluating Large Language Models in Theory of Mind Tasks.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents Evaluating Large Language Models in Theory of Mind Tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.423054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.423054Z digest=sha256:30050793b998f7578ebcbceb46231dcfdb899a333617de51f0fe0520297dd718

Observation 6eb69524-e602-42c6-9335-744ef968d22f · outbound

This paper cites Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.427752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.427752Z digest=sha256:ef350eb68784e470ebfa840ca3c0d223e19f9fdf73ca57c6a028c7db19ed7454

Observation 6ae4e850-b523-4a30-8178-090802812d79 · outbound

This paper cites Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.433253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.433253Z digest=sha256:9e04e92aa6fe7bfd37ece56d7f94d439386ec168ce6af891e96f94f547695933

Observation 5d3bde82-2bb4-4935-b1ff-416b2c91e569 · outbound

This paper cites RoleMRC: A Fine-Grained Composite Benchmark for Role-Playing and Instruction-Following.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents RoleMRC: A Fine-Grained Composite Benchmark for Role-Playing and Instruction-Following

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.438941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.438941Z digest=sha256:69e26444dce20b4a10462559e49eab6b51655def80d1816fe1240f113022a091

Observation 4f720ed5-cb13-4286-845c-a4e5e3b0d86e · outbound

This paper cites emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.444581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.444581Z digest=sha256:b958a167b1735cdb1347956bbe1510152635fdc0715431e9aeee3f4d27300371

Observation d0ab30de-2910-480f-a0ed-12d34126c37e · outbound

This paper cites CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.450675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.450675Z digest=sha256:cd3749a466fdb0fb04987ff7e4da267cd5c21b18fbb68f28fb2c1dc60c8cc22c

Observation 49028659-6b2a-41c8-8c9e-c9be2fd46ec0 · outbound

This paper cites arXiv preprint arXiv:2502.09082.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents arXiv preprint arXiv:2502.09082

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.457095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.457095Z digest=sha256:53ab55dd96f26a536e4da8cf344c8d19a7113de09c26647aa0c59cbf6b99aa6d

Observation db3614dc-7a7f-4451-be46-e0fce240432e · outbound

This paper cites RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.461904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.461904Z digest=sha256:f0aa98e049c46afc9b1fe1956b1b916cd049f43818f8bfa42166e730b39fd61d

Observation 83a874ab-64a4-4e31-bbaa-4c5a3015a8e2 · outbound

This paper cites Qwen2.5-Omni Technical Report.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents Qwen2.5-Omni Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.466980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.466980Z digest=sha256:7d3574789040fda1b9143919153979a38c8a0ecc50086e0006870c082a87eff5

Observation d17a3726-3a7b-4486-a695-6b0d4d77566b · outbound

This paper cites URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.472390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.472390Z digest=sha256:fd1a5aad0d31e6f91dfd0ffa2f5744eba7b8f8831376157c93e6036b61b8f26c

Observation 8c5927f2-6892-4b69-8a73-7dd5735b9669 · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.477526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.477526Z digest=sha256:806b6d1b803591681ae02e171cdab5283c4e1d9516a1f1463fbc2b0655ca52e8

Observation a4a6a580-96cf-40c4-b852-4d92d09dcbcc · outbound

This paper cites OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.482387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.482387Z digest=sha256:648860761e0ad2cbae677fd45617226bf3cceecc83d9ecb476f6637db294497e

Observation dfb5b3fe-a44e-4673-a7ba-aa74c33934b5 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents BERTScore: Evaluating Text Generation with BERT

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.487611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.487611Z digest=sha256:cacb537f77b1c993c1ec3cde400d3c719d9501c3855fbb53254dcfc02542da5e

Observation 00b818e6-1cb8-42b1-ad9b-23c49ca56607 · outbound

This paper cites GPT-4 Technical Report.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents GPT-4 Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.387483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.387483Z digest=sha256:3f66e69429487e89f813d45b1b836b3019eef96c2e11c9d592be9434a40ee559

Observation 449e0b3b-926a-412a-aac5-ca68e7b7b30c · outbound

This paper cites In 2024 IEEE Spo- ken Language Technology Workshop (SLT), 818–824.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents In 2024 IEEE Spo- ken Language Technology Workshop (SLT), 818–824

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:36:06.930583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:36:06.392409Z digest=sha256:6b6df56f2acdb114a55338350025187553925956a83f7a4cc01263645dca4c62

Observation 10e02dce-04fd-4a16-b62d-ee6201480ee7 · outbound

This paper cites Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.412825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.412825Z digest=sha256:df5eb47a2e4d89d5084e0e5063b49b0138b9b23869328df3cc5637bf6a147623

Pith citing papers

Observation 211cbeb2-0bfd-4997-86a0-86c4add2122c · inbound

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning cites this paper.

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:25:30.165181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T14:22:25.660785Z digest=sha256:eb1ffb012a4acd8024a19769cbced91d28260beebe0db34a6da0e0d9de4c1281

Observation f3ba0f25-1e94-4539-9d43-587ae4f8c0d5 · inbound

DeSRPA: Decoupled Speech Role-Playing Agent via Inference-Time Intervention cites this paper.

DeSRPA: Decoupled Speech Role-Playing Agent via Inference-Time Intervention VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-03T22:39:01.899081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T23:23:47.291534Z digest=sha256:f396b0e5de33d1eb0a09a4dc871258988cdd6b2403310adf532b37796960cca1