Pith. sign in

Paper Citation Record · LEDGER

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability

As of 23 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2601.22661.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.22661 v2

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T06:36:33.887146Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 64fed19a-4d58-49a6-867f-8977988efecd · outbound

This paper cites TTS-1 Technical Report.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability TTS-1 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:31.523597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:31.523597Z digest=sha256:2474b2ab69bf0140efe544a13ea122474cbf7382225815bc79a69de4bf173382

Observation 187d4839-c0fb-4345-b711-fafa91a4708a · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:32.017309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:32.017309Z digest=sha256:b71294272925daf9658451fd0e041f8d24e0c720b9184df63685872c0c611b18

Observation 5669cd96-e95c-4b59-a58c-f0a29b7d705e · outbound

This paper cites Differentiable Reward Optimization for LLM based TTS system.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability Differentiable Reward Optimization for LLM based TTS system

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:32.240560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:32.240560Z digest=sha256:dfb375cf6c0a9d0110c8e0a65480ec45530cdf5779a456f6b015900f290aa21e

Observation 38d72d7a-b07b-4921-b6cc-eb25bd3d5477 · outbound

This paper cites Prompttts: Controllable text-to-speech with text descriptions.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability Prompttts: Controllable text-to-speech with text descriptions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:32.326580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:32.326580Z digest=sha256:e074485bf065925c3a76fc6f3abef1aa8fe4a4d8c00a4032547da4ecaeed574d

Observation 87f751dc-da9c-46cb-9382-42db3a39acff · outbound

This paper cites Qwen3-TTS Technical Report.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability Qwen3-TTS Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:32.393867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:32.393867Z digest=sha256:ce92a7f2ffd59462c1da9b8a448a6f794c819639075e81c48d0b0d8cb94f2f91

Observation 7a67601c-c652-4d21-ab6c-cf428f010520 · outbound

This paper cites Speechrole: A large-scale dataset and benchmark for evaluating speech role-playing agents.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability Speechrole: A large-scale dataset and benchmark for evaluating speech role-playing agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:32.453974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:32.453974Z digest=sha256:ff5fd354d2114f48440554132ce0b140f203147961ff4068b48d95a8b07b15fc

Observation 41bc5a40-4de2-4557-a59c-22cda5da7818 · outbound

This paper cites PromptTTS 2: Describing and Generating Voices with Text Prompt.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability PromptTTS 2: Describing and Generating Voices with Text Prompt

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:32.513360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:32.513360Z digest=sha256:cfc0be88ee662d76ffbf799773f555bab9a5e29e709cd11033fad03b3f3b19bf

Observation 7f05627a-4879-44d3-83b5-c725b6760a3d · outbound

This paper cites AudioRole: An Audio Dataset for Character Role-Playing in Large Language Models.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability AudioRole: An Audio Dataset for Character Role-Playing in Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:32.600302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:32.600302Z digest=sha256:2eb595753f6acc0f0fe4ec477e96e7f4d01a1e1d106d2e63020c5744842b3c47

Observation 8c2c6535-9aa0-4a03-8ede-cde0e03d33f1 · outbound

This paper cites Reinforcement Learning for Emotional Text-to-Speech Synthesis with Improved Emotion Discriminability.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability Reinforcement Learning for Emotional Text-to-Speech Synthesis with Improved Emotion Discriminability

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:32.668134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:32.668134Z digest=sha256:9b38b0394dc02309a1e92757f24a1d0a741d9a7814d7868e26f62be9afdffbbd

Observation 15b907bd-a0de-4e76-a710-3cdf41b45d8d · outbound

This paper cites Hybrid transform- ers for music source separation.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability Hybrid transform- ers for music source separation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:32.814773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:32.814773Z digest=sha256:73e1d9e7f362c61d7b9d8947f7b38ddcd9fca7ea7e2fbf9df1231c55af2c77fd

Observation 924b277b-9b4b-4a20-8414-1fb637b15a98 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:32.939198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:32.939198Z digest=sha256:42b725fb3ef8bc5b02a60a0ae9b1fb1ea67cc9419a67ef5c190367506a9bb039

Observation 75574330-b66f-4077-b7de-0e26c3cb36da · outbound

This paper cites Speech-drame: A frame- work for human-aligned benchmarks in speech role-play.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability Speech-drame: A frame- work for human-aligned benchmarks in speech role-play

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:33.055238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:33.055238Z digest=sha256:4aef353590a99bdb011d52b43d287aed037494c99f386cfb858f9324e60448e5

Observation 790dd417-4554-47a2-8075-15009e99c47e · outbound

This paper cites Rrpo: Robust reward policy optimization for llm-based emotional tts.arXiv preprint arXiv:2512.04552,.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability Rrpo: Robust reward policy optimization for llm-based emotional tts.arXiv preprint arXiv:2512.04552,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:33.244749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:33.244749Z digest=sha256:42935297ed25d2a5d773563a942286468e90bc6e0dc5a75e9251898a816b7239

Observation 2c2610fc-6b84-4fc3-a315-671642161b80 · outbound

This paper cites Step-Audio 2 Technical Report.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability Step-Audio 2 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:33.364746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:33.364746Z digest=sha256:ddfa12436b9a9f45b43d6eee0ce68c4e6cba5c1e19426f5c73a1c3a98ff0f3bf

Observation 351a3903-35eb-426c-973b-757458d38f9f · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:33.475394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:33.475394Z digest=sha256:695b9e4843d078239d1b8944a8bba0ace84c1dd8d0b294f5a9a94c7a993a19a6

Observation 0a09d3b0-8e4b-48a4-8c89-aa5866e2b82a · outbound

This paper cites Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:33.674758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:33.674758Z digest=sha256:f1f81fbb809c9b4f49a8ddf18c1e267e870d40de19018b9265682bfcf50e3413

Observation 5f075c68-f66a-4551-bea9-ff9ec60c0b40 · outbound

This paper cites Mimo-audio: Audio language models are few-shot learners.arXiv preprint arXiv:2512.23808, 2025a.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability Mimo-audio: Audio language models are few-shot learners.arXiv preprint arXiv:2512.23808, 2025a

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:33.830007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:33.830007Z digest=sha256:14cfd7e215856065001dc248ba195e8e77024d3bfde9ea2b106a805065390127

Observation dde0822b-38ab-4c94-9f46-3701b307eb16 · outbound

This paper cites drama_name.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability drama_name

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:33.887146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:33.887146Z digest=sha256:a19dee68de25a7767202521f69669052d1eb4e8ffe0d11fe51bbfac9b896b4f3

Observation c7d6f970-913a-41b2-9fca-53b48d88bd28 · outbound

This paper cites Ov-instructtts: Towards open-vocabulary instruct text-to-speech.arXiv preprint arXiv:2601.01459,.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability Ov-instructtts: Towards open-vocabulary instruct text-to-speech.arXiv preprint arXiv:2601.01459,

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:32.729198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:32.729198Z digest=sha256:605ed78307ee1ac1472462027178e4e8db2a0a6e4153bc119890de95e5f995c9

Observation 78d2d6a9-4041-44da-9dfb-4a344fb57ac6 · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:33.115934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:33.115934Z digest=sha256:0589b74b6d52b905b0047dd83e612807bd76238d5cb2209320714a3366b753dd

Observation 69053db0-0cbd-4e31-9b02-cea958287e84 · outbound

This paper cites Flexivoice: Enabling flexible style control in zero- shot tts with natural language instructions.arXiv preprint arXiv:2601.04656,.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability Flexivoice: Enabling flexible style control in zero- shot tts with natural language instructions.arXiv preprint arXiv:2601.04656,

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:31.747286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:31.747286Z digest=sha256:55598eb2c5855c309df8893656ad8ec79e537213e520a48e6cac3497fe2b3e2d

Observation 013a0647-317f-4345-9b7f-f67b38f00c34 · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:32.118897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:32.118897Z digest=sha256:303b919c33eb72cb688c6dddb37d734bcad734810eb5a775b95fbd4c48665fee

Observation 11aa3de7-df12-4be4-9338-347592404301 · outbound

This paper cites Qwen3-VL Technical Report.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability Qwen3-VL Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:31.646881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:31.646881Z digest=sha256:2331b505e9d2b5693b1eea5f5e52fa52148760a6608d8b7f67dc06e3be564d64

Observation 21fdd2b4-abf1-44b8-a7b8-baf9025ebf8e · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:31.871558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:31.871558Z digest=sha256:ec1a36ff40d189b11303c3f24643789cf3a41dc6606f7650931cf746e02de753

Pith citing papers

No inbound Pith citation observations are available.