Pith. sign in

Paper Citation Record · LEDGER

EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2308.05725.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2308.05725 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:23:40.366858Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:38:28.620667Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0d5f7645-2d04-4d74-be77-c2ff20279371 · inbound

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training cites this paper.

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:27:25.612037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T05:27:25.425188Z digest=sha256:645cec6ee99174b644a3f0ffeb872441dcca2c71f3572250b02cb5250cde4cc7

Observation 742884e6-2cf3-47cd-8ac0-ae6d3c53620f · inbound

WHISTRESS: Enriching Transcriptions with Sentence Stress Detection cites this paper.

WHISTRESS: Enriching Transcriptions with Sentence Stress Detection EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:40.366858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:40.366858Z digest=sha256:5992cad70b627aa1c0411c97502dffe6bebc746e5b484c79f3d5697cef5011f2

Observation 5ab751e6-8225-4415-82ca-520f8cb83fcc · inbound

StressTest: Can YOUR Speech LM Handle the Stress? cites this paper.

StressTest: Can YOUR Speech LM Handle the Stress? EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:02:18.259217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T13:00:23.002962Z digest=sha256:f2950d27b638c3c361093c1e4c9653693039c3e07ac794af9775ad44a28266ba

Observation a5e2b151-c0e1-4e2c-beb8-922275a56e54 · inbound

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion cites this paper.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.345371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.345371Z digest=sha256:8e929c9fa1010c9ea44638b9bbf7d096f32e67e73f50fea976d3cdad1416e4c7

Observation 79602e4a-bfe0-4eec-9975-1136cfc0f26c · inbound

NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech cites this paper.

NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T16:33:52.438991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:33:52.438991Z digest=sha256:da9e11b728587ea20314a4e17452c674a5247b4a399956232259809be2f5d3a8

Observation 4402f487-40f7-461f-be2d-02e40eaa173d · inbound

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations cites this paper.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:31.393989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:31.393989Z digest=sha256:f5bf970dc8302c1fb1d7b1951d3f04a4a3f6c5afb08bef1e869f8db78e00bab5

Observation 2bbcfe48-1e89-402a-ad67-8a90aa46b3da · inbound

Computational Narrative Understanding for Expressive Text-to-Speech cites this paper.

Computational Narrative Understanding for Expressive Text-to-Speech EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:31:47.053839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T19:30:53.250247Z digest=sha256:b0a63d9a33be105b72333a47e6fe5bf1fa37d47fd77ed21e5448441219201a5d

Observation 8a8b3a10-c5d3-42aa-bd5c-3168ab4bdcfe · inbound

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs cites this paper.

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:01:24.344834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T12:57:04.450462Z digest=sha256:2887292515371ee993b31325851b9b1d467212eb659a581ba9c0952b3d2231eb

Observation 0252de93-6edd-4e86-b63c-26a9a7e93e21 · inbound

AudioSAE: Towards Understanding of Audio-Processing Models with Sparse AutoEncoders cites this paper.

AudioSAE: Towards Understanding of Audio-Processing Models with Sparse AutoEncoders EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T04:28:01.655676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:28:01.655676Z digest=sha256:478a2c6a2615b31c5d12061e2cbce4de6dacb9402c64a1138f8b7c3fee29e1e3

Observation d8ea44ac-40cb-4828-8173-552871608972 · inbound

Voxtral TTS cites this paper.

Voxtral TTS EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:39:35.896911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:8277c5acc52bd476d85cee05cd96740f98c524b1e0b82479ba8f8496fc5efd57

Observation 549b17d1-640d-431b-8730-8f5d3ea72df8 · inbound

JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions cites this paper.

JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:01:08.380039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T16:43:33.397158Z digest=sha256:37e323576c5bed15c1a74b43c980f3d7971c0afc593439adb652dd49a51d0df8

Observation 0661bbb5-e44f-4264-8a08-e077c33939ac · inbound

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech cites this paper.

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:33:55.471353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T02:32:26.122526Z digest=sha256:7df837186ac5cb99a2c0d4a0db4912eb72da80db37d46f7fe545df3ef3cedf1b

Observation ff5f36a2-5395-4982-b1d5-1428c3a7f1f3 · inbound

Toward Natural Emotional Text-To-Speech System with Fine-Grained Non-Verbal Expression Control cites this paper.

Toward Natural Emotional Text-To-Speech System with Fine-Grained Non-Verbal Expression Control EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T20:53:57.753462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T20:53:47.252916Z digest=sha256:26ec18d214cb405de6bf276e4adb2a734e8ff5bf48c16c2b89bb53f6be50d8c7

Observation b73b78d3-4db9-431a-b7d5-8d14dca35cc9 · inbound

Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS cites this paper.

Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:16:11.068632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T21:31:09.399885Z digest=sha256:7cdf31439a5c32cef1a158c1a2bac05b709030e3704189e2c0ce76a24a346dc6

Observation 0ef35b23-59a6-4ae5-b80e-62ec4283707a · inbound

UniVocal: Unified Speech-Singing Code-Switching Synthesis cites this paper.

UniVocal: Unified Speech-Singing Code-Switching Synthesis EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T00:46:24.556094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T13:17:13.510587Z digest=sha256:27adba1bb1742f57dbcb17bdd915464ca83c856e660038c48fa400458756fe6d

Observation 483f3ab7-4cda-43af-8e3c-9493413f6c44 · inbound

Unlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning cites this paper.

Unlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:38:28.622473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T14:35:17.004680Z digest=sha256:a1af9539e0198e4e3093c7d203e5c399185675148c4fe4bfaa0d0ae45103bc7a

Observation 612af17a-827e-4187-b885-0b712fcafe7d · inbound

Dialogs: a studio-quality expressive conversational Russian speech corpus for dialog assistants cites this paper.

Dialogs: a studio-quality expressive conversational Russian speech corpus for dialog assistants EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T02:30:34.179334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:30:34.179334Z digest=sha256:b681a353728c2cd30f1856e3fe5adca8942409362918bddda23117edac514cc7

Observation 256e3517-1a63-4ca0-b573-407c72110a94 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:10.474782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:10.474782Z digest=sha256:c1c9dfd24c5c1b23b8dab510227af42258506fc98a9ba2b98bd1337c23510623

Observation fec8ed68-93ac-427b-a97e-d3dea908ede9 · inbound

Beyond One-Size-Fits-All: Personalized and Culturally Adaptive Emotional TTS via Interactive Optimization of Individual Emotion Perception Spaces cites this paper.

Beyond One-Size-Fits-All: Personalized and Culturally Adaptive Emotional TTS via Interactive Optimization of Individual Emotion Perception Spaces EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T00:40:33.274250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:40:33.274250Z digest=sha256:597811dc90c24a5f32a85daae8b949dc78f55d7e556119d281ddca7365066768