Pith. sign in

Paper Citation Record · LEDGER

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark

As of 3 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 1 inbound Pith citation observation for arXiv:2604.10580.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.10580 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:09:21.027609Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:09:21.027609Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-11T11:06:05.422750Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact7
  • verified fuzzy12
  • unresolved3
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 98a0d030-2482-442f-a86a-cac2d6234697 · outbound

This paper cites The manager booked the flight.

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark The manager booked the flight

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:36:04.352902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:09:21.027609Z digest=sha256:682104eab3e209498d5ee4efb5fe913f48e37f40a20cdd41aff6a7892cb582a6

Observation e1449494-7048-4ae3-987f-528810084254 · outbound

This paper cites Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark.

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T11:06:05.434356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:09:21.027609Z digest=sha256:b6438d123ff164d7b629713c1c05d75d820fee0e73a0b814aa29bc592d83fd3b

Observation ded03893-93b2-446e-988f-64b603df93ea · outbound

This paper cites Task We consider the task of context-conditioned stress generation in TTS.

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark Task We consider the task of context-conditioned stress generation in TTS

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-05-18T07:36:04.344178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:09:21.027609Z digest=sha256:75b53acc53e2f2941c2a66633f97e69e980861499991edb41e933020397a027b

Observation 075ccf6a-4690-4403-8881-43a00aafda5b · outbound

This paper cites Evaluated Systems We evaluate a diverse set of TTS systems, each under the con- ditioning modes it natively supports.

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark Evaluated Systems We evaluate a diverse set of TTS systems, each under the con- ditioning modes it natively supports

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-05-18T07:36:04.350200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:09:21.027609Z digest=sha256:91939569d7ea6fea3c01b434bc6709d59cb0736f87c7c29b24a96e0019a1285c

Observation 547eada3-d7f8-453d-bd04-fb3e04500cdc · outbound

This paper cites Context-Aware Stress Realization Table 3 summarizes performance across models and input modes.

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark Context-Aware Stress Realization Table 3 summarizes performance across models and input modes

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:36:04.341271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:09:21.027609Z digest=sha256:f8f018d4bae077a04866d63e174cbd70f6325ffa2ca5a20b90189b7c704f66b6

Observation e10ed520-3fc1-4794-bab0-082e2f951aba · outbound

This paper cites an unresolved cited work.

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-18T07:36:04.323800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:09:21.027609Z digest=sha256:0799664290f90dec56373f8f43b05c5788a86876e94a59b83a310c42107ccba9

Observation f80c4d76-71a6-4c09-b5ae-acf6b9d21eea · outbound

This paper cites an unresolved cited work.

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-18T07:36:04.332836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:09:21.027609Z digest=sha256:68479949d5347c8122207417927dbc8c8762b37cb774b9485e857bb209b837f1

Observation b8b7b5fb-45b5-429c-98fc-834b96af2015 · outbound

This paper cites an unresolved cited work.

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-18T07:36:04.326473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:09:21.027609Z digest=sha256:aaa3cca6c9c75c4cbd4d3dfee5a58f06daf0471251bbcfbb5b5f2ec3fd604bb1

Observation 0f0e4e28-0312-4133-b543-9dcf3a28bb9e · outbound

This paper cites A theory of focus interpretation.

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark A theory of focus interpretation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:06:05.454407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:09:21.027609Z digest=sha256:ae5449862a9d00107bdafb9929227e3ac4e99cae484b95cb09d63e4e5791440c

Observation 05afb797-44a8-45ea-b75d-1c2fa24603c6 · outbound

This paper cites In Christopher Hitchcock & Alan Hajek, edi- tors: Oxford Handbook of Probability and Philosophy , Oxford University Press, pp.

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark In Christopher Hitchcock & Alan Hajek, edi- tors: Oxford Handbook of Probability and Philosophy , Oxford University Press, pp

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:06:05.497214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:09:21.027609Z digest=sha256:c59eca399013652e912fe63a870f28833df77a822d5569316b3e4b67388100f0

Observation 4ffa384e-b624-47c7-ae8d-f7ede0ab4e56 · outbound

This paper cites Accent is predictable (if you’re a mind- reader).

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark Accent is predictable (if you’re a mind- reader)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:36:04.338449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:09:21.027609Z digest=sha256:077ee703edf9e18f5e533084ba42d04d0adee4c1c93a417b3a2709c4df7e88b5

Observation b3e1d252-be36-442f-9592-2b39002a1514 · outbound

This paper cites StressTest: Can YOUR Speech LM Handle the Stress?.

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark StressTest: Can YOUR Speech LM Handle the Stress?

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:06:05.389620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:09:21.027609Z digest=sha256:49b12e64afa6ee26738b50b6dae8f210dd8c23a014df85141595585bcceddda5

Observation 906bf905-4e4c-410b-a420-194026ab11a6 · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:27:25.722747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:09:21.027609Z digest=sha256:9e63938e5a0515ec4fb4b4447e30684e8b23ce9271ba296d91c24639a7bb0a6c

Observation 84ab07b3-8eb8-4a62-a79e-06bf7b1ff994 · outbound

This paper cites Chatterbox-TTS.

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark Chatterbox-TTS

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:36:04.355717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:09:21.027609Z digest=sha256:dab916c17fc9fd01c76e3d4c30bdb40ccf845c939333f5c9cb62aaa6fb547446

Observation 74885d97-a12e-47a7-8354-057f72d76eb5 · outbound

This paper cites Qwen3-TTS Technical Report.

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark Qwen3-TTS Technical Report

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:24:56.189411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:09:21.027609Z digest=sha256:4beac5d8b7e99aeeaafc22bc7c0a912ab7d20435495a917fc2fbe922c302c772

Observation 550de9f9-faa1-4d9b-87a6-05b02ba7445e · outbound

This paper cites WHISTRESS: Enriching Transcriptions with Sentence Stress Detection.

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark WHISTRESS: Enriching Transcriptions with Sentence Stress Detection

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:06:05.375927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:09:21.027609Z digest=sha256:acc80ee065ada56be284b8ade6c2fc18819a3abdc5a14c66b1d80aa0a12607e2

Observation 8526020e-8b7c-4a8b-b960-a073633542f7 · outbound

This paper cites Crowd- sourced and automatic speech prominence estimation.

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark Crowd- sourced and automatic speech prominence estimation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:36:04.347152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:09:21.027609Z digest=sha256:0f5bc7dbe2896c2f4ccdcc5c50170630dc7ad0ef594b21fbb6d13de04c4d6943

Observation 8065e57a-1b1a-46ed-8387-a1c4b255cb0c · outbound

This paper cites Em- phassess: a prosodic benchmark on assessing emphasis transfer in speech-to-speech models.

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark Em- phassess: a prosodic benchmark on assessing emphasis transfer in speech-to-speech models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:36:04.308827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:09:21.027609Z digest=sha256:d4d42d11db448478a23d5e65c16f4d87cc2f6551d8d14c492876521b799c659a

Observation 20785ebe-2c82-4783-8f3d-c216c66663c5 · outbound

This paper cites ProsAudit, a prosodic benchmark for self-supervised speech models.

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark ProsAudit, a prosodic benchmark for self-supervised speech models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:06:05.381772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:09:21.027609Z digest=sha256:4d0e4c9b582636536ae8ed5a979bb6e62f0e57ca99c5dcf039733ab1208c0c96

Observation 06d34669-4640-456c-ab76-18cada32fb39 · outbound

This paper cites Style tokens: Un- supervised style modeling, control and transfer in end-to-end speech synthesis.

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark Style tokens: Un- supervised style modeling, control and transfer in end-to-end speech synthesis

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:36:04.321271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:09:21.027609Z digest=sha256:fdeb43196be0e13f6cd90a7a13c34525db83e126fb3ec13594faba6d7bb8e9dc

Observation b8974647-ad6e-4e36-8b35-22f81580d60e · outbound

This paper cites Towards End-to-End Prosody Transfer for Expressive Speech Synthesis with Tacotron.

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark Towards End-to-End Prosody Transfer for Expressive Speech Synthesis with Tacotron

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:06:05.530718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:09:21.027609Z digest=sha256:a04b2b6ec2f29b6817b500051772fa6fb87d331aba13f071dda114848efe7300

Observation c2d0302b-4536-4d70-8943-892e226fe0e3 · outbound

This paper cites Styler: Style factor mod- eling with continuous prompt for expressive speech synthesis.

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark Styler: Style factor mod- eling with continuous prompt for expressive speech synthesis

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:36:04.335436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:09:21.027609Z digest=sha256:e6b784036d3a38baca063b6003acb0196e78516fe8c78112e1a906d4234fc704

Observation 1487332f-97cc-4c8b-b90b-49402ebc78d8 · outbound

This paper cites To- wards end-to-end prosody transfer for expressive speech synthesis with tacotron.

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark To- wards end-to-end prosody transfer for expressive speech synthesis with tacotron

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:36:04.318342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:09:21.027609Z digest=sha256:f68a937d5b0036299a7dd45fa0064553d93ab433b33d2f58260f9eed9cd645b5

Observation 97511ef8-c786-4a96-b0d5-b7d9850b2d9d · outbound

This paper cites Word-level text markup for prosody control in speech synthesis.

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark Word-level text markup for prosody control in speech synthesis

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:36:04.329765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:09:21.027609Z digest=sha256:d1c25447f1a01396670b8ca6246c8c4dee33552a2e4d4b79c437de6356efe169

Observation efaec3fd-8b3e-4d59-b8ac-ba03e12acd23 · outbound

This paper cites Higgs Audio V2: Redefining Expressiveness in Au- dio Generation.

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark Higgs Audio V2: Redefining Expressiveness in Au- dio Generation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:36:04.314988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:09:21.027609Z digest=sha256:18cc79fd8c2ab4d204b0ca997592b446e8498d08617115905918e6bebb383702

Observation 21d564f0-e726-4d35-bd93-13ab8c870b96 · outbound

This paper cites Phonology, phonetics, and signal-extrinsic factors in the perception of prosodic prominence: Evidence from rapid prosody transcription.

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark Phonology, phonetics, and signal-extrinsic factors in the perception of prosodic prominence: Evidence from rapid prosody transcription

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:36:04.312119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:09:21.027609Z digest=sha256:f0b00629b1a01c5943a140c4b23420b121345d585297894f6a471e4f6acc5bb4

Pith citing papers

Observation e1449494-7048-4ae3-987f-528810084254 · inbound

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark cites this paper.

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T11:06:05.434356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:09:21.027609Z digest=sha256:b6438d123ff164d7b629713c1c05d75d820fee0e73a0b814aa29bc592d83fd3b