Pith. sign in

Paper Citation Record · LEDGER

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

As of 22 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 13 inbound Pith citation observations for arXiv:2506.16381.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16381 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:32:31.021219Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:14:43.938774Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 1c4c5413-a55b-420f-8993-e950f3df6c01 · outbound

This paper cites Natural language guidance of high-fidelity text-to-speech with synthetic annotations.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:32:30.935359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:32:30.935359Z digest=sha256:e62f58be93c9f54e0aba8777b57660fc65f7a2ea51334ce6ee101e534545b71b

Observation 91eef112-d517-4018-9592-b53ce9e3043a · outbound

This paper cites LLM Evaluators Recognize and Favor Their Own Generations.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems LLM Evaluators Recognize and Favor Their Own Generations

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:32:30.939918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:32:30.939918Z digest=sha256:1e31eaae017dcd9425550e1dbc058fb4d875cea80dcabc5720b0e95ad5aa8099

Observation d03331a2-b034-4e69-b52b-507ea25d091f · outbound

This paper cites when saying …, raise your voice.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems when saying …, raise your voice

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:32:31.254377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:32:30.971061Z digest=sha256:d615d292dc96b721d85e3b3d5e8bc80db48f6ac09a15f6a99afeda6fa0b51374

Observation 5f200a37-6991-4fba-bf76-442b572898eb · outbound

This paper cites like,” “imagine,.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems like,” “imagine,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:32:31.185406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:32:31.001543Z digest=sha256:aa61c95331da8c5afde18844d97e18907919eaf0ae53cbf574bc3cbb702b0289

Observation 0db05b81-2181-44f9-8ddd-52823d3163cf · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:32:30.952653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:32:30.952653Z digest=sha256:ffd2832cac37838ae26c5a7a1ff8e13c266edb28f59844a2be57b49815956775

Observation e0b56e1c-db55-4849-9e07-15ac14b7bf04 · outbound

This paper cites high female pitch.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems high female pitch

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:32:31.301900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:32:30.956652Z digest=sha256:bc6b9d81b23af6a092041cd88ed65365dcd3ab7fe7cd1d842add46ed1df2268a

Observation f06d8854-60c0-45f2-9703-7a8dc91a2b40 · outbound

This paper cites hoarse,” “furious.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems hoarse,” “furious

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:32:31.289749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:32:30.974641Z digest=sha256:e0d6be1013927353b2010295d9cbf5a7e21f2117bfcaf682b1c3982020321035

Observation 3e853945-9f3a-45b8-bd27-fc6d7bd45681 · outbound

This paper cites • Use a rich palette of synonyms and idioms; employ different grammatical structures (e.g.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems • Use a rich palette of synonyms and idioms; employ different grammatical structures (e.g

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:32:31.278001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:32:30.978278Z digest=sha256:97dee74a4f0ad6876606b571224a350cc16a21ab7fff7a1054dfa517c82840d1

Observation 08a20cb0-8be8-44d6-86a1-0022dd76d1c2 · outbound

This paper cites Imagine,.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems Imagine,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:32:31.265910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:32:30.981877Z digest=sha256:d5ed306b38726aa755191c6267d5eb6b8054116eb94dc2abaf66671f700ce294

Observation 4a11d729-592a-4dd6-ae7a-f2bbe45bfe67 · outbound

This paper cites when saying …, raise your voice.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems when saying …, raise your voice

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:32:31.242921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:32:30.985106Z digest=sha256:8cd98cbfd2a5cc3bbbf40b39c5a07514bc52990a6e104ae0b2d0ee4e39d13bc3

Observation 301e4e57-180f-4a97-b8b9-4d01c89b3a7c · outbound

This paper cites an unresolved cited work.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:32:31.232396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:32:30.988091Z digest=sha256:5d3c3564882b40e45fc10a30b374e9a94bf530f638d0a77d2fc80fd88415282d

Observation 7a02cb83-833e-472c-80e7-bfd981841e77 · outbound

This paper cites an unresolved cited work.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:32:31.221501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:32:30.991344Z digest=sha256:a956a943c19cf635b40933f006edd11dcb058c9c93dd31832dd2cc42d92f8365

Observation efc7ce45-d0a9-41f6-8a0a-b410d502896c · outbound

This paper cites an unresolved cited work.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:32:31.209734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:32:30.994683Z digest=sha256:64803b350259a64e4762f9cd1d497c2350c67ca4a17082268ddf6b925224b7ca

Observation 7675d19b-e74c-4928-bfc7-f8ed52a14284 · outbound

This paper cites an unresolved cited work.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:32:31.196676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:32:30.997735Z digest=sha256:91f379af25ec972c43276a19d1087f1490f209f731569936cfb5a40d4d2bd126

Observation a05a3d84-339f-4ea6-877c-c13a9104167e · outbound

This paper cites general,.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems general,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:32:31.171745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:32:31.005980Z digest=sha256:642345437f8d71ec80bc46222faea1c4bd4b737df55b7317f33998a128e8f2f4

Observation c9d1562b-3952-48f5-8daa-02443c1c74bc · outbound

This paper cites Character + a single minimal action or speaking manner.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems Character + a single minimal action or speaking manner

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:32:31.159846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:32:31.009393Z digest=sha256:0cee5acf10f318ff13642767a9befd46307bab2b6260ef8460ab1da73f6752d4

Observation 211fd641-cfb0-489b-b32a-60e1d9dbb735 · outbound

This paper cites an unresolved cited work.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:32:31.146783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:32:31.012990Z digest=sha256:4e353c637728de292b6d2e1201e90d10f7cdea92772dd1b9393851adc8c893f9

Observation 3bb0264f-b38d-42c3-a195-c929f1072b7c · outbound

This paper cites All three instructions must differ and must not start the same way.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems All three instructions must differ and must not start the same way

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:32:31.134944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:32:31.016944Z digest=sha256:101560ad97d59f3a3a4e5f7c0c03c283380319d4a5a373b6abef91e846958e14

Observation 3ac62693-adb2-4219-b1ee-9ac68704f33b · outbound

This paper cites female high voice,.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems female high voice,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:32:31.123074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:32:31.021219Z digest=sha256:3509680803c49ece663c9951fc784087580677e32811ba3d5e84355a040d744c

Observation d5e5367b-6e78-4ca4-b051-14f110120e45 · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T19:32:30.944300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:32:30.944300Z digest=sha256:938100b01408b2d2910da5e8fdbc75434eab036dbd6dd9b132d698fbc23e6eae

Observation 266c7ed0-129f-4467-bd76-66023d4bf56f · outbound

This paper cites PromptTTS 2: Describing and Generating Voices with Text Prompt.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems PromptTTS 2: Describing and Generating Voices with Text Prompt

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T19:32:30.930017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:32:30.930017Z digest=sha256:0dc33412f0effa50c749f95b3d684be918b848cd8156831ad15d1b14d01d0eaf

Observation dd1aede6-4060-4a96-bf98-5a7372a53934 · outbound

This paper cites an unresolved cited work.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:32:31.316876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:32:30.948787Z digest=sha256:99e372ecb549afe863e98ee1bc3f08239b0788775655612f9dcf16fe9cd59c2b

Pith citing papers

Observation af1056ca-f61a-48c9-af15-a2cdf23ec4e0 · inbound

Qwen3-TTS Technical Report cites this paper.

Qwen3-TTS Technical Report InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:24:56.122236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T19:24:56.057631Z digest=sha256:93993a549803a1dc52f70d47241b7355842ffbf7b2d38f1831669e2c032cb280

Observation d868c2ba-6f12-4b3a-8397-71b4977185e2 · inbound

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation cites this paper.

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:16:00.213627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T17:15:28.918204Z digest=sha256:9dcc1d01b9805362521ff6e4c0c1902b46339e3e13016932f00bfee2722ff777

Observation 88bbdbc7-9e49-4a19-a837-94f66f121876 · inbound

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations cites this paper.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:47:12.979580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:ee9412ace182ef3f852c8dbfcd983a0f520694da488926fa9d3b90889ca842f1

Observation e78d9862-8b72-4ad6-bed9-2b81c554d029 · inbound

MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech cites this paper.

MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T04:04:47.258265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T04:03:38.919545Z digest=sha256:be2263f50b707daf17121d8aef3dc3521d6550964e9bdf426e35b216d0158b78

Observation 19987e2b-a1cb-44cc-b390-ade3ae5f29da · inbound

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech cites this paper.

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:19:03.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T21:14:58.814362Z digest=sha256:2753b54e716dd82cc2cf01b0c0a12d34bdf0000734c2685a72740428ea3b5ad5

Observation b976e3ac-df8c-4e34-89b3-0513ac085e4a · inbound

ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge cites this paper.

ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:00:05.346556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-25T22:07:10.531784Z digest=sha256:d7301e1ff9a5b872009ba569e68ccb5b2d5a3d0665e0c737131b9906863f0b0c

Observation db9b52e4-04f1-464d-9c3b-854d8b45cc3f · inbound

Is Natural Always Appropriate? Investigating Naturalness and Appropriateness Across Different Domains for TTS Evaluation cites this paper.

Is Natural Always Appropriate? Investigating Naturalness and Appropriateness Across Different Domains for TTS Evaluation InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:55:42.971946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-01T03:32:23.838961Z digest=sha256:f5c02cbf3e2a841c793121cf5e213d2fe8d4be9873fc48cf079d5b7a1f983f09

Observation 2cc58875-3fa4-4603-9419-b2cabe24bced · inbound

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models cites this paper.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:03.452297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:03.452297Z digest=sha256:d9cec48ca160633dfdbcbddfbb606a65a06e9c907704e5d488ae7d499c7d3e05

Observation ce6e5120-c663-4fa3-805a-fd0067f1c597 · inbound

Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation cites this paper.

Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T05:08:04.544761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:08:04.544761Z digest=sha256:46bec8caf3bfb00cbb7846ab70efbebe947667148018251ce2d007f5e68e5825

Observation e2f04486-eb39-4b01-920a-b8217a575738 · inbound

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems cites this paper.

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:37.741292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:37.741292Z digest=sha256:c5df7f7e08fdd357e289e247ceb3db448530ef9a0cf890331c25090d9baf4bc4

Observation 5e39b400-d8f9-48ee-812f-f0de8a735222 · inbound

Beyond Prompt Adherence: Auditing Attribute-Level Voice Control in Speech Generation cites this paper.

Beyond Prompt Adherence: Auditing Attribute-Level Voice Control in Speech Generation InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T00:43:26.122470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:43:26.122470Z digest=sha256:de2ac8674d66fb6052d4dd82d6d0703ab65ed44058d82a0ab05faaff22d44b62

Observation 1e05fdea-1195-4406-989a-5ef9c53d79dd · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:22.277483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:22.277483Z digest=sha256:f31c341fa884b475353bdf867adc98328675b2cf60320096e7a41f0fee325cc7

Observation 8ad53429-5f47-4f33-9b3a-5a55b1bb9063 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:43.938774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:43.938774Z digest=sha256:893bc90efc457778db82c937bba11f8172858cd1f0fd2953fe5692abbb1a31a3