Pith. sign in

Paper Citation Record · LEDGER

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2506.16381.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16381 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T00:43:26.122470Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation af1056ca-f61a-48c9-af15-a2cdf23ec4e0 · inbound

Qwen3-TTS Technical Report cites this paper.

Qwen3-TTS Technical Report InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:24:56.122236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T19:24:56.057631Z digest=sha256:74ad388b8e0d53e1d5c1f85eec6b30e2aeed13c4db3329203adee15d8a784c4d

Observation d868c2ba-6f12-4b3a-8397-71b4977185e2 · inbound

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation cites this paper.

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:16:00.213627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T17:15:28.918204Z digest=sha256:2067a94eb7653b321c144c9abe1bda9a6dd5f359c5d65e6ce187b7508d0d35bf

Observation 88bbdbc7-9e49-4a19-a837-94f66f121876 · inbound

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations cites this paper.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:47:12.979580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:85f2b3cf8731e21affba4ce869d3bfdf91a0641db4b7602cf4e4b6aa824afeb0

Observation e78d9862-8b72-4ad6-bed9-2b81c554d029 · inbound

MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech cites this paper.

MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T04:04:47.258265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:03:38.919545Z digest=sha256:fe331751c494022bd28d474f2d79cc1f168bddf7c79e805a70bc2d225c297b76

Observation 19987e2b-a1cb-44cc-b390-ade3ae5f29da · inbound

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech cites this paper.

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:19:03.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-20T21:14:58.814362Z digest=sha256:ae44a165a3abf3fdc28c9d8bbea1085bf4ed23414cef420fbf37cbe850cb184e

Observation b976e3ac-df8c-4e34-89b3-0513ac085e4a · inbound

ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge cites this paper.

ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:00:05.346556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-25T22:07:10.531784Z digest=sha256:ac607145d10c105dc72057b2e0be801b659da8d068a45dc0f69466a41322d6e6

Observation db9b52e4-04f1-464d-9c3b-854d8b45cc3f · inbound

Is Natural Always Appropriate? Investigating Naturalness and Appropriateness Across Different Domains for TTS Evaluation cites this paper.

Is Natural Always Appropriate? Investigating Naturalness and Appropriateness Across Different Domains for TTS Evaluation InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:55:42.971946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T03:32:23.838961Z digest=sha256:98bbdeb2b2212d40aca8ac4c8e2f5427f2622e3acd2d4a5f9824a67274e5538c

Observation 2cc58875-3fa4-4603-9419-b2cabe24bced · inbound

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models cites this paper.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:03.452297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:03.452297Z digest=sha256:4706f72d8f21bc69792493a646f57423f2a4e59ca9eb2cb7cd5fbd03677e0dc7

Observation ce6e5120-c663-4fa3-805a-fd0067f1c597 · inbound

Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation cites this paper.

Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T05:08:04.544761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:08:04.544761Z digest=sha256:0c81dd3bc6b315e6739560d71cd4a13605db279de82c6a557f9802d129b291dd

Observation e2f04486-eb39-4b01-920a-b8217a575738 · inbound

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems cites this paper.

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:37.741292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:37.741292Z digest=sha256:88e60db666f93611b92b1b01b1339fd8973ffbb9485c665f0d8be074a504064c

Observation 5e39b400-d8f9-48ee-812f-f0de8a735222 · inbound

Beyond Prompt Adherence: Auditing Attribute-Level Voice Control in Speech Generation cites this paper.

Beyond Prompt Adherence: Auditing Attribute-Level Voice Control in Speech Generation InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T00:43:26.122470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:43:26.122470Z digest=sha256:5ebceaed60fd3ce6f5b9e7cc8a7496291fa792432d8f665340e09b332049205a

Observation 1e05fdea-1195-4406-989a-5ef9c53d79dd · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:22.277483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:22.277483Z digest=sha256:03d9318a2b5e644034bae859b8ffc24a33c25906610e786801d8534b70472139