Pith. sign in

Paper Citation Record · LEDGER

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese

As of 20 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 1 inbound Pith citation observation for arXiv:2505.11200.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11200 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:02:45.995503Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-08T12:03:11.243693Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T19:21:09.992559Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact4
  • verified fuzzy26
  • unresolved17
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d12e0eee-0575-4604-92f3-acec6972d62a · outbound

This paper cites The kendall rank correlation coefficient.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese The kendall rank correlation coefficient

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.916722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.452290Z digest=sha256:cc8685d5f3488970cdc8436c6f3f2dcbb782f3b100af671d54986d4eb51b0848

Observation 78e42f87-1c18-4a06-bbc6-b9bf84f4f46d · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.460507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.460507Z digest=sha256:fad3dca104225440ec8a5aafc5d9eb55cee62a342a96ee136d71870df4984f3c

Observation 49ab2a14-7b56-4a04-961c-c6bec3e765b2 · outbound

This paper cites The t05 system for the voicemos challenge 2024: Transfer learning from deep image classifier to naturalness mos prediction of high-quality synthetic speech.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese The t05 system for the voicemos challenge 2024: Transfer learning from deep image classifier to naturalness mos prediction of high-quality synthetic speech

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.889436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.474393Z digest=sha256:ec5af6d8e8b0d5850bee99e7e81629e3a5388685ed6da8b1c0537ed5baabd4b3

Observation 27821350-9168-49bc-855e-bdb93b82c704 · outbound

This paper cites Generalized linear mixed models: a practical guide for ecology and evolution.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Generalized linear mixed models: a practical guide for ecology and evolution

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.854107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.485384Z digest=sha256:66947407773729f36dc4c2d9ee914ab61aecef1f457bd489525d5160ca0336e9

Observation 2feb102f-9eb0-4a89-9a40-03be0124b9db · outbound

This paper cites Why we should report the details in subjective evaluation of tts more rigorously.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Why we should report the details in subjective evaluation of tts more rigorously

Reference 5

Resolution
verified exact
doi, observed 2026-08-15T21:02:46.147655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.499141Z digest=sha256:1ebe32eae3d9b31b353c7f859781b13a16ea3d441a0a56530cfdfe1eb1b57ff6

Observation 34a5caeb-448c-4616-b27d-41c5bec6b7c2 · outbound

This paper cites Qwen2-Audio Technical Report.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Qwen2-Audio Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.507556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.507556Z digest=sha256:4f7727a46bf86f7446402768a97816dd9ba22d8a85a6c2ab76504b59374bb20b

Observation 54564007-9d02-4775-8abd-164db0a99811 · outbound

This paper cites an unresolved cited work.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.519126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.519126Z digest=sha256:67092c03cd233a76cddf06eabb945e4e6b4515abd3fc953320efaba161a936ef

Observation bd94a023-0738-4b32-85ff-82892eab48c6 · outbound

This paper cites Disambiguation of Chinese Polyphones in an End-to-End Framework with Semantic Features Extracted by Pre-trained BERT.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Disambiguation of Chinese Polyphones in an End-to-End Framework with Semantic Features Extracted by Pre-trained BERT

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.527668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.527668Z digest=sha256:5a2102b5674021a15ef96640b972363730d4300aabaf8748bb0f348e37e3da53

Observation c987edd5-03fb-4ff8-bef5-4dd278e08814 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.536218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.536218Z digest=sha256:6a077de30a6849b90ddeec020618f47362943b453e6f09afadcf49f1435eb75a

Observation 939b2ce8-a515-4760-b107-dd68c9163cea · outbound

This paper cites Assessing the impact of contextual framing on subjective tts quality.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Assessing the impact of contextual framing on subjective tts quality

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.824182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.544926Z digest=sha256:ba99a38695376c4b438b98ae5b7e18a8d1098fca3b554f3fb71cd622a24c4f2c

Observation 075d394f-af28-4007-8479-c4a61a262681 · outbound

This paper cites The turing test: the first 50 years.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese The turing test: the first 50 years

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.789140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.556597Z digest=sha256:9f3f754f2bbcd6f9319814c61241e48ca43732f707dd886d0f191081745eec43

Observation 0bcd6c3c-53d4-4c3a-ae1a-db140529cb68 · outbound

This paper cites Analysis of speaker similarity in the statistical speech synthesis systems using a hybrid approach.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Analysis of speaker similarity in the statistical speech synthesis systems using a hybrid approach

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.754777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.566322Z digest=sha256:c7c768b25a7ec5a8c0189d500292fcb28a285968511af9bf8aa970bcd695c2c2

Observation 3f9bdd01-ed8a-49c2-83cc-865a0096be17 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.572675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.572675Z digest=sha256:5ba7f5859206e6496bcfe4bc355ca77697ce54044a60aff674df563f69a9e7b3

Observation 838d049d-4211-44f4-8e20-80ca2dd798ed · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Lora: Low-rank adaptation of large language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.715922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.582227Z digest=sha256:4683f44c79d5e4365821af1d15c4b65509340c6da07ed7b63c4daf487f5df39a

Observation 0873eeb2-a2ab-40a1-97d6-4a8611358ce0 · outbound

This paper cites Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.591082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.591082Z digest=sha256:66008b406c3c708fa9346eeb662a9465f4536268d599801097fc0242ab7e3cbf

Observation a2c57a22-089a-4124-a412-99066bacda59 · outbound

This paper cites Mm algorithms for generalized bradley-terry models.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Mm algorithms for generalized bradley-terry models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.599714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.599714Z digest=sha256:5063f4884e126a35acc771b120319c30c00926dbaf000647029e6b75ebb777cf

Observation fc5ef76a-f15e-4309-ab88-086bec1649de · outbound

This paper cites GPT-4o System Card.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese GPT-4o System Card

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.608954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.608954Z digest=sha256:f0b35258a8b8a22f52b2d85c63c1b33d6e529a9ec3d4609681eafcb9e502ed76

Observation c87f2b6c-f133-4c6f-9ddb-07c0a1ee6fa1 · outbound

This paper cites Method for the Subjective Assessment of Intermediate Quality Level of Audio Systems, 2015.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Method for the Subjective Assessment of Intermediate Quality Level of Audio Systems, 2015

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.634139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.617482Z digest=sha256:80464687d22d4b1b23b183b7e2ebb78592a37d96db1593266773abcb9da855ea

Observation 71107992-3aec-49a1-a264-a01315bd3f0b · outbound

This paper cites Subjective evaluation of speech quality with a crowd- sourcing approach, 2018.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Subjective evaluation of speech quality with a crowd- sourcing approach, 2018

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.609343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.628694Z digest=sha256:44243214db16c4eaf2100cc4f5e06823749b0938df832965597abc9b2089decb

Observation 8918ce2d-48dc-44f8-83a0-fb5157157b10 · outbound

This paper cites Compact Neural TTS Voices for Accessibility.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Compact Neural TTS Voices for Accessibility

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:02:46.619507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.639186Z digest=sha256:bf988b2df53f80ea3df407872eb877a17e1202dfba825898341bb2504f45bb06

Observation 4fe8b601-769d-4554-adff-150867f3c357 · outbound

This paper cites Stuck in the mos pit: A critical analysis of mos test methodology in tts evaluation.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Stuck in the mos pit: A critical analysis of mos test methodology in tts evaluation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.569430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.671587Z digest=sha256:88bff46ce0c607f075413aa11d9d0e65a94aad426e72b3bdd18bbfd97b3a4887

Observation 78dc6c6c-af68-4195-ae55-065e6914a91e · outbound

This paper cites Issues in chinese prosody: conceptual foundations of a linguistically-motivated text-to-speech system for mandarin.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Issues in chinese prosody: conceptual foundations of a linguistically-motivated text-to-speech system for mandarin

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.540103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.679355Z digest=sha256:d736c409ef908d76a6ea4adc0e64dfafac15558e3d8cfc479775439683ac8691

Observation 1ddbb4e4-8785-4c2b-9de7-9bd1c2c331a4 · outbound

This paper cites The limits of the mean opinion score for speech synthesis evaluation.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese The limits of the mean opinion score for speech synthesis evaluation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.496958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.686752Z digest=sha256:b18a78b9ef61df1210c2e43a4b987d4498b841432f8e1487ecf1c674f1611ca3

Observation 25e06993-00be-4795-9ff1-7e7478a64e9e · outbound

This paper cites StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.723124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.723124Z digest=sha256:fc44d03e37d6c2c27e6c849bc90d8a78e36a4b626cb0dadf23b91408cd6c99fb

Observation ee85daee-7f91-4448-bb93-d3020af5c143 · outbound

This paper cites Hyper-realistic, multi-emotion generative speech model speech-01.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Hyper-realistic, multi-emotion generative speech model speech-01

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.462948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.736670Z digest=sha256:50b17e8605a5f6f9ad0dbe290e7f5aa932501c69e0d537ba0ce5578665361bbd

Observation 23f73e61-8485-4b1d-b7a4-b63c8ebd9b16 · outbound

This paper cites Speech Quality Assessment in Crowdsourcing: Comparison Category Rating Method.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Speech Quality Assessment in Crowdsourcing: Comparison Category Rating Method

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:02:46.404802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.765487Z digest=sha256:348ecd7bd728a65f28e726982dc70ade3b4211ff89f76aff18ad505c8c3ee9e1

Observation 653dc60e-7403-48af-a90e-c356fc618f18 · outbound

This paper cites The blizzard challenge.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese The blizzard challenge

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.427943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.773130Z digest=sha256:571ac415d28e758dfa88aba26a99a99f1e7ac5ffa20a3025f09944b2b5b5f3b0

Observation 5ecee893-2d36-4eca-a73a-a04941c003c4 · outbound

This paper cites Dnsmos p.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Dnsmos p

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.401200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.796723Z digest=sha256:976b6605b8710901118f8e767d7ea0d207e3a815948c9f339d3007f334628b8c

Observation a8de6606-9795-4a78-bbc8-7041e0f01c01 · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.817444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.817444Z digest=sha256:bc990604f60536bd7b9e89e2dab8d100bfc0e5c7da07acff2974f3d5a53772e8

Observation 7ec24e69-2d0f-447d-9579-242ec86fb6cf · outbound

This paper cites Mean opinion score (mos) revisited: methods and applications, limitations and alternatives.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Mean opinion score (mos) revisited: methods and applications, limitations and alternatives

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.357552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.827440Z digest=sha256:e9dab0b94b3a907570126c37947feac17d4cd1b591337eaefb83c19bd627fc68

Observation 2d921818-f727-435a-b522-005c737f94ca · outbound

This paper cites Rethinking MUSHRA: Addressing Modern Challenges in Text-to-Speech Evaluation.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Rethinking MUSHRA: Addressing Modern Challenges in Text-to-Speech Evaluation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.838352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.838352Z digest=sha256:36378b541a165d2dd1840de4246b047ab9b3805cc8829fd29cb108bed0b89399

Observation eca95645-6b86-4f9c-8307-7e66d4703efa · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.850786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.850786Z digest=sha256:8518f601cde4c9658f0f3c47260cedbcad214a57ebc707b89f4d7bad5fda15ac

Observation ed67bc23-3e89-4bb0-bb63-f9fcfed3ee86 · outbound

This paper cites Contextual interactive evaluation of tts models in dialogue systems.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Contextual interactive evaluation of tts models in dialogue systems

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.320286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.859506Z digest=sha256:d18ec80e6f244bc842246383b21230feb86e990fb32f024c9db341973d4af38d

Observation 320ec6cd-47ca-4755-b7f4-245507619faa · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.866945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.866945Z digest=sha256:aae424144b0f98f7a176d20a6c736cc42a0b7ca146d5c0949e97df1caf5f0e59

Observation dfbe0640-77c6-4aff-81a3-0eb3fbfc9e74 · outbound

This paper cites Bilingual and code- switching tts enhanced with denoising diffusion model and gan.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Bilingual and code- switching tts enhanced with denoising diffusion model and gan

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.271146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.875017Z digest=sha256:b8e3f2e27e7626c392b83ec5c5434ffac6736115f84c4520493fa8ba8b309e36

Observation 7d5914c6-f227-43b5-876d-6e43de4d5959 · outbound

This paper cites Talk2care: An llm-based voice assistant for communication between healthcare providers and older adults.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Talk2care: An llm-based voice assistant for communication between healthcare providers and older adults

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.883992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.883992Z digest=sha256:57a20ab8a40eefc13a337f4bb72db5d185699349a3996a55f8d873d6cfc7ed7e

Observation 0e0df59f-fc67-4124-91d2-ab4338032a0c · outbound

This paper cites Dialog modeling in audiobook synthesis.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Dialog modeling in audiobook synthesis

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.228324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.896901Z digest=sha256:578da66ac58bf1b24dc09c80c543a11729662d087b95850cdd7349cd687265c4

Observation 730a5761-7d6e-48b7-b7f6-fc5a6f450973 · outbound

This paper cites orders of magnitude larger.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese orders of magnitude larger

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.196380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.906044Z digest=sha256:2d097c3b9730963a7a2c38538caac8513e76901495f596089b8a9815e556babc

Observation cdbc3f11-774e-4e02-b821-670445f8fc91 · outbound

This paper cites Pure machine voice.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Pure machine voice

Reference 41

Resolution
parse uncertain
raw_fallback, observed 2026-08-15T21:02:47.161308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.913721Z digest=sha256:d2f24d7df86d164dbe27c320ff60fc36de485a2c2249101614dab983159decc1

Observation a6cc27a3-e201-445b-8c9d-6e51f2a837f3 · outbound

This paper cites The i m i t a t i o n of human speech is too forced.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese The i m i t a t i o n of human speech is too forced

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.135018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.927549Z digest=sha256:a90dc27bcf8016414f1c267d2ddcbc71d44374e3bfa9fdb73eb7c9a7f5095baf

Observation 64a8d3f5-6982-4628-b4c9-dae1e669ed74 · outbound

This paper cites O b v i o u s l y a machine tone - doesn ’ t sound like a real person.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese O b v i o u s l y a machine tone - doesn ’ t sound like a real person

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.098970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.934317Z digest=sha256:fdffbe731c13361035e5e7de3ea68191440bb8744dc086470d8173c8a168b677

Observation da151159-e479-4e0f-999d-208f638af4c1 · outbound

This paper cites Sounds like a late - night radio host.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Sounds like a late - night radio host

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.071983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.948407Z digest=sha256:64c44c96c29d489c943e876bdbdbef312118e4f6044de60f6733cfabe2cf7759

Observation ac974039-4221-4c92-afeb-5000f6d26a52 · outbound

This paper cites Many thins.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Many thins

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.047847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.960734Z digest=sha256:903f07378b8bac46cd6d08260ebbaec4f3ecee71ebb3162fe2656daadf4c63a0

Observation cc2d41db-d096-4f5e-b35b-038e7cb39e49 · outbound

This paper cites an unresolved cited work.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:02:46.992756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.971102Z digest=sha256:bf9f66624c3d18307db825c49afc1ad437a85693a02c078641480a4005168f3e

Observation 69e6cc48-396f-48e5-ae18-3bec302d2417 · outbound

This paper cites go away.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese go away

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:46.948182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.978046Z digest=sha256:aab452966e19924948c591a3a68a17ce0c4f678cc7c4950be8167068faaa7edb

Observation 730e009a-eede-46dd-b78a-3a67896de368 · outbound

This paper cites angry ,.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese angry ,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:46.921239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.995503Z digest=sha256:a7db84544a446f0516af49326c0109992e1d316e515134148beb9c07f7bd4e77

Observation 45f9dab7-753f-4a68-a415-96c670f0e988 · outbound

This paper cites doi: 10.21437/Blizzard.2023-1.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese doi: 10.21437/Blizzard.2023-1

Reference 2023

Resolution
verified exact
doi, observed 2026-08-15T21:02:46.092266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:02:45.789007Z digest=sha256:d1dbe62423aab77652692ff3bad6ee09507df02dad3d77301e30550e9f47e08b

Observation 1aed6274-3448-4c8d-afe2-64eb5efbe64b · outbound

This paper cites URL https://www.sciencedirect.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese URL https://www.sciencedirect

Reference 2308

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.696157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.696157Z digest=sha256:08f5efbcca5079d8d1472e7fcf1d299f1a8c8a4daa7ce6e5a84ecce1fdd8ab38

Pith citing papers

Observation d1e43662-3b5f-4805-b13a-e882f93a10a2 · inbound

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis cites this paper.

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:09.999306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:03:11.243693Z digest=sha256:88fe9a0253db3db57e5fc8d2248289af24c78bd71a17d2ba54015417ea08e0c7