Pith. sign in

Paper Citation Record · LEDGER

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese

As of 23 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 1 inbound Pith citation observation for arXiv:2505.11200.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11200 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:02:45.995503Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-08T12:03:11.243693Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T19:21:09.992559Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact4
  • verified fuzzy26
  • unresolved17
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d12e0eee-0575-4604-92f3-acec6972d62a · outbound

This paper cites The kendall rank correlation coefficient.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese The kendall rank correlation coefficient

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.916722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.452290Z digest=sha256:88ecd652ef4e8a64db85c885173e5db31ed696c118a2fff74e5c10509265036c

Observation 78e42f87-1c18-4a06-bbc6-b9bf84f4f46d · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.460507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.460507Z digest=sha256:07e7a9b5def694c718502f52118c660128ef33a4b5e59e5c5f205b0b415b1699

Observation 49ab2a14-7b56-4a04-961c-c6bec3e765b2 · outbound

This paper cites The t05 system for the voicemos challenge 2024: Transfer learning from deep image classifier to naturalness mos prediction of high-quality synthetic speech.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese The t05 system for the voicemos challenge 2024: Transfer learning from deep image classifier to naturalness mos prediction of high-quality synthetic speech

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.889436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.474393Z digest=sha256:dc61c46c06bcab54dcb2dc830215d5d95d65cb869b19bed2495b708d53f9a2c8

Observation 27821350-9168-49bc-855e-bdb93b82c704 · outbound

This paper cites Generalized linear mixed models: a practical guide for ecology and evolution.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Generalized linear mixed models: a practical guide for ecology and evolution

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.854107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.485384Z digest=sha256:5033d3c46ba8eb185f04ed4a018787518e950cb51033dd36c967edff9eb864a7

Observation 2feb102f-9eb0-4a89-9a40-03be0124b9db · outbound

This paper cites Why we should report the details in subjective evaluation of tts more rigorously.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Why we should report the details in subjective evaluation of tts more rigorously

Reference 5

Resolution
verified exact
doi, observed 2026-08-15T21:02:46.147655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.499141Z digest=sha256:e842ebba26f04f562cef67e930ade8b79f6cb9e6463e2bf5376293e6d5a89d60

Observation 34a5caeb-448c-4616-b27d-41c5bec6b7c2 · outbound

This paper cites Qwen2-Audio Technical Report.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Qwen2-Audio Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.507556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.507556Z digest=sha256:459e953f105624bedae4841e9eb540dfd965c8015ace1a7a4e07df6f915efe74

Observation 54564007-9d02-4775-8abd-164db0a99811 · outbound

This paper cites an unresolved cited work.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.519126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.519126Z digest=sha256:30b72c2f4c8003030d5b2c653f0aede6fdc57a14efd09d20a98257cf45048eef

Observation bd94a023-0738-4b32-85ff-82892eab48c6 · outbound

This paper cites Disambiguation of Chinese Polyphones in an End-to-End Framework with Semantic Features Extracted by Pre-trained BERT.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Disambiguation of Chinese Polyphones in an End-to-End Framework with Semantic Features Extracted by Pre-trained BERT

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.527668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.527668Z digest=sha256:96bfa8cc8ac7b71af5321f3f8722869b10415bb1606677ea3ad490fd56eeca10

Observation c987edd5-03fb-4ff8-bef5-4dd278e08814 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.536218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.536218Z digest=sha256:1694708bb8da98dd33797e1c8f2cec68c14348593eed6ae1559592a6f155af38

Observation 939b2ce8-a515-4760-b107-dd68c9163cea · outbound

This paper cites Assessing the impact of contextual framing on subjective tts quality.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Assessing the impact of contextual framing on subjective tts quality

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.824182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.544926Z digest=sha256:9f359eecb0806230124e7cc7e93d425f7f0f34ea3c117865982d5ef0979ab761

Observation 075d394f-af28-4007-8479-c4a61a262681 · outbound

This paper cites The turing test: the first 50 years.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese The turing test: the first 50 years

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.789140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.556597Z digest=sha256:ae941752d92b5fca95cb3cddda502fbf3ff54448f9bc7402cc9b1f33b1a04f4d

Observation 0bcd6c3c-53d4-4c3a-ae1a-db140529cb68 · outbound

This paper cites Analysis of speaker similarity in the statistical speech synthesis systems using a hybrid approach.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Analysis of speaker similarity in the statistical speech synthesis systems using a hybrid approach

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.754777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.566322Z digest=sha256:e1baa9c8f155096fcc0c86205fadca89130a2facff19d88cb6a689f96330e568

Observation 3f9bdd01-ed8a-49c2-83cc-865a0096be17 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.572675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.572675Z digest=sha256:0c8e452169592b3651898c6ddc01a72caeb329ef733b11f6df7d0f43bb730d61

Observation 838d049d-4211-44f4-8e20-80ca2dd798ed · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Lora: Low-rank adaptation of large language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.715922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.582227Z digest=sha256:535db1b522ff0f6f56e31e87d731512f0f3625380775ddaf6abea178032aee40

Observation 0873eeb2-a2ab-40a1-97d6-4a8611358ce0 · outbound

This paper cites Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.591082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.591082Z digest=sha256:7b5c544a57ab5b873a802bc7d28c58cc8fcca2c203feee311db5df3337e91a3e

Observation a2c57a22-089a-4124-a412-99066bacda59 · outbound

This paper cites Mm algorithms for generalized bradley-terry models.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Mm algorithms for generalized bradley-terry models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.599714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.599714Z digest=sha256:ac7655bfe62a2235bbf280605732d089b8c03dfcee48cb7f61138fd495fc2b76

Observation fc5ef76a-f15e-4309-ab88-086bec1649de · outbound

This paper cites GPT-4o System Card.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese GPT-4o System Card

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.608954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.608954Z digest=sha256:12059468efb8a37e0732b93aedcad09a611e11ed1f10f8d39d5f0b0b1d404a5f

Observation c87f2b6c-f133-4c6f-9ddb-07c0a1ee6fa1 · outbound

This paper cites Method for the Subjective Assessment of Intermediate Quality Level of Audio Systems, 2015.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Method for the Subjective Assessment of Intermediate Quality Level of Audio Systems, 2015

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.634139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.617482Z digest=sha256:20580a6aefa873b4a9cbdec2ee44126f3173f1824dc6a35de5500b9740b71892

Observation 71107992-3aec-49a1-a264-a01315bd3f0b · outbound

This paper cites Subjective evaluation of speech quality with a crowd- sourcing approach, 2018.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Subjective evaluation of speech quality with a crowd- sourcing approach, 2018

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.609343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.628694Z digest=sha256:44adc5f40951f0045a6ee82dd2d57e445e9547fd6e71a1e7729c36defdd474c7

Observation 8918ce2d-48dc-44f8-83a0-fb5157157b10 · outbound

This paper cites Compact Neural TTS Voices for Accessibility.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Compact Neural TTS Voices for Accessibility

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:02:46.619507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.639186Z digest=sha256:81076d17a283adcefaa5748b65c1f90f30c942550629440c38a3cf088c433d39

Observation 4fe8b601-769d-4554-adff-150867f3c357 · outbound

This paper cites Stuck in the mos pit: A critical analysis of mos test methodology in tts evaluation.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Stuck in the mos pit: A critical analysis of mos test methodology in tts evaluation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.569430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.671587Z digest=sha256:59b884d6614f5d6bf8f8d5d942ab7acf9ceaeeaccff406a0d97fe8bb2c526cd9

Observation 78dc6c6c-af68-4195-ae55-065e6914a91e · outbound

This paper cites Issues in chinese prosody: conceptual foundations of a linguistically-motivated text-to-speech system for mandarin.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Issues in chinese prosody: conceptual foundations of a linguistically-motivated text-to-speech system for mandarin

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.540103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.679355Z digest=sha256:95f32fc43510c5fcb2f7c4c79a4308be5dc5a4bda69b069a9cea044249ce71eb

Observation 1ddbb4e4-8785-4c2b-9de7-9bd1c2c331a4 · outbound

This paper cites The limits of the mean opinion score for speech synthesis evaluation.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese The limits of the mean opinion score for speech synthesis evaluation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.496958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.686752Z digest=sha256:c17976dd95dd4dd4137fdae371d60f9dd4065db78a8b726fea2ad79d3068dbfe

Observation 25e06993-00be-4795-9ff1-7e7478a64e9e · outbound

This paper cites StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.723124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.723124Z digest=sha256:f45c354d46e5cc84ce2f269b0b34dbc3b6f81467f484cae686e2a54eb84a1a1e

Observation ee85daee-7f91-4448-bb93-d3020af5c143 · outbound

This paper cites Hyper-realistic, multi-emotion generative speech model speech-01.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Hyper-realistic, multi-emotion generative speech model speech-01

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.462948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.736670Z digest=sha256:1101b2e207891e90c9797c9eae22042574fc48398eb8f50f70f226cbcbecab2e

Observation 23f73e61-8485-4b1d-b7a4-b63c8ebd9b16 · outbound

This paper cites Speech Quality Assessment in Crowdsourcing: Comparison Category Rating Method.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Speech Quality Assessment in Crowdsourcing: Comparison Category Rating Method

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:02:46.404802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.765487Z digest=sha256:c44c917064b16e6a57e4feb9eb720ea0c0f93dfa4d4a3180aa61cf78c9b92352

Observation 653dc60e-7403-48af-a90e-c356fc618f18 · outbound

This paper cites The blizzard challenge.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese The blizzard challenge

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.427943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.773130Z digest=sha256:260d0272298ef46fc5b29f1b5456a7c42305647fa8202d1618567c94cf603e47

Observation 5ecee893-2d36-4eca-a73a-a04941c003c4 · outbound

This paper cites Dnsmos p.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Dnsmos p

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.401200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.796723Z digest=sha256:6e4647cfbc5d9b41d5a8deacbdb6ddb069ed3707a8532c8d1b454459731f283d

Observation a8de6606-9795-4a78-bbc8-7041e0f01c01 · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.817444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.817444Z digest=sha256:84b1430ebdc14c14652992ca48904a1a7697d83e57595e6b4d72f437d0137255

Observation 7ec24e69-2d0f-447d-9579-242ec86fb6cf · outbound

This paper cites Mean opinion score (mos) revisited: methods and applications, limitations and alternatives.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Mean opinion score (mos) revisited: methods and applications, limitations and alternatives

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.357552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.827440Z digest=sha256:ef3610c513f9db3a43946916fa1968e12fe82469d17b6c4855157b7f5fe4d890

Observation 2d921818-f727-435a-b522-005c737f94ca · outbound

This paper cites Rethinking MUSHRA: Addressing Modern Challenges in Text-to-Speech Evaluation.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Rethinking MUSHRA: Addressing Modern Challenges in Text-to-Speech Evaluation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.838352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.838352Z digest=sha256:1904611ec6f9e16c3e163dfd346cbff4d34f2ad1906c3aed318d9f14a4d7e607

Observation eca95645-6b86-4f9c-8307-7e66d4703efa · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.850786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.850786Z digest=sha256:2cf3fd62a53847fc7c40637ea4b30858e23085ec5829aff7a7961abba4af2b90

Observation ed67bc23-3e89-4bb0-bb63-f9fcfed3ee86 · outbound

This paper cites Contextual interactive evaluation of tts models in dialogue systems.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Contextual interactive evaluation of tts models in dialogue systems

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.320286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.859506Z digest=sha256:1e6a62d124412ceade2c0c2cf64147877cb1134aa05e7947887f2b198cfa6b19

Observation 320ec6cd-47ca-4755-b7f4-245507619faa · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.866945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.866945Z digest=sha256:63b47d19ec8ee16f94d351eb1a12c4a57305eea42ba686981bfd3a49e0e747b0

Observation dfbe0640-77c6-4aff-81a3-0eb3fbfc9e74 · outbound

This paper cites Bilingual and code- switching tts enhanced with denoising diffusion model and gan.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Bilingual and code- switching tts enhanced with denoising diffusion model and gan

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.271146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.875017Z digest=sha256:98238a5765c195eb93d8a32f3e8f6833aab0fea4cf18dd401569c97a95711b35

Observation 7d5914c6-f227-43b5-876d-6e43de4d5959 · outbound

This paper cites Talk2care: An llm-based voice assistant for communication between healthcare providers and older adults.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Talk2care: An llm-based voice assistant for communication between healthcare providers and older adults

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.883992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.883992Z digest=sha256:464a2d19e8c539b875ba1416af07737bc2e511eea3bf1d51af9f1460fad29612

Observation 0e0df59f-fc67-4124-91d2-ab4338032a0c · outbound

This paper cites Dialog modeling in audiobook synthesis.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Dialog modeling in audiobook synthesis

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.228324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.896901Z digest=sha256:d98ef1d8a3ce4f8f436cb103c4ffcb92a21e9d7a96ea144efd20a8c073be3ed6

Observation 730a5761-7d6e-48b7-b7f6-fc5a6f450973 · outbound

This paper cites orders of magnitude larger.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese orders of magnitude larger

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.196380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.906044Z digest=sha256:0dc1c001783459096c2a1f8351763454d9d243df965b8fb359a583cdd48058ff

Observation cdbc3f11-774e-4e02-b821-670445f8fc91 · outbound

This paper cites Pure machine voice.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Pure machine voice

Reference 41

Resolution
parse uncertain
raw_fallback, observed 2026-08-15T21:02:47.161308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.913721Z digest=sha256:bce5ee53e7ed86aef020a657daa6cb1b48aba06f31b162ea0632d9e9faf02e8b

Observation a6cc27a3-e201-445b-8c9d-6e51f2a837f3 · outbound

This paper cites The i m i t a t i o n of human speech is too forced.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese The i m i t a t i o n of human speech is too forced

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.135018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.927549Z digest=sha256:4dbc044eae06c865707017de5b1f90cb6cb6ddd49790336d6ad016fc61544906

Observation 64a8d3f5-6982-4628-b4c9-dae1e669ed74 · outbound

This paper cites O b v i o u s l y a machine tone - doesn ’ t sound like a real person.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese O b v i o u s l y a machine tone - doesn ’ t sound like a real person

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.098970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.934317Z digest=sha256:6099c317cd9aee600e7f372a9be98727f671055abc98a0bedfb4a0cbbc90dfa2

Observation da151159-e479-4e0f-999d-208f638af4c1 · outbound

This paper cites Sounds like a late - night radio host.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Sounds like a late - night radio host

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.071983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.948407Z digest=sha256:791394c071491bbde0f9771d9f71097ad9b32f0747eb8c6b68bc614254604df2

Observation ac974039-4221-4c92-afeb-5000f6d26a52 · outbound

This paper cites Many thins.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Many thins

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:47.047847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.960734Z digest=sha256:8c0fc687cc4f85fff52ef054a8100ef0ec4f9dfe1ada6edcd77ce2453a978585

Observation cc2d41db-d096-4f5e-b35b-038e7cb39e49 · outbound

This paper cites an unresolved cited work.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:02:46.992756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.971102Z digest=sha256:61775f322f2e921c6e7f73b25a85a786d3f623359f3bde0d9532a7a525649522

Observation 69e6cc48-396f-48e5-ae18-3bec302d2417 · outbound

This paper cites go away.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese go away

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:46.948182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.978046Z digest=sha256:9bcf53414382ff1d4ee23901bd29fe9f0f1c816684c15f971447400dd16c8d72

Observation 730e009a-eede-46dd-b78a-3a67896de368 · outbound

This paper cites angry ,.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese angry ,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:02:46.921239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.995503Z digest=sha256:273397e502e5898dccc17a3040897f2abdd8e2eab14441e51e5bebd4d7e093b1

Observation 45f9dab7-753f-4a68-a415-96c670f0e988 · outbound

This paper cites doi: 10.21437/Blizzard.2023-1.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese doi: 10.21437/Blizzard.2023-1

Reference 2023

Resolution
verified exact
doi, observed 2026-08-15T21:02:46.092266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:02:45.789007Z digest=sha256:ca4c8bb2f10ecc7dc89dd68d5b32a010eebf3ba472dad0169aff66d2fcbd1086

Observation 1aed6274-3448-4c8d-afe2-64eb5efbe64b · outbound

This paper cites URL https://www.sciencedirect.

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese URL https://www.sciencedirect

Reference 2308

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:45.696157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:45.696157Z digest=sha256:d53f838725f3e58394a67d6f533ecf95af30ec83db38546265d4402da682e073

Pith citing papers

Observation d1e43662-3b5f-4805-b13a-e882f93a10a2 · inbound

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis cites this paper.

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:09.999306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T12:03:11.243693Z digest=sha256:5f3a5470bd0a3b7647d72a7deb5ff068af7958271045d406933615537e3ed0e5