Pith. sign in

Paper Citation Record · LEDGER

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations

As of 12 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 3 inbound Pith citation observations for arXiv:2604.16211.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.16211 v2

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T07:37:44.393592Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-03T00:29:39.365107Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T00:37:29.602606Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact18
  • verified fuzzy23
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation feabe402-ad24-4795-96dc-7bdb0f956477 · outbound

This paper cites NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-10T07:47:12.998128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:6a9f23b093da4e13a08e69680202a76aa1ae1fb141479bfcbf1091f128b0ab40

Observation 6b538caa-07c4-45ff-a6ae-0b99b6162072 · outbound

This paper cites he might cough.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations he might cough

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:47:12.995342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:8f4e4d6991d70c58eee16e8623f0868f0a672d46f37f5d7891601545968957b8

Observation 27a5d746-ae0c-419a-96c4-2d0f50078ed0 · outbound

This paper cites We benchmark 15 TTS systems, including 7 prompt-based and 8 tag-based systems, offering a diverse evalu- ation spectrum.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations We benchmark 15 TTS systems, including 7 prompt-based and 8 tag-based systems, offering a diverse evalu- ation spectrum

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:04:11.489824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:7e496215afe61ee426e657817faacc7228caae000bc5c0e7790ad1406272ff71

Observation 1c9e028c-ee24-4c0e-8859-0b8e9adcb724 · outbound

This paper cites hallucinate.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations hallucinate

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:04:11.487306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:7924d34ba21200b63b81855cf732f0ef7304faf2d080d627b6de2abc0707fc4e

Observation 0d710bc8-9dd7-4af6-8160-3aedcd3979f1 · outbound

This paper cites an unresolved cited work.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-21T13:04:11.447477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:4c779c3312c0e8cf6ba74d3c03e3328cc46dff359be1309ce35df7d957532eff

Observation 3a4da3e1-523a-4da2-a628-e8578b8e794e · outbound

This paper cites First, LLMs were used in benchmark dataset generation to draft candidate texts and speech captions, which were then reviewed, filtered, and finalized by the authors.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations First, LLMs were used in benchmark dataset generation to draft candidate texts and speech captions, which were then reviewed, filtered, and finalized by the authors

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:04:11.450312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:73f7124cf5b745969772926c46a7a3f9303df6bf51c3fe42fa9834a032a19419

Observation 04bd37d6-f0a5-4f3e-8662-e90d85a359e8 · outbound

This paper cites Recent advances in speech language models: A sur- vey.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Recent advances in speech language models: A sur- vey

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:04:11.452804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:7f45280a7f6c1e5cd159bdb8325166c6002bb87ce6e5207c0a450bbe15ddf1ac

Observation 7759d619-2cfb-40c0-9a1f-c3b129291c2a · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:d1f5126ae824cbaa9ba9f8754589cbc1bff011e5109cd4816e4e2265af4eb038

Observation 7d19110e-2473-4b44-9719-1d061dc64ff5 · outbound

This paper cites Hierarchical semantic-acoustic modeling via semi-discrete residual representations for expres- sive end-to-end speech synthesis.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Hierarchical semantic-acoustic modeling via semi-discrete residual representations for expres- sive end-to-end speech synthesis

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:04:11.449582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:a6713c9b5ef3ca26106a896c261c291243eebf44d562eba3c35bb6535d9116d7

Observation c380a552-f753-4b53-9953-c07da931c3c6 · outbound

This paper cites The geneva minimalistic acoustic parameter set (gemaps) for voice research and affective computing.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations The geneva minimalistic acoustic parameter set (gemaps) for voice research and affective computing

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:04:11.456159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:96989b2440d49d3d2b2afb2f75a488a49ede9d12df57d5b29c0c031801158550

Observation 8205ffe1-9fd3-4fb5-b6de-5ea8ce83faf6 · outbound

This paper cites NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:47:12.987713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:9d53212c1786bcaac791200e076e29ef24e974f5e278f0dd752519394bdd1d11

Observation c649467e-f95c-4a0a-a0cc-2321bdb0e766 · outbound

This paper cites A scalable pipeline for enabling non-verbal speech generation and understanding.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations A scalable pipeline for enabling non-verbal speech generation and understanding

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:47:12.985035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:525b086de8f08296c4d1f7207e46b94251c45b67f1a84dce6414f2aabc9ab09a

Observation af5e7410-cce0-45a0-8c47-a190754cd14a · outbound

This paper cites SMIIP-NV: A multi-annotation non-verbal expres- sive speech corpus in mandarin for llm-based speech synthesis.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations SMIIP-NV: A multi-annotation non-verbal expres- sive speech corpus in mandarin for llm-based speech synthesis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:04:11.439499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:b77fbebdc62879eac5a4bc2113727d7544bce679d79416659d7b7b52fc9eb756

Observation 610ee6e7-2250-4543-8895-9e5295d4fd57 · outbound

This paper cites NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:47:13.005685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:fd43ab1e4130ed369acdf6237695c47192bc58817b2965f0a6fa26987b1bfbb1

Observation 1a300b4d-37a8-4b0e-b5f9-196622d66bca · outbound

This paper cites Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T07:47:13.008385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:1836d5a9659c6d9426a8fd3fe2dc3057d8e5a4166d8e127379389c6c2094d563

Observation c73c6368-2941-41f0-a830-699ceb01ff36 · outbound

This paper cites Wesr: Scaling and evaluating word-level event-speech recognition.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Wesr: Scaling and evaluating word-level event-speech recognition

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:47:13.003199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:be926ca605e9467e5a3d80d2699ded34eb5ece0f85ac23ab94f3c452168ab6bb

Observation 9300312c-b63f-415f-9138-e551ad49a83d · outbound

This paper cites ParaLBench: A Large-Scale Benchmark for Computational Paralinguistics over Acoustic Foundation Models.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations ParaLBench: A Large-Scale Benchmark for Computational Paralinguistics over Acoustic Foundation Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:47:13.000667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:ea61fb60fba025a9d66ba631dc49bca48ebec88de39d5089112e86a322ef1adf

Observation 88bbdbc7-9e49-4a19-a837-94f66f121876 · outbound

This paper cites InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:47:12.979580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:ff6f6c631758f0875347b5710bbb29b0b2536bb3579812027fbd6261e2dbdb73

Observation daef691f-79c1-488f-912b-0c260306671c · outbound

This paper cites S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:40:33.395224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:e4440edf505aa4d4284307ed8b3cef0ef4042087ff46d4f86a81e78e75edf7ba

Observation 393b341a-ea7e-4bca-a3b6-38667f2a7984 · outbound

This paper cites Paras2s: Benchmarking and aligning spoken language models for paralinguistic- aware speech-to-speech interaction.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Paras2s: Benchmarking and aligning spoken language models for paralinguistic- aware speech-to-speech interaction

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:47:12.990186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:faea2305f561819efd6272628219616b4bbf9a909c14a6e88809357d06aa6924

Observation 8bd03de5-5c01-4ae1-b85d-d1ef69b6ca18 · outbound

This paper cites Wavbench: Benchmarking reasoning, colloquialism, and paralinguistics for end-to-end spoken dialogue models.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Wavbench: Benchmarking reasoning, colloquialism, and paralinguistics for end-to-end spoken dialogue models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:47:12.977079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:34280bc484017f4458a56fd66197772c250437350db17e4fdedf0f016b76e277

Observation 1c39a316-e22e-4a4b-8b19-459207c9c1ba · outbound

This paper cites Nv-bench: Benchmark of nonverbal vocalization synthesis for expressive text-to-speech generation.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Nv-bench: Benchmark of nonverbal vocalization synthesis for expressive text-to-speech generation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:47:12.992878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:d4e5fd65dc60ef5b04ea94364a50eca25a28199abe3481322d6d4fa8c57a6847

Observation de8e9c7b-de27-4630-beed-ec1dcf749c1d · outbound

This paper cites ChatTTS: A generative speech model for daily dialogue.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations ChatTTS: A generative speech model for daily dialogue

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:04:11.437330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:1dffe98a4d0d219a4a8f88e9ded71a60216cfd843e6e9879edcf919d54bf1998

Observation b2ec56f0-efba-4e04-a19a-27d88d035169 · outbound

This paper cites Higgs Audio: Text-audio foundation model from bo- son ai.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Higgs Audio: Text-audio foundation model from bo- son ai

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:04:11.454989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:0bfa584b69ab62e4592725d7a7835d9a569af8fc41a0b0b1f8b4900f95a95834

Observation d074fde2-2cb9-4c57-8f15-3aca76c04028 · outbound

This paper cites Bark: Text-prompted generative audio model.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Bark: Text-prompted generative audio model

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:04:11.457529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:8fe9997fd1d397fcaf92639050e41dea73fe97d8ad8c733f60c4b0aa6b07a8d9

Observation e0349b6f-db83-4856-a82b-da36a5808452 · outbound

This paper cites Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:47:12.969453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:941e771e36b4e5fffa8362c571c2770236ab3125c83027a85f8e55516faa4764

Observation 5e60e40e-065c-47b2-aa30-845c6eb56194 · outbound

This paper cites Orpheus-TTS: Towards human-sounding speech.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Orpheus-TTS: Towards human-sounding speech

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:04:11.459856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:ad63805eb2e58d6f2df9a50775914bee15584d603d59653a3db5068f840a048f

Observation 3e8b8a52-24ab-41f1-9ce3-a631bd20d96a · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:19:09.677777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:73280aac3b3b64922340faf44f0595d20262418246779c60a6082b4ae9ef3781

Observation 2b61a442-c3d0-4e66-a538-8b5a990ed55b · outbound

This paper cites Elevenlabs documentation: Models.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Elevenlabs documentation: Models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:04:11.492174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:d2e8706ca31aced93912f4b8e166d0fabe48c3de5e0a668eee5892113982c102

Observation 62ac1692-a77a-4f85-994e-a04cb5541b5b · outbound

This paper cites Dia: A TTS model capable of generating ultra- realistic dialogue in one pass.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Dia: A TTS model capable of generating ultra- realistic dialogue in one pass

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:04:11.484376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:8a0491ba82f04b1bd392d78a729af292998ad0e8c48746ef6e4d34013ee153d9

Observation afc391e5-679b-4016-80c9-e2ccbf4b9025 · outbound

This paper cites SynParaSpeech: Automated synthesis of paralinguistic datasets for speech generation and understand- ing.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations SynParaSpeech: Automated synthesis of paralinguistic datasets for speech generation and understand- ing

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:04:11.489607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:abc573c199d8178e2b4822ac85bc864fde0933f4373375a6e81246ca185b84ee

Observation 8d936b37-7484-4d40-83bc-5c1754f5b965 · outbound

This paper cites of of the idea that has been the same idea for a thousand years that they believe that—.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations of of the idea that has been the same idea for a thousand years that they believe that—

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T07:47:12.964311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:72c1ff4d0ef26cf42fef2c492efd3293f84bfa7195cea0e4662c223f2dbf3f0b

Observation f8f4ee65-caee-4fec-a210-27da1115ccfb · outbound

This paper cites Acoustics of breath noises in human speech: De- scriptive and three-dimensional modeling approaches.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Acoustics of breath noises in human speech: De- scriptive and three-dimensional modeling approaches

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:04:11.491992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:9e1c177e5ef251d6f3d472cae50072dfa8c24b96d954878065f1fecab5ca3e93

Observation 38c3fd28-657b-44fc-8612-776e5fe1c62b · outbound

This paper cites V oices without words: the spectrum of nonverbal vocalisations.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations V oices without words: the spectrum of nonverbal vocalisations

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:04:11.494119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:6e74cff89a86f620eb8c33a56c467c447ce7734c9e44f404236cbc94735be03b

Observation d92b28e9-47f9-4735-8982-569207d58fa1 · outbound

This paper cites The acm multimedia 2022 computa- tional paralinguistics challenge: V ocalisations, stuttering, activ- ity, & mosquitoes.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations The acm multimedia 2022 computa- tional paralinguistics challenge: V ocalisations, stuttering, activ- ity, & mosquitoes

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:04:11.477551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:4732a1a6d78611b6faa92bdb8199fdc25ff63f32e0cadc59ea468cf69077af00

Observation c4ab0335-7a53-4f59-ac6e-2cc86afa9dd4 · outbound

This paper cites Acoustic analysis of several laughter types in conversational dialogues.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Acoustic analysis of several laughter types in conversational dialogues

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:04:11.469699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:3c49d8e0a6b8ad4ce71a341013d7d848ee854c150dafc5da8bfa6774d1874c18

Observation 32093497-db68-4a84-9949-c057dbc08bcf · outbound

This paper cites An acoustic-prosodic analysis of laughter types.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations An acoustic-prosodic analysis of laughter types

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:04:11.472508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:f8276a1d6b163307f60d2dfdfe6f51621e9f01719557ce731d466366ea3f9208

Observation ceebf8f1-919a-40ef-86c2-c857d926823a · outbound

This paper cites DNSMOS P. 835: A non-intrusive perceptual objective speech quality metric to evalu- ate noise suppressors.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations DNSMOS P. 835: A non-intrusive perceptual objective speech quality metric to evalu- ate noise suppressors

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:04:11.470009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:f543964fdd89f45348625cbe0307960eda37a59569e9c0b1bde6de2dada36b07

Observation 98f242e3-30bc-42a4-90d1-512c6f41ff9c · outbound

This paper cites Clap learning audio concepts from natural language supervision.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Clap learning audio concepts from natural language supervision

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:04:11.475025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:b2b403121e3618e1455775c4510d3e16637f9ccea502a4b6be84a76348f2e421

Observation 44d53595-dc07-436a-a513-d4467bccf832 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Robust speech recognition via large-scale weak supervision

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:04:11.479788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:9cc7c9eb86bbfe2fdae13c15256d5eef221ab9c11f523528000dae0cf4b24c7d

Observation 92dd31a9-a5c4-4108-8991-cfbc5136c7b3 · outbound

This paper cites Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to- end speech recognition.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to- end speech recognition

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:04:11.478547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:b5766b11631d3c1c6b0d40b3872910fed66ab86020825ebf7b8ebe91d974f85b

Observation daff7a43-b05c-4642-a9e5-39320103ca43 · outbound

This paper cites SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-10T07:47:12.961800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:725a6f7a121917ee063b1b500f199735c25057c191af3fc9ffe7904231524485

Observation 28fd3e27-1afb-49b1-aa3a-c8274b0c770e · outbound

This paper cites Audio-Aware Large Language Models as Judges for Speaking Styles.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Audio-Aware Large Language Models as Judges for Speaking Styles

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:47:12.959163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:3e86ca557570d5fac30484c31c9adb64ed1a4747547923115d47d9e251103635

Observation ce21878a-9bd4-4981-b48d-c6e2bf057531 · outbound

This paper cites Qwen3-TTS Technical Report.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Qwen3-TTS Technical Report

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:24:56.189411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:305fa3796e7c8e927a7c47dd726312a6b0c44e9b5478bfe3b7cb682ea808e076

Pith citing papers

Observation feabe402-ad24-4795-96dc-7bdb0f956477 · inbound

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations cites this paper.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-10T07:47:12.998128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:6a9f23b093da4e13a08e69680202a76aa1ae1fb141479bfcbf1091f128b0ab40

Observation 930b5166-b391-40eb-a030-bdd7ab00b9fd · inbound

MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech cites this paper.

MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-11T12:11:08.463100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T04:03:38.919545Z digest=sha256:c3c4e2b5a5baf93191da1f8d9d2e589474aba440539c21ff69595c41e40c076f

Observation 2c026c6e-ddcd-4cd7-bd09-4bcea3cab481 · inbound

Beyond Words: Towards Effective Modeling of Non-Verbal Vocalizations in ASR cites this paper.

Beyond Words: Towards Effective Modeling of Non-Verbal Vocalizations in ASR NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:37:29.603796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-03T00:29:39.365107Z digest=sha256:6a0e022cb0db907da938ff5d95eafe300851bcda19ffd9fcbfb123a3a67b3434