Pith. sign in

Paper Citation Record · LEDGER

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

As of 21 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 8 inbound Pith citation observations for arXiv:2508.04195.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.04195 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:51:32.702267Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:40:38.072150Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T14:25:02.434510Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact2
  • verified fuzzy6
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a7329c7c-d665-42d1-83e9-0297e80d850d · outbound

This paper cites , " * write output.state after.block = add.period write newline.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:29.419290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:29.419290Z digest=sha256:6a8fcd89b03c34191a675f8f2615bad82f58b62024a1b8f1f02531e0a60b795b

Observation 72272a94-9074-45d3-9444-56577ac6960d · outbound

This paper cites write newline.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:29.470752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:29.470752Z digest=sha256:bd631e038308bca97578be150dd262968a84334636f4eea45dcf212467038248

Observation 001b74c3-9d4b-4ac2-b2fd-de1a1b7f24b0 · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:29.524524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:29.524524Z digest=sha256:500f33ba55c0e585bfe968488c0ba8944dbfe4239ec24e0520636773f1114c7b

Observation bc2d1ba8-8745-4564-bf8f-7f57b24c0aa9 · outbound

This paper cites Humane Speech Synthesis through Zero-Shot Emotion and Disfluency Generation.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Humane Speech Synthesis through Zero-Shot Emotion and Disfluency Generation

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T00:51:33.246231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T00:51:29.579149Z digest=sha256:ee8f83de7ea5a6f4d883211d99233e8344e5862f3a97669fa42cdf2e1996083a

Observation f7251430-0160-4158-b836-a7496021ba59 · outbound

This paper cites BEATs: Audio Pre-Training with Acoustic Tokenizers.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:29.678062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:29.678062Z digest=sha256:f64de7f0d061ef1cfa0ede81dd64aa2b5b54bdac5ebc3bb1d3ef5895059dc80a

Observation c2d1e50e-b985-4320-bbcb-8f6a59ffe200 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:29.737750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:29.737750Z digest=sha256:e1b95cc923c8e71a48f0aab8d5be1bcfd5fb6c463c4d5ee79d1a920312d36b1e

Observation 9ca561e2-bdbc-49fa-8171-9c771a950e8b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:29.827602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:29.827602Z digest=sha256:32378b20e3a03953b6513265a98ee6a17e7383d3f8c274163f098949f2df0cf9

Observation c3fbd9e7-19a0-47af-93b3-4c819a8ee3e5 · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:36.556797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T00:51:29.891439Z digest=sha256:2b025d59f3516c6a29eb5a5e23026c574819e6c543b244c9c45af853e85b26d7

Observation ea8f7aa4-6cb6-4a43-a588-7b5dfc2e2fbf · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:29.957770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:29.957770Z digest=sha256:32196ed9a603bc7f46e5da6e5e9cf7be628f3d0fa22356e717c9e42e8cb765ff

Observation b679bf31-2d50-4682-94bf-78ffa0bfacc2 · outbound

This paper cites FunASR: A Fundamental End-to-End Speech Recognition Toolkit.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:30.009070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:30.009070Z digest=sha256:12cc85754ac49812c5b001104baa73dbde91fa1dc17167869b0318446732bcc8

Observation 15c7e2f3-41c4-4e96-ba2e-5ba58fe90610 · outbound

This paper cites Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:30.106597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:30.106597Z digest=sha256:5827c134ca487f24305873107c5ed8588555d49a7dac6e40f3347a41535ea060

Observation 89df5f50-df60-457b-bc39-d774e3dc31b9 · outbound

This paper cites F.; Ellis, D.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations F.; Ellis, D

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:51:36.375230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T00:51:30.179764Z digest=sha256:b57fc1aa91227df4dfa117ee87ed2ebfd301083caec066c1537fc5b1e47f118d

Observation 331361af-2510-447b-bd02-e88df15243df · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:36.216252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T00:51:30.195673Z digest=sha256:a75409df09bb122875ae616da4fa819af6516e605c7320308753c827784d15d6

Observation 0deb0d0e-f8ba-4bf9-bad0-278bc699f7b0 · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:36.042699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T00:51:30.296507Z digest=sha256:723fdfd8a79c3a4bb4eeb57b300682ecc8d49a5c78fa5aa2e01fbaddd729a2c5

Observation c671d657-7cf7-46fc-a3ff-b7c02cdf8e9e · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:30.398007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:30.398007Z digest=sha256:3b5fc2c5061c537ed2580cebec22871862864eda6103440ba8c1d23dec108db0

Observation 4b1c2fe0-0dfc-4965-84d1-e33b4bc65832 · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:35.882856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T00:51:30.527531Z digest=sha256:a13cba32ba45d01e433dd4cb10afb52a3ee6463a0ac16a01340fa62158658e02

Observation edc42f4b-f738-489b-bf02-57c465a9e932 · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:30.658204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:30.658204Z digest=sha256:7c73459e9afed8f0b4617fdc937f176751ab366869accfa058b8d90042ea2ddc

Observation ee47a100-26eb-4ef5-ba87-838361fe724a · outbound

This paper cites Making Flow-Matching-Based Zero-Shot Text-to-Speech Laugh as You Like.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Making Flow-Matching-Based Zero-Shot Text-to-Speech Laugh as You Like

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:30.769766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:30.769766Z digest=sha256:c1513dfdad082981a581df9f9db392a7dbd4da16193837ff170eb9b39a889362

Observation 3b946c59-8ff3-4d57-96c4-14438a6cac74 · outbound

This paper cites S.; and Aslin, R.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations S.; and Aslin, R

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:51:35.727651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T00:51:30.924211Z digest=sha256:2502c1ccf05bc575d028d59444ba04bf4b9150a304c04dfcf43389864547cda9

Observation 6f58bd19-c46a-484a-ba33-d99fb23b5c44 · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:35.582626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T00:51:31.118101Z digest=sha256:d3582ae069a80910a9df4022212888eb490f1845598c3663a4e9e77140bff981

Observation c48991fa-e5de-4190-a353-fa8ff95e2e32 · outbound

This paper cites M.; and Nass, C.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations M.; and Nass, C

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:51:35.440905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T00:51:31.272438Z digest=sha256:256f31f136c8a9fb6e2250ba9cb3a8cec8814dda007ab52cd7da944e67e65b57

Observation 7d22d4cb-c8e1-4550-a9c7-783d4027daf7 · outbound

This paper cites E.; Rohde, H.; and Corley, M.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations E.; Rohde, H.; and Corley, M

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:51:35.276295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T00:51:31.335779Z digest=sha256:e65f1980e71f7f9fe2fe68524a73435a87fb253d01b5c5e9da4e3d4ba2f4f276

Observation 4402f487-40f7-461f-be2d-02e40eaa173d · outbound

This paper cites EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:31.393989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:31.393989Z digest=sha256:13462b95b2165f8bda84a63f4b3bf6e055b9c702650e1d9c67c34b91e9b8c689

Observation 90a046bd-2328-42d7-be3b-53404aae77a4 · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:35.131919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T00:51:31.467268Z digest=sha256:885ab5ae73169d954df87dddfcd64627b16584433988345e7c41f25dc440e43a

Observation 4d568a58-48f6-438a-bb20-7a83a32d2c25 · outbound

This paper cites W.; Xu, T.; Brockman, G.; McLeavey, C.; and Sutskever, I.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations W.; Xu, T.; Brockman, G.; McLeavey, C.; and Sutskever, I

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:31.540103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:31.540103Z digest=sha256:035ac9e7de453af457580928bc89007841815d7d589cbbdb03c5ba947cbd991b

Observation e61ee274-39f3-4365-b648-63172564e471 · outbound

This paper cites M.; Li, G.; and Du, C.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations M.; Li, G.; and Du, C

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:51:34.968392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T00:51:31.638027Z digest=sha256:4ffd1e233b22590d72e212fcf88a2df718179d787cb74018317fa117efdf45fe

Observation f1c1672f-6204-4d98-aea6-49563bb65895 · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:34.794508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T00:51:31.749455Z digest=sha256:9d47e944cb69a674c4ba9ce174030f6990efe70de4c5637e2eb31c63965e2e29

Observation 2ac5f453-ff97-419e-b111-48e6799535e2 · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:31.866776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:31.866776Z digest=sha256:9226b69d25500cacb9623440a7f6277461f9d99923ed8424d2604c319e97cb56

Observation 350cd4e5-6116-4810-85fa-692f7834aa6d · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:34.623200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T00:51:31.939646Z digest=sha256:98e3be47c00875ee4f1463c68ec3c6b7e9b1dcda57d2cce2107fdd109f3a8c40

Observation 85795783-f8c7-4e9e-9ce3-5e4595ae444a · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:34.396959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T00:51:32.025999Z digest=sha256:4f3709fc0aa4ea47652b1bdc8abac3c9a5fe03b0d5df7690ff35d2e9eb6b8dae

Observation c679a6ea-d81e-483d-a059-f530be6fd47f · outbound

This paper cites A Survey on Neural Speech Synthesis.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations A Survey on Neural Speech Synthesis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:32.131382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:32.131382Z digest=sha256:3cf813e8d011f0edd244da8c1e9f090fc5068a9a89fb5d3ca664c0df4f227e04

Observation c2447670-cbd6-4f8c-9fac-3720882b3d2a · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:34.094456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T00:51:32.233349Z digest=sha256:457823441fd89db7acc138e8cb1dfc5201152f8053152f51d2d9d704a26d1883

Observation 24c5b1c2-1b12-4404-9cba-5c97d514214d · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:33.881745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T00:51:32.333796Z digest=sha256:08407a040720b7cedceb4c999c49b576d30e6d3bc9b6e67a88236e6f404bfe3e

Observation 701de64b-c30b-4506-bb16-5bd8000d593d · outbound

This paper cites DisfluencySpeech -- Single-Speaker Conversational Speech Dataset with Paralanguage.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations DisfluencySpeech -- Single-Speaker Conversational Speech Dataset with Paralanguage

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-06T00:51:32.888098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T00:51:32.388095Z digest=sha256:aa0c21467b24b10cecabcc7e76bb494a92901d07806b49130bddcd0014a95f01

Observation fe9c6c71-7323-4996-bf15-f3cd272cfa16 · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:33.665539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T00:51:32.480445Z digest=sha256:86a4ed50d12470f588edb0c90dd444857680c6b3f5cd79bf117bee6cc6f093aa

Observation e1eb9f8d-3177-42f9-a402-81b86bc9775d · outbound

This paper cites E.; Thakker, M.; Tompkins, D.; Tsai, C.-H.; Li, C.; Xiao, Z.; Zhao, S.; Li, J.; et al.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations E.; Thakker, M.; Tompkins, D.; Tsai, C.-H.; Li, C.; Xiao, Z.; Zhao, S.; Li, J.; et al

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:51:33.505371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T00:51:32.557490Z digest=sha256:bd174d209ee667a5a2b9bc4db2a05edafedfd6fe0a83c844b0745d39d7630251

Observation 352d1629-f1fb-4857-a0c5-0c74c2bd5a9d · outbound

This paper cites Open Source MagicData-RAMC: A Rich Annotated Mandarin Conversational(RAMC) Speech Dataset.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Open Source MagicData-RAMC: A Rich Annotated Mandarin Conversational(RAMC) Speech Dataset

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:32.620801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:32.620801Z digest=sha256:f92ffbaa0b3180cc0deb3172d710f401425c03ac71125691e301f0582d753bd9

Observation b3f16aa5-609e-407f-8418-baa6bc143a82 · outbound

This paper cites an unresolved cited work.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:51:33.368415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T00:51:32.702267Z digest=sha256:644786dcec23ecf53ba6360ba8cffd3356ee900a5a3d71d307cc55553bb196e9

Pith citing papers

Observation 4a89afb3-bc58-402c-a22f-e4a9c46b3b84 · inbound

Flavors of Moonshine: Tiny Specialized ASR Models for Edge Devices cites this paper.

Flavors of Moonshine: Tiny Specialized ASR Models for Edge Devices NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T16:40:38.072150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:40:38.072150Z digest=sha256:fe2bd736a38bab58224d876e1d2db411162ad165f848e7b43048cd016c2a4b04

Observation dc465a96-f3af-4c05-bdc3-c67a12c398db · inbound

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models cites this paper.

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:03:24.826968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T23:00:18.720371Z digest=sha256:5b1c8ef3317671649714021e74425999472cc5a2ad151c06f67f786d0edaae8c

Observation 610ee6e7-2250-4543-8895-9e5295d4fd57 · inbound

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations cites this paper.

NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:47:13.005685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T07:37:44.393592Z digest=sha256:97465919c7940eb5e153445c3acdc0f5b12a557a77a8a80d36932f8e807240b6

Observation ecc84b56-edce-420c-b44d-25ffcba4899c · inbound

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis cites this paper.

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:09.976948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T12:03:11.243693Z digest=sha256:30ba7d1ea188dde553c0b961f658ae03d7291292899cd208a0a05ef203c72e42

Observation 16435fef-12da-4d26-819a-82ef45fa8ac9 · inbound

Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Model cites this paper.

Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Model NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:07:13.507801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T04:03:27.608638Z digest=sha256:62e0c577f12836a34a6f6844b78d096cd19aae939ee443f308d2b3aa83bacd9e

Observation 8eb7ae9c-ae45-4db1-b2f8-41d6733a6577 · inbound

Beyond Words: Towards Effective Modeling of Non-Verbal Vocalizations in ASR cites this paper.

Beyond Words: Towards Effective Modeling of Non-Verbal Vocalizations in ASR NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:37:29.604600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T00:29:39.365107Z digest=sha256:7a73af0fedef572f72a5cab1e1705c62873dd90b241c868f1122f33884f3d4ed

Observation 780fd3bc-2880-4093-a9f3-601039e19f43 · inbound

TriA Pipeline: A Large-Scale Automatic Audio Annotation Pipeline For Audio Classification In Specific Scenarios cites this paper.

TriA Pipeline: A Large-Scale Automatic Audio Annotation Pipeline For Audio Classification In Specific Scenarios NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-08T14:25:02.436023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T14:20:24.828073Z digest=sha256:303ebd81764f993a599c6ecc1aff597da608da1a59c28799f82c1ba0761792da

Observation d6960b66-d470-4ba4-9ba0-c4a9026b2694 · inbound

Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing cites this paper.

Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T14:01:52.696984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:01:52.696984Z digest=sha256:3ebc7c4c5b49bde088d1f3a2b643f728534dd8aa074fdc7cf6809164f673bd6b