Pith. sign in

Paper Citation Record · LEDGER

Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2407.05361.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.05361 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:58:07.862140Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T03:37:35.221767Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0847645e-1357-4457-be39-54962b7a6d8c · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.461429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:385ec5bc5d4820ead2d3485b6bb082551a1befc0535667822a1823a496d16f9b

Observation 619ad11f-4d92-493b-ac30-edf1ab09139c · inbound

Ming-Omni: A Unified Multimodal Model for Perception and Generation cites this paper.

Ming-Omni: A Unified Multimodal Model for Perception and Generation Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:07.862140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:07.862140Z digest=sha256:868241c7e901ad42e67dde9cd9cbb88c9572a277c274c21090516ae2cb48e7a7

Observation 9984cf1e-8e2c-41fd-8878-4c2e019d5754 · inbound

StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding cites this paper.

StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:33:09.346129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:33:09.346129Z digest=sha256:5c3fdda851acbe9f79ef256ceb1c092655827caf2648291fc689d2318743bcf9

Observation 89261381-9e78-4f79-bcef-a759ffdab20e · inbound

NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech cites this paper.

NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:33:52.444640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:33:52.444640Z digest=sha256:ddd3e26b245e47221e7b487975282999294929e6c0e0b94fe2933bbaeb0b2554

Observation d4a39cdd-a493-4d96-90f3-f2082407f8dd · inbound

Whilter: A Whisper-based Data Filter for "In-the-Wild" Speech Corpora Using Utterance-level Multi-Task Classification cites this paper.

Whilter: A Whisper-based Data Filter for "In-the-Wild" Speech Corpora Using Utterance-level Multi-Task Classification Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T12:36:16.250981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:36:16.250981Z digest=sha256:39ea3b7834e7302fbcd82bfc4ed61524004ae8a70072cc2e1562ed53aeebe3ab

Observation 2dd37568-2bd1-4ee7-bb30-6599b6d618ad · inbound

REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers cites this paper.

REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T23:41:51.083471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:41:51.083471Z digest=sha256:114b2e82fce7ea35bbf21190a0c207632b04ffaa878ffd01b96599ade93612ab

Observation 8d757418-74ce-4c4a-bd83-d8a955f0829b · inbound

AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation cites this paper.

AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T11:41:00.064681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:41:00.064681Z digest=sha256:86c305b977f8840a4647cd5ee1e4ddc297a5b4f8eb6303223cf642fdb90cf825

Observation 8623222a-d6e0-4eec-9ba4-f132b042c462 · inbound

Over-the-Air Adversarial Attack Detection: from Datasets to Defenses cites this paper.

Over-the-Air Adversarial Attack Detection: from Datasets to Defenses Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T19:27:38.516808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:27:38.516808Z digest=sha256:254cd41ffd32efdb4ead55adcdeb05b9a56f26d6163da05d2ad5e7043d6b22b3

Observation 67778385-97e8-4920-b666-a32c0c30aee5 · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:32.656117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:32.656117Z digest=sha256:1a35c8097aaa0291f298a73d66500cc587f9f2a9951bc206cb1f9677be17f992

Observation 74e8b097-9b56-4718-84c0-761387fba4d7 · inbound

Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training cites this paper.

Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T18:03:42.435007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:03:42.435007Z digest=sha256:895d561aa61f1bc03b391e604f02218ef565a0ae1982cf6406838ce9a0812ceb

Observation 608f90e4-bc54-4082-943e-6bf1fa17f6bf · inbound

StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration cites this paper.

StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:24:00.893750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T22:15:17.804936Z digest=sha256:74a6d102be62ff6737cea5dcce26afcc81bfd034fad11af95eee78ac850a364f

Observation e34932ed-ecdb-42da-a5f3-21cdaa1d3d9f · inbound

Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS cites this paper.

Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:16:11.062977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T21:31:09.399885Z digest=sha256:9ce54c2cf8292b49a35186458d405818a74c3ec67566104ffc9089df4978153f

Observation d62537a3-9ff4-49c5-9a60-64ed195b9888 · inbound

BareWave: Waveform-Native Flow-Matching Text-to-Speech cites this paper.

BareWave: Waveform-Native Flow-Matching Text-to-Speech Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:37:35.223279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T15:14:03.681833Z digest=sha256:81574125a73d13689b1608d39abd14e78074303cb80affce336e7195d716a3fb

Observation 7de01159-c38e-444b-aa96-f909a7d722ed · inbound

Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders cites this paper.

Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:27:30.042911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T17:09:36.845217Z digest=sha256:f199845582c1463c6ffa23df8f5c33dd017de991a4fabbc1087b01757da7f03b

Observation 1524d8be-3023-4edb-aeea-5ff6d095c286 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 170

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:45:47.191378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:e985249e1eccc5d9bc52f532846d875f55d8f52aceccb71e5eb7fd6d1a4b220e

Observation dcfe1bb0-7916-43a6-ad3b-06d4a79bf3d3 · inbound

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis cites this paper.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:14.162197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:14.162197Z digest=sha256:6bb5d71ee46fa5e676e70acf76142fb98f57516780c0eb346bda959e0544592b

Observation 5cca810b-9b01-4a15-9bcd-fb0ddc527fdc · inbound

Investigating Codec-Internal Latent Audio Watermarking for Neural Codec Robustness cites this paper.

Investigating Codec-Internal Latent Audio Watermarking for Neural Codec Robustness Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T08:25:28.128920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:25:28.128920Z digest=sha256:8d1aa0f20a6fb5a0658ab7e21ca30a7c31860744d06ce6e09cc0e91632275848

Observation 552f8c7f-7728-4854-8caa-0392f27ae99d · inbound

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis cites this paper.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:37.291105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:37.291105Z digest=sha256:7a8dfb66896e5a5e363cfa831b837bbacbc09eb8cf591e94f8ea468baaa7e90f