Pith. sign in

Paper Citation Record · LEDGER

MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 34 inbound Pith citation observations for arXiv:2502.18924.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.18924 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 34 of 34 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:37:04.938708Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:50:11.241591Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2e77d247-b361-440c-8a45-8d12cfda49e6 · inbound

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech cites this paper.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.112300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.112300Z digest=sha256:2f21b4c2e5c6bac599718aeee2bb37dbf851f1f990b66fed1b109147a8085716

Observation 450f5f95-ce57-49e9-964b-de015244eaee · inbound

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder cites this paper.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T22:16:17.265205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:16:17.265205Z digest=sha256:f229e1f916d2e00815e7a64b13212333ab945357327132441d6f8e6b84b3e3a2

Observation cdbe0ec9-7038-4163-a31d-7bc3d19ebd5f · inbound

TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis cites this paper.

TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:30:34.140760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:30:34.140760Z digest=sha256:79b4598fd2ea6674f3d01790c0f16026430e727b2efef588f21dfe89399e8aaa

Observation 587953c3-962d-4b96-8ae5-507bfbef7791 · inbound

ACE-Step: A Step Towards Music Generation Foundation Model cites this paper.

ACE-Step: A Step Towards Music Generation Foundation Model MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:16.366580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:16.366580Z digest=sha256:6f4f7ade9924c7b2237597d58cabbf754e3cceb56d5f83ac47756737dbd46046

Observation 099d9eb2-9a58-43c6-b9b0-a7be49a96e24 · inbound

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching cites this paper.

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:55.278735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:55.278735Z digest=sha256:4d3425e220496535ab55404ec1f87ddfc101ba3628f3708f2ea07fefd9fd3890

Observation f6bd46b6-78ae-4182-87a8-0c239981490a · inbound

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching cites this paper.

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:50:51.053768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:46:39.196042Z digest=sha256:4a37043202a96ef5bf7531a26e71d0fe312563d4342a88dbfbf978be8e121e30

Observation 03cecab6-b8f0-4506-8765-6f7727eddc1e · inbound

The Thin Line Between Comprehension and Persuasion in LLMs cites this paper.

The Thin Line Between Comprehension and Persuasion in LLMs MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:07:07.323651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T06:05:44.600781Z digest=sha256:4fcf8ef4a92957b6fefdb7b105b112be83d815e4c8d69925a5f1c19c2739316b

Observation a800ec70-a113-4f3c-8d23-8a6a1ba1b0f0 · inbound

Exploiting Leaderboards for Large-Scale Distribution of Malicious Models cites this paper.

Exploiting Leaderboards for Large-Scale Distribution of Malicious Models MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T18:16:10.849757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:16:10.849757Z digest=sha256:d3e228d46c093ca97f7c54cf3d4c09d9c62d0427422bd43f57efc953cf29f0a0

Observation 693a491d-ead1-4089-a240-50699ac46ec3 · inbound

JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment cites this paper.

JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T13:14:37.524679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:14:37.524679Z digest=sha256:d48e522c4139ee39063581672c962fd526c431b93a97e59486654750cf306355

Observation c783b992-901b-485c-ae29-8d885f5d5cd6 · inbound

Attention2Probability: Attention-Driven Terminology Probability Estimation for Robust Speech-to-Text System cites this paper.

Attention2Probability: Attention-Driven Terminology Probability Estimation for Robust Speech-to-Text System MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T16:21:56.269739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:21:56.269739Z digest=sha256:4fa6cef92a5ff0dcc849d7f742138035a2e5b59bd5de700671f64be57cca8a86

Observation 7dd14df9-6bc9-44e4-a47a-e65792cea62a · inbound

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration cites this paper.

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T19:11:59.530961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T19:11:59.530961Z digest=sha256:0d9f23985aae3c89e13bf1611eb0c597b47baac45667b2789116d7d83fcdb656

Observation e4ed716a-13e0-4bc3-a677-661395c9a330 · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:32.995304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:32.995304Z digest=sha256:26db8d9f8fb3376c64396566eca80515c896071a81561e908e11a9451b55dcef

Observation 9625e5cf-2220-4ca0-8e39-cb1d27fd56b5 · inbound

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models cites this paper.

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:03:24.776569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T23:00:18.720371Z digest=sha256:a0c81675f0d05f9ca0c22426574187dce5f4ab88ca8da9f991596c2e6b100019

Observation 1662e086-ac3a-4e49-a673-016f13136596 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.172616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:e5f45765052ab022147e6bd9e0aa0481a34e0690c77183791b6ce5570c6d43e8

Observation a4c39120-64ca-43b2-befd-ce6d88be8442 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 88

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:a8de90af97b98686ea6835317d809534ace920419e2b3e24a2e5b711dc87957a

Observation 04768d21-8e5a-4997-9c49-9a26dedad9ca · inbound

CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing cites this paper.

CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:51:00.828282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T15:46:58.010112Z digest=sha256:d9a0c2a7cce5867786a47411f02033263b08e31046dd472c45ddbd70d1c13feb

Observation 4099f967-69ac-4231-b82a-512ba525982c · inbound

From Seeing it to Experiencing it: Interactive Evaluation of Intersectional Voice Bias in Human-AI Speech Interaction cites this paper.

From Seeing it to Experiencing it: Interactive Evaluation of Intersectional Voice Bias in Human-AI Speech Interaction MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:59:50.716425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T07:59:22.416259Z digest=sha256:94e0227104a3d6d48b87f3fdf648083ee9920e135ae153cfa2f508367413ca7f

Observation 2590a872-a3f4-425c-86d7-f6e915558813 · inbound

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching cites this paper.

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-22T03:00:58.998461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T02:56:06.910758Z digest=sha256:3375a5e4055b998430c3a8e22b00f64b3ddd49b78b9e4264be5173b8c3e24c61

Observation 3ad9fbba-a588-4172-919f-eb6559d93ab0 · inbound

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching cites this paper.

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T13:29:53.318898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:29:53.318898Z digest=sha256:202c814eff387903fe1893eaf1455d7ec1f0c88996bf9836c54c6c90ebc5a8ce

Observation f700083f-6c0b-4c7d-afe0-7f7c12ce8eba · inbound

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue cites this paper.

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:13.109018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T21:05:54.061395Z digest=sha256:301de203d2811fdea211805f9831ec3c0514381cf62094288326e0999c33f70e

Observation 768e4f15-4a57-45e2-adb1-22e3139c3de8 · inbound

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling cites this paper.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.706809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:db1014563311befc6e9b77b709c26521f3fe08a773737ee1f358dad8c5914f73

Observation 111fa7d6-a449-4200-a13e-c0646cbd6737 · inbound

VoxCPM2 Technical Report cites this paper.

VoxCPM2 Technical Report MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:47:19.748586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T21:18:22.911332Z digest=sha256:35114fa380785fe0f5410c5c6d6a2a7e27bdb690868cba70c94ab2b92dcfe04a

Observation d62cfc91-7398-45ee-acda-3dc9648f1de1 · inbound

dots.tts Technical Report cites this paper.

dots.tts Technical Report MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:20.051412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:c7c4041fb30d39085c44165a2917728d58b82e40fd6dab25d3318c25e54fb5a4

Observation 6ee9bf96-8091-4db3-a043-d2fa1fc3cf9b · inbound

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs cites this paper.

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:19:02.929572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T22:46:47.786929Z digest=sha256:ce947a075ae2b6715f4d3dc8f686cf28bdcb89b467d0b7c8e76d0b9f2564e9a1

Observation 2bcc5c2b-2ddf-4239-a728-12ede820e97e · inbound

Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors cites this paper.

Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:24.876078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T19:04:25.542629Z digest=sha256:17ca72ee675085cb5934d1018b5ff95dc78246c112fd5253d65ce816bc0cd38a

Observation 9ab8c22f-96c7-43be-8eaa-9bd5e66b0d74 · inbound

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis cites this paper.

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:47:30.566017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T17:01:13.972071Z digest=sha256:be5ef59c5feec0c608b7f3a274d1f3e1ee6046c5fa54cd6124fcf9acc31e9515

Observation 910e558f-fe07-4e39-b823-61df4e107646 · inbound

Joint Residual Reweighting for Classifier Free Guidance in Flow-Matching Zero-Shot TTS cites this paper.

Joint Residual Reweighting for Classifier Free Guidance in Flow-Matching Zero-Shot TTS MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:50:11.243408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-25T19:35:52.631543Z digest=sha256:159f01e72a73f6ec89295d54dd7dc936e81940182954fd90fb40face0425c357

Observation 763bbe90-8755-42a1-9e6e-4e16b0463491 · inbound

Joint Residual Reweighting for Classifier Free Guidance in Flow-Matching Zero-Shot TTS cites this paper.

Joint Residual Reweighting for Classifier Free Guidance in Flow-Matching Zero-Shot TTS MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:44:40.076521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T10:07:48.053378Z digest=sha256:5dfbbda1d0f8674fd1729948f772501adaf3a125bfe65d55f481d0c710cb820b

Observation 4d72292d-bc6e-491b-a2d7-affa35327ce3 · inbound

DETECT-3B-Omni is Agnostic of Content and Demographics cites this paper.

DETECT-3B-Omni is Agnostic of Content and Demographics MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T02:38:17.163398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T02:38:17.163398Z digest=sha256:97e97dd71f445242bd22b4ee1c4895bd48063ae15efbe0ac3a098c00cf1a84b4

Observation 966fa3d7-be6a-434e-9f6b-365d6f941014 · inbound

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models cites this paper.

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 82

Resolution
unresolved
no resolver link, observed 2026-07-13T05:10:26.667731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T05:10:26.667731Z digest=sha256:e67e754e5eabd2f0c1502994a06e675049ce5c0862b3eea96c642dbe934e332e

Observation 1562ad06-8805-48b2-a8ad-4d7b16698ae6 · inbound

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis cites this paper.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:36.903349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:36.903349Z digest=sha256:01884f322b63ac97f9d4bd36d4d4412d942f9bba9016bb8e15d20e24c47b1714

Observation 5f8c3cba-3765-47a0-b2de-46647600cdaf · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:23.503407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:23.503407Z digest=sha256:8019790823bc3bbd5b7bf16305a4c488fc10722543e94882eb333365f897408c

Observation bf3d068e-ceb0-4675-abbf-063188f005fa · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:44.716159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:44.716159Z digest=sha256:beeb64b4f785dbfd5d6c6ac09c257089c1c14daae534b8e473d4cfd5c55b538d

Observation 38fea359-7336-4eca-8aec-8ec2fa84dc1d · inbound

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization cites this paper.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.938708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.938708Z digest=sha256:d88b855a9fd168fded0ba8af240fff031f7ebe919a6c5fc7d298b63efad04ca8