Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-09T19:17:09.247932Z
Paper Citation Record · LEDGER
As of 25 July 2026, this Paper Citation Record lists 42 of 42 outbound references and 1 inbound Pith citation observation for arXiv:2605.00329.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-09T19:17:09.247932Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-07-25T06:30:59.84592+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-14T11:52:50.598080Z
A source-named dated measurement, never combined with another source.
Source: cited_works
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0663f690-9b7f-48ed-a116-c353285dea62 · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 88d0bdc9-221b-434c-b9b7-cbc5271fc023 · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Analytic-DPM: an Analytic Estimate of the Optimal Reverse Variance in Diffusion Probabilistic Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 44b831ed-a124-4380-917d-6fc54cf0ea0d · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation The Cramer Distance as a Solution to Biased Wasserstein Gradients
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation d115dd5f-bd38-4580-a00b-cb411e473a62 · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation SoundStorm: Efficient Parallel Audio Generation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation a16c3f7f-e9ca-40e2-bb3e-6f23844d1146 · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 9ad26cf3-fc44-4db2-88c2-495bf41a42b2 · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation FMA: A Dataset For Music Analysis
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 2dcd9ac9-5278-48a6-87c5-5d891fd2540b · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Audio Retrieval with WavText5K and CLAP Training
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation b06a06d0-96a2-4d06-b089-3466c113a9f0 · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Clotho: An audio captioning dataset
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation e2cea1fd-b2c2-48a6-beb6-07a389a3d921 · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation D-AR: Diffusion via Autoregressive Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 899a539f-3246-4b00-86f5-064ca22ed729 · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Mean Flows for One-step Generative Modeling
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 3b1d4a25-fb5e-4e64-b0d5-26e232db8224 · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation ef66f485-ee69-4016-8a72-159f0ede2af4 · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Distilling the Knowledge in a Neural Network
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation ebf1c355-e682-4cc3-8eff-7761ef0424ba · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Classifier-Free Diffusion Guidance
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 80b227e9-feab-41f6-9f16-c5d79b4727ba · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 2a3f17f9-1852-4239-ae25-f01a34248660 · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 90dfe4d6-8b09-4e3b-a6a5-dc334e82d3c0 · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 4e448ce7-d7b4-4467-a516-16c129be942e · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation D., Kim, B., Lee, H., and Kim, G
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 39d88a97-f0b0-48b8-9ac7-7861846c0866 · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation dc385cff-e34c-4667-ad7d-253fca6ea776 · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Pseudo Numerical Methods for Diffusion Models on Manifolds
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 432160ed-03b1-4f91-a0c6-ce44519b2ecb · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Efficient speech language modeling via en- ergy distance in continuous latent space.arXiv preprint arXiv:2505.13181
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 96e78ea2-aa46-438e-ad0d-31167e781a3b · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Likelihood-Free Inference with Generative Neural Networks via Scoring Rule Minimization
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 359e4848-4225-4bdd-bacc-5d6ac371349e · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Diffused Responsibility: Analyzing the Energy Consumption of Generative Text-to-Audio Diffusion Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation b0d4c629-b08d-42e0-ad9a-b2a7ddd94947 · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation FitNets: Hints for Thin Deep Nets
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 13b3eb1d-f8e4-487f-8801-14cafbf05cee · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Progressive Distillation for Fast Sampling of Diffusion Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation e6569c8a-08f6-4014-819d-d37a0819cc39 · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Consistency Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 44dd3aee-e895-49b9-a19f-10385be2c025 · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Patient Knowledge Distillation for BERT Model Compression
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 7bd1ae22-7544-4de1-90fc-a0f62d34b403 · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Multimodal Latent Language Modeling with Next-Token Diffusion
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 1c96f253-6e4a-4ca4-881f-7f45e0b53fcd · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Contrastive Representation Distillation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 1d819b93-7627-48bd-b92b-619c12b3a0ee · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 47dce838-5b98-45e7-8a9f-9a71731b854b · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 6e270299-926a-4d1b-97f0-4b6fab95171f · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Comparing discrete and continuous space llms for speech recognition
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation ebd68a4c-3872-4398-9419-19b1b8df4f0f · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation a3e5a554-578a-4010-801a-8c6f61e5d8d6 · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 0bee8e78-bee5-4d02-9283-0c283a87ae61 · outbound
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 28852967-1c73-45bd-92a5-d1d156638e8e · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation arXiv preprint arXiv:2505.07344 , year=
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation e1c9d323-e989-4711-81ed-a990793ad6f7 · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 9dd64dcb-bdc1-4abb-82d1-a8071cb81648 · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Energy-distance The following content lists out the definitions and theorems required to prove Corollary 1, stated as follows
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation bd54c241-ca2e-439a-bb82-080877b47316 · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation The equality holds if and only ifP=Q
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 22c3f3d2-1116-41bb-b205-815112fa4521 · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Text Embeddings Table 6 examines how different text embedding choices affect the performance of our one-step energy-scoring model with representation distillation
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation ceea380d-752a-4c9e-8915-e86a52833eec · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation stdev” stands for standard deviation. “stderr
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 9cdd082d-6473-41a4-8531-54fa61461d04 · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 16bfc16b-5280-4629-9ebe-b30f486e06dd · outbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation This 1024-dimensional vector is duplicated 78 times to form the conditioning sequence
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.
Observation 79a2d7b3-b564-4331-a444-46dfa7549aa6 · inbound
FdAudio: MeanFlow-Anchored Fr\'echet-Distance Post-Training for One-Step Text-to-Audio Generation Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.