Pith. sign in

Paper Citation Record · LEDGER

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation

As of 25 July 2026, this Paper Citation Record lists 42 of 42 outbound references and 1 inbound Pith citation observation for arXiv:2605.00329.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.00329 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-09T19:17:09.247932Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-25T06:30:59.84592+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T11:52:50.598080Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact26
  • verified fuzzy9
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0663f690-9b7f-48ed-a116-c353285dea62 · outbound

This paper cites ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:46.557602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:47f3a1f3acecde3feed03d4665a6606612cac5b3d5dec2978eb990b8846d9505

Observation 88d0bdc9-221b-434c-b9b7-cbc5271fc023 · outbound

This paper cites Analytic-DPM: an Analytic Estimate of the Optimal Reverse Variance in Diffusion Probabilistic Models.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Analytic-DPM: an Analytic Estimate of the Optimal Reverse Variance in Diffusion Probabilistic Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:46.867408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:45d5088c2d2d284b240bb0c8bbbc883318df6abdca4eccc6ac5310764ca3243a

Observation 44b831ed-a124-4380-917d-6fc54cf0ea0d · outbound

This paper cites The Cramer Distance as a Solution to Biased Wasserstein Gradients.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation The Cramer Distance as a Solution to Biased Wasserstein Gradients

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:46:47.137466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:4d3789d7e29ceab0d8a4e621402e6b9bbc2918263ea65ae1acd7f4e54acdd208

Observation d115dd5f-bd38-4580-a00b-cb411e473a62 · outbound

This paper cites SoundStorm: Efficient Parallel Audio Generation.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation SoundStorm: Efficient Parallel Audio Generation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:48.083389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:6a5c1b3179075cdf04bbc2c42d2607a2733a6f7e78b86a7a00879bf3beac70c5

Observation a16c3f7f-e9ca-40e2-bb3e-6f23844d1146 · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:46.275282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:5594cc2625913b70c4fdc6c06f8784b73a42e0fa73d2db44919c161a70c8a7f4

Observation 9ad26cf3-fc44-4db2-88c2-495bf41a42b2 · outbound

This paper cites FMA: A Dataset For Music Analysis.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation FMA: A Dataset For Music Analysis

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:47.351380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:b82e8caefcac4d6727bced5d617c5bae968adc6f463710963cfc5c95bc58bb8e

Observation 2dcd9ac9-5278-48a6-87c5-5d891fd2540b · outbound

This paper cites Audio Retrieval with WavText5K and CLAP Training.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Audio Retrieval with WavText5K and CLAP Training

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:46:47.697378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:4df437ceaaa0a8148ed56178cbfc1ff03cd2546a94fc494bf6f1600d358f4239

Observation b06a06d0-96a2-4d06-b089-3466c113a9f0 · outbound

This paper cites Clotho: An audio captioning dataset.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Clotho: An audio captioning dataset

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:51:16.679073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:1dfb037f413d26664fb28f45a376c5f653492f50bd2a9b3742c0f5dbd7579076

Observation e2cea1fd-b2c2-48a6-beb6-07a389a3d921 · outbound

This paper cites D-AR: Diffusion via Autoregressive Models.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation D-AR: Diffusion via Autoregressive Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:47.878360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:7e116f768004449406a09dad189845f4e6c3b37172fd98935e68e40060f9c0f0

Observation 899a539f-3246-4b00-86f5-064ca22ed729 · outbound

This paper cites Mean Flows for One-step Generative Modeling.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Mean Flows for One-step Generative Modeling

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:46:44.154479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:ee8a02ecab76edbeacb66a8ad528e8537b15d5bee8d22387ea9dee12c20b80f3

Observation 3b1d4a25-fb5e-4e64-b0d5-26e232db8224 · outbound

This paper cites EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:45.778288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:3737c4be4cae4dbfbd1bf985e9ef2c6502fd51f934f41ec15d1dd6fd7b788530

Observation ef66f485-ee69-4016-8a72-159f0ede2af4 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Distilling the Knowledge in a Neural Network

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:46:45.044460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:83c81f759f0ce86f19b3c3984735a030e4c6e23549a289006bce8d95590874ee

Observation ebf1c355-e682-4cc3-8eff-7761ef0424ba · outbound

This paper cites Classifier-Free Diffusion Guidance.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Classifier-Free Diffusion Guidance

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:46:42.935094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:b773a3a76ced24a2e77cfcb4987215f5ad6428261fd2dc3283c3a8c5fac60a59

Observation 80b227e9-feab-41f6-9f16-c5d79b4727ba · outbound

This paper cites Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:43.397088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:f45f53ff6d9d99255e4909d0a8997706e53fe66bf8606910f7e1b186ea732b3b

Observation 2a3f17f9-1852-4239-ae25-f01a34248660 · outbound

This paper cites TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:46:44.885236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:a7c368cafc152d57e7d1503e5ea78bf1d526e6959492027d292aaf21b103cb2d

Observation 90dfe4d6-8b09-4e3b-a6a5-dc334e82d3c0 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:44.686232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:861d0fea57145ecf77d8d71f65d043a35de87957349d9fce63aa5c38d176d7a0

Observation 4e448ce7-d7b4-4467-a516-16c129be942e · outbound

This paper cites D., Kim, B., Lee, H., and Kim, G.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation D., Kim, B., Lee, H., and Kim, G

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:51:16.650211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:129878887167cd74d58c1658a9076cd1a4ab261c53a0f05d6ea52f77fee14c0b

Observation 39d88a97-f0b0-48b8-9ac7-7861846c0866 · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:44.540336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:fff8e12ad90dc29a71173e675e53b20148954e14546c67dd366a8aedf81f1945

Observation dc385cff-e34c-4667-ad7d-253fca6ea776 · outbound

This paper cites Pseudo Numerical Methods for Diffusion Models on Manifolds.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Pseudo Numerical Methods for Diffusion Models on Manifolds

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:46:45.210173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:f5454100830c8b07b3361fb215129da026b35554e79006107febb9524a5d158e

Observation 432160ed-03b1-4f91-a0c6-ce44519b2ecb · outbound

This paper cites Efficient speech language modeling via en- ergy distance in continuous latent space.arXiv preprint arXiv:2505.13181.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Efficient speech language modeling via en- ergy distance in continuous latent space.arXiv preprint arXiv:2505.13181

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:43.954400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:c4af62cfa4dcc88ca9232883b5f922e15a4398921ee8386f407fc526ef789515

Observation 96e78ea2-aa46-438e-ad0d-31167e781a3b · outbound

This paper cites Likelihood-Free Inference with Generative Neural Networks via Scoring Rule Minimization.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Likelihood-Free Inference with Generative Neural Networks via Scoring Rule Minimization

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:42.685402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:00c8511c3447fb70aef9dc283c5aae520f003650444e9e1809e887302db14e8f

Observation 359e4848-4225-4bdd-bacc-5d6ac371349e · outbound

This paper cites Diffused Responsibility: Analyzing the Energy Consumption of Generative Text-to-Audio Diffusion Models.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Diffused Responsibility: Analyzing the Energy Consumption of Generative Text-to-Audio Diffusion Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:42.234377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:dbbdf5174389e483477b94477058d9b5d6a6617b1c019ffff87ab29b0ab3244f

Observation b0d4c629-b08d-42e0-ad9a-b2a7ddd94947 · outbound

This paper cites FitNets: Hints for Thin Deep Nets.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation FitNets: Hints for Thin Deep Nets

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T01:46:30.917443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:494d687d4b4ba645d78746d9279187b03e06da9e899fbcfc68292c36ac5cca5f

Observation 13b3eb1d-f8e4-487f-8801-14cafbf05cee · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Progressive Distillation for Fast Sampling of Diffusion Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:46:43.118360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:1680e727466f1de69c41b6efddaa5e759ee344f3d18c4106ec82b377eefd4f02

Observation e6569c8a-08f6-4014-819d-d37a0819cc39 · outbound

This paper cites Consistency Models.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Consistency Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:47:29.003852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:428708960c5deeedaffc984d3ff9558a1443dda907c44dc4cf6368e78f54060e

Observation 44dd3aee-e895-49b9-a19f-10385be2c025 · outbound

This paper cites Patient Knowledge Distillation for BERT Model Compression.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Patient Knowledge Distillation for BERT Model Compression

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:46:42.823215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:345870128273bcf2d0393449ec6b352e0ee2964aaaefd9fd474d9eff2765b6a3

Observation 7bd1ae22-7544-4de1-90fc-a0f62d34b403 · outbound

This paper cites Multimodal Latent Language Modeling with Next-Token Diffusion.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Multimodal Latent Language Modeling with Next-Token Diffusion

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:44.347336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:b8f665accaf2cdb793e92452b5d63ccaacbe33273b1c96af96fc879f6c114f67

Observation 1c96f253-6e4a-4ca4-881f-7f45e0b53fcd · outbound

This paper cites Contrastive Representation Distillation.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Contrastive Representation Distillation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:45.999848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:aaf803bf10a40945aedcc5d029c3f2096cd19d32b2367c1f86e1e3c4738ae4b1

Observation 1d819b93-7627-48bd-b92b-619c12b3a0ee · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:41.933376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:0972fe7a6f2d2a1169ca9b07bd09ba93378848abff0fe301ca4ca30cc4a33532

Observation 47dce838-5b98-45e7-8a9f-9a71731b854b · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:51:16.662379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:21f74e38871ef8084ef05d9b8160fbe15e71a54c23770fdd655c0cd3ce586e39

Observation 6e270299-926a-4d1b-97f0-4b6fab95171f · outbound

This paper cites Comparing discrete and continuous space llms for speech recognition.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Comparing discrete and continuous space llms for speech recognition

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:51:16.674760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:fb8761972b3b0421d1f864cdff36dfd40cfeb88bdc9eb2b2ca381aa94b92156d

Observation ebd68a4c-3872-4398-9419-19b1b8df4f0f · outbound

This paper cites Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:42.075264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:ed1988f32cf1c5c9de5c1e229442f0ea6197d0b6c165092c994273139055470c

Observation a3e5a554-578a-4010-801a-8c6f61e5d8d6 · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.697648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:619baedd507d9ebf0474350305b4974728f8a0ae0ce34335d214bb56ba962eb9

Observation 0bee8e78-bee5-4d02-9283-0c283a87ae61 · outbound

This paper cites and Chen, Y.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation and Chen, Y

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:51:16.653895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:a9710673c3e2117f12312d2e80d6ab8c6c2a1f396fb4367d51f5ea8fd795173c

Observation 28852967-1c73-45bd-92a5-d1d156638e8e · outbound

This paper cites arXiv preprint arXiv:2505.07344 , year=.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation arXiv preprint arXiv:2505.07344 , year=

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:45.575659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:a7d530c316892f2db22404695d2514ab1452c39323a47c7729be1dfff437ba79

Observation e1c9d323-e989-4711-81ed-a990793ad6f7 · outbound

This paper cites AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:43.615417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:4acb58c8d2ccbf0ceab67b6774048ae8492cc9587ed98b8c5755db3c18b12ad1

Observation 9dd64dcb-bdc1-4abb-82d1-a8071cb81648 · outbound

This paper cites Energy-distance The following content lists out the definitions and theorems required to prove Corollary 1, stated as follows.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Energy-distance The following content lists out the definitions and theorems required to prove Corollary 1, stated as follows

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:51:16.666883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:dd2b25e3772a15b6b4e0996e6c4879eb04e753c29bf6c1f6f0adc6b4ca99b615

Observation bd54c241-ca2e-439a-bb82-080877b47316 · outbound

This paper cites The equality holds if and only ifP=Q.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation The equality holds if and only ifP=Q

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:51:16.646255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:e8e1bdceb5d0de532d00a606cdb162fb03c8a947b26cb022eb8e4fb908b320cb

Observation 22c3f3d2-1116-41bb-b205-815112fa4521 · outbound

This paper cites Text Embeddings Table 6 examines how different text embedding choices affect the performance of our one-step energy-scoring model with representation distillation.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Text Embeddings Table 6 examines how different text embedding choices affect the performance of our one-step energy-scoring model with representation distillation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:51:16.658264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:69df97dbe59a1867817ad64015ca818a68611b5781c0b0a74609526b2e76fddb

Observation ceea380d-752a-4c9e-8915-e86a52833eec · outbound

This paper cites stdev” stands for standard deviation. “stderr.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation stdev” stands for standard deviation. “stderr

Reference 40

Resolution
malformed identifier
raw_fallback, observed 2026-05-25T21:51:16.670780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:530f96e579e9a169cbb093fe54585093b1059bbadf8492259dd9bec622a0c1fc

Observation 9cdd082d-6473-41a4-8531-54fa61461d04 · outbound

This paper cites an unresolved cited work.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Unresolved cited work

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:45.473367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:be02ce740932a89d5063a1230e972afc61d5276341f6b0de25c8f882e21c265e

Observation 16bfc16b-5280-4629-9ebe-b30f486e06dd · outbound

This paper cites This 1024-dimensional vector is duplicated 78 times to form the conditioning sequence.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation This 1024-dimensional vector is duplicated 78 times to form the conditioning sequence

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:51:16.642501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-25T06:30:59.84592+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:3e0b45581712fd8c6cbcf1361534b3024bca67574c87aaef432ffc4d7bd93d49

Pith citing papers

Observation 79a2d7b3-b564-4331-a444-46dfa7549aa6 · inbound

FdAudio: MeanFlow-Anchored Fr\'echet-Distance Post-Training for One-Step Text-to-Audio Generation cites this paper.

FdAudio: MeanFlow-Anchored Fr\'echet-Distance Post-Training for One-Step Text-to-Audio Generation Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T11:52:50.598080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:52:50.598080Z digest=sha256:62826cfa359416cfb59c90b373360fec4b38ea228e5b49764e26ddf422c9b8f1