Pith. sign in

Paper Citation Record · LEDGER

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation

As of 23 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2606.31259.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.31259 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-01T03:45:40.636290Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy51
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5ea8592d-e260-4e9d-846d-c89534dc9c97 · outbound

This paper cites Make-an-audio: Text-to-audio generation with prompt- enhanced diffusion models,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Make-an-audio: Text-to-audio generation with prompt- enhanced diffusion models,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.254303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:779ff432c3131ccd9233485e6d6c95ec5b51c005553f466a1fbf9d18a014486f

Observation e832f22f-5a78-44ff-b0b8-063574399ea2 · outbound

This paper cites Audiogen: Textually guided audio generation,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Audiogen: Textually guided audio generation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.256235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:b12d0d87db1bebd21d259569ad074169a0f8d1beb07f3535b1abefe3cca4662e

Observation 4a620aa2-6883-4dea-aee5-5764793845de · outbound

This paper cites AudioLDM: Text-to-audio generation with latent diffusion models,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation AudioLDM: Text-to-audio generation with latent diffusion models,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.249336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:74e0f8aad45a60a875b79df3406285bc0922de568b4f5d2d40ed1e5dca3d65c8

Observation 5b884f80-2ba9-48b0-abf5-378989279e80 · outbound

This paper cites Audioldm 2: Learning holistic audio generation with self-supervised pretraining.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Audioldm 2: Learning holistic audio generation with self-supervised pretraining

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.245047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:88c0911488183b0c2d74775a9dade66e0e0ff0f7244ff403b199fcbbc8979c24

Observation 414dd74d-f4dd-452a-8876-22be12576b5d · outbound

This paper cites Any-to-any generation via composable diffusion,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Any-to-any generation via composable diffusion,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.263868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:2f79a94ecee80a9e64ae2de0ccafec56006d38589357bcb76017f7b2bfe5fbd2

Observation 49ad1c5a-e797-4ad2-92b8-5d4e199f14f4 · outbound

This paper cites Auffusion: Leveraging the power of diffusion and large language models for text-to-audio generation,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Auffusion: Leveraging the power of diffusion and large language models for text-to-audio generation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.258221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:d0c982393b01e06519fa697dd30539b2c2d96b80263df5cd883fade0c1986f07

Observation 44652a9a-ebf0-4e02-a6cd-f58d8276cf61 · outbound

This paper cites Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.332480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:8ac2bb2bbcde8149f067feff9ffab53792e01109a9a64c81480bdb6d5d384a2c

Observation 28a7277b-f8bf-4cfc-8ef4-3f017db3587e · outbound

This paper cites Denoising diffusion probabilistic models.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Denoising diffusion probabilistic models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.279657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:2b80e845e480d72aeb5807386368632a0fa157516672aab4bfa9a40859242b09

Observation 4d8ee700-e503-4295-9a56-83ddbf834e01 · outbound

This paper cites Denoising diffusion implicit models.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Denoising diffusion implicit models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.283418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:6f05074ac729d10e125cd702372eb300f6268f4280a13e9d5573b2993df90836

Observation ea950e4a-e0df-44d3-b54e-605c408b205c · outbound

This paper cites Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.330382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:d206d0f58e47ee1c6d448f071794f4f411a472424845b5b45a20bacb4bb2a57b

Observation e81d83d3-b117-46d2-8dc2-ed5211fa980c · outbound

This paper cites Dpm-solver++: Fast solver for guided sampling of diffusion probabilistic models,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Dpm-solver++: Fast solver for guided sampling of diffusion probabilistic models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.334663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:8c636f10cd270a50b989a475553c749028c6e65eca9e7aba334b07c91a222651

Observation 1c8c73f1-8741-4dfb-8095-ac27e46cd834 · outbound

This paper cites Elucidating the design space of diffusion-based generative models,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Elucidating the design space of diffusion-based generative models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.338922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:bc2722db73245af320c32fe46962780361f075157036ea4552c9b6221e327ff3

Observation b85ed6f7-fc2e-45d7-b3a3-a855b02a826e · outbound

This paper cites Analytic-dpm: an analytic estimate of the optimal reverse variance in diffusion probabilistic models,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Analytic-dpm: an analytic estimate of the optimal reverse variance in diffusion probabilistic models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.323764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:cadaf90c660b8293c1ba1e43cef7a273c0ef2b8e15f8a327cd59ed00ed18ccb9

Observation 8a6807da-9f65-4cc3-b049-c459aae6b750 · outbound

This paper cites Consistency models,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Consistency models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.353838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:0d2499155484db877d99317e214cb3fd2865cf1f46c0c6424c8f8b374f9450b9

Observation 100b490a-db43-47b8-a455-8d17d3293fc5 · outbound

This paper cites Consistencytta: Accelerating diffusion-based text-to-audio generation with consistency distillation,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Consistencytta: Accelerating diffusion-based text-to-audio generation with consistency distillation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.325851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:1658eb1170a238c1e6f948b2398633b8feaff94cb0086b09238e21e633a405d1

Observation 7e86ca38-588f-42c1-b013-dbf301d508dd · outbound

This paper cites Audiolcm: Efficient and high-quality text-to-audio generation with minimal inference steps,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Audiolcm: Efficient and high-quality text-to-audio generation with minimal inference steps,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.348917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:132da1504485cd29ce935f8857e9b665a80187cb884c11a4005c63436c22bf81

Observation 29755e60-187a-4df3-8d9b-4a82e9429a5b · outbound

This paper cites Instructpix2pix: Learning to follow image editing instructions.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Instructpix2pix: Learning to follow image editing instructions

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.310263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:f9f2af75564d1a62c3a7f06430176b0796f9e5e750fa5646d830bc5e7e254000

Observation 0ce01166-648a-45b2-9f0d-a3d5aa2231c5 · outbound

This paper cites Ts2f: Text-assisted speech-to-face generation,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Ts2f: Text-assisted speech-to-face generation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.321940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:971e49c3b89355c82ceb06ec8bb7aee575132f9a8d6bfc8034ff60f36a766d05

Observation 147d604e-9d83-49ac-b59a-af2cdb335604 · outbound

This paper cites AudioCaps: Generating captions for audios in the wild,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation AudioCaps: Generating captions for audios in the wild,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.319977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:24bc9cc20c0135659c109fc1afe7149cd09139d7abd97f034d816ef470b70da1

Observation 4c4083aa-ab37-4cd7-9df1-2c2dc7bbe03d · outbound

This paper cites WavCaps: A ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation WavCaps: A ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.315584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:dc8c1c5f72037d3e509de620ec4022df561128b43bf328467fa4b511273549e8

Observation d91e1fe3-7281-4a62-bf7f-fdb63386e5b0 · outbound

This paper cites Audiosetcaps: Enriched audio captioning dataset generation using large audio language models,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Audiosetcaps: Enriched audio captioning dataset generation using large audio language models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.317779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:62405df913fd70deac5ad0978c3963920b0dc45aa03ba99b7b2bd790c44316be

Observation adc3b1f5-3410-4d24-9607-0b7b30dd42f7 · outbound

This paper cites Journeydb: A benchmark for generative image understanding,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Journeydb: A benchmark for generative image understanding,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.328463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:a655158ff9e0915a73b30df5b641bcbb7b991716c78dca5b42bb3ca62cfcbd5c

Observation 94db4798-f07f-480b-98ea-bb6f2d572426 · outbound

This paper cites Laion- 5b: An open large-scale dataset for training next generation image-text models.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Laion- 5b: An open large-scale dataset for training next generation image-text models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.336702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:2c8288506660d885510b570c8d97a2ef43a2f1b408bf420504a0448e47ad669a

Observation 026bd4cc-cd01-408d-a812-8b1271a30ada · outbound

This paper cites Prolific- dreamer: High-fidelity and diverse text-to-3d generation with variational score distillation,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Prolific- dreamer: High-fidelity and diverse text-to-3d generation with variational score distillation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.341108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:8323851456fcf2a65578c072f95ae0ea031052adec823033284e67a6e4423560

Observation a234b00f-cd1a-41e5-87d0-0ac3ff12296f · outbound

This paper cites Swiftbrush: One-step text-to-image diffu- sion model with variational score distillation,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Swiftbrush: One-step text-to-image diffu- sion model with variational score distillation,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.304868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:0c70b7490fb84858598168667b7a60831124de4ae269889755bc02105d668552

Observation aaa167fc-e9b2-4255-bd6e-f0f614ed01c2 · outbound

This paper cites Swiftbrush v2: Make your one-step diffusion model better than its teacher,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Swiftbrush v2: Make your one-step diffusion model better than its teacher,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.304672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:2633d7640ab4246892212223fd284ff9204ac431bfca33916c833a0fec9dd8da

Observation 0c131e16-0f6b-41d6-b6c7-5afd49369a9a · outbound

This paper cites Clotho: An audio captioning dataset,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Clotho: An audio captioning dataset,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.309270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:0ef98da006810c6841af8a41b57b4e77acfff88ba8e4c7ac3203a35c722039e3

Observation bc36acda-8e77-44d1-a3f6-0dcf8547c4e7 · outbound

This paper cites Audiolm: A lan- guage modeling approach to audio generation,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Audiolm: A lan- guage modeling approach to audio generation,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.343219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:9b2b5907d16bf12da519e0421c9dbac71b571f17330cdaa1bbf52e2b6028c5eb

Observation 1d1c1db6-e338-40be-8b43-b8c7fdd30f06 · outbound

This paper cites Diffsound: Discrete diffusion model for text-to-sound generation,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Diffsound: Discrete diffusion model for text-to-sound generation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.301107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:ffbeb9294a31b38c5421f1bd2c4e4f285f88fc38ca4a75f684f22dfff7ac48eb

Observation 3ab76f2b-28f4-4b3c-9d41-7c62a4abd398 · outbound

This paper cites Improved techniques for training score- based generative models,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Improved techniques for training score- based generative models,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.306635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:ce75ba0c4ccb990601ff441860c65dd751dbdab1c39e50dcb83716dafc26166b

Observation ce9a2467-275f-41c7-8a8c-ceb53551769b · outbound

This paper cites Text-to-audio gen- eration using instruction guided latent diffusion model,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Text-to-audio gen- eration using instruction guided latent diffusion model,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.294478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:b907219c089d25955ec9a549c339a93b8fea57bad844c6dd7b712b654d3ba80f

Observation da459659-1138-4f96-9d95-2aca7d617869 · outbound

This paper cites Diffwave: A versatile diffusion model for audio synthesis,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Diffwave: A versatile diffusion model for audio synthesis,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.302321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:eb6be07845958fd75909d69faaced4dc45166352890732fd2b934af8cb0b0f3a

Observation 789d25a4-d230-4375-8818-5a5c360ce065 · outbound

This paper cites Wave- grad: Estimating gradients for waveform generation,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Wave- grad: Estimating gradients for waveform generation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.307006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:68320709022161e4b105bbe7bd66f4060ca5b07d4516d2bb14f2dac2a0744f6c

Observation 87601a5d-2d35-48db-ab41-d9713bb0306b · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation High- resolution image synthesis with latent diffusion models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.296572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:8c92388c1641cd9209f64af2cae142f17cab3f864219c8b7ecd7c575730913b5

Observation be6dbe49-c11d-4e59-b529-b57578e5adfd · outbound

This paper cites Pseudo numerical methods for diffusion models on manifolds.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Pseudo numerical methods for diffusion models on manifolds

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.308408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:5011e3227c0fe7302ea574e277b203022700592cc524285ed2784654a610e76d

Observation a0c3e79c-78bb-4cf8-9702-3bcc8732884e · outbound

This paper cites Progressive distillation for fast sampling of diffusion models,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Progressive distillation for fast sampling of diffusion models,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.284468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:aaeaac9b3a07f8fdc9b0d635f100f4bace36cbe57310a6b09300e76b9a5583b2

Observation 223603cc-ab08-4380-889e-67d9ff11200d · outbound

This paper cites Dreamfusion: Text-to-3d using 2d diffusion,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Dreamfusion: Text-to-3d using 2d diffusion,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.286544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:91fe341bd951a7959915f0102d47037df5bc80dc8161b5bd6cab21bd33766145

Observation 0d4044a5-481a-4642-a0e2-76a719ae5961 · outbound

This paper cites Score jacobian chaining: Lifting pretrained 2d diffusion models for 3d generation,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Score jacobian chaining: Lifting pretrained 2d diffusion models for 3d generation,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.288395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:78e7cfee90f0e2853351b84714189dbc5db3ec4ee5a29f1015948df0540989c3

Observation bcfc807b-7570-4dd2-ae27-26f39dc238c4 · outbound

This paper cites Magic3d: High-resolution text-to- 3d content creation,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Magic3d: High-resolution text-to- 3d content creation,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.278698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:46a227bfb5f9902e1b40050718cd5e054272d1d9f3a994711e8e2866ef907224

Observation 77f13f2e-3f87-4f52-82fe-41b860d7d387 · outbound

This paper cites Latent-nerf for shape-guided generation of 3d shapes and textures,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Latent-nerf for shape-guided generation of 3d shapes and textures,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.259937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:cb46df95c7769c746dc5742c2794b915a4f01922df0d4371720efab44d47dc28

Observation dd4afe92-b547-4ade-b482-e9efa7288a82 · outbound

This paper cites Fantasia3d: Disentangling geometry and appearance for high-quality text-to-3d content creation,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Fantasia3d: Disentangling geometry and appearance for high-quality text-to-3d content creation,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.280578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:a996d8c64ae654739661ad55be46d81128d1267c9a312609caa1fd068cb85ffc

Observation 99fd5030-4c2e-4868-8fcb-2642dc5ad66b · outbound

This paper cites LoRA: Low-rank adaptation of large language models,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation LoRA: Low-rank adaptation of large language models,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.274475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:ae284c6bc4382d710eb230ca7c0042820eecce6d6cb5718dbb3973a1099c4b3a

Observation a4f16c29-6ad6-4037-9475-9f61d0ec565a · outbound

This paper cites Nonlinear total variation based noise removal algorithms.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Nonlinear total variation based noise removal algorithms

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.261798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:5ee55ebf73589a27983b72a15df1dd3a82103c9820b8b3043496d2b16437dfb7

Observation 96f26093-8399-48bd-8371-45709339df8a · outbound

This paper cites A duality based approach for realtime tv-l 1 optical flow.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation A duality based approach for realtime tv-l 1 optical flow

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.351427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:b26c2234bf4238008af68577a19cb7e8e5d3f8c7d3bfddf27b3a52314e54c946

Observation 0203c610-71f9-4a7f-a7ed-88409025fc71 · outbound

This paper cites Classifier-free diffusion guidance.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Classifier-free diffusion guidance

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.272520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:7fdb4601d9481ab5a8762e9d473cc112f1b699144acd825f358b8208a45c54a9

Observation 8fffec2a-fe23-46f6-8929-0262afc4f7b1 · outbound

This paper cites Decoupled weight decay regularization,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Decoupled weight decay regularization,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.268441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:e12c6cdc42838d862cf66b893c1d1ab04e00b9ab1533deb767294bb01cacf594

Observation 01c32da1-7a78-4783-a7e8-5acf733b18d1 · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Panns: Large-scale pretrained audio neural networks for audio pattern recognition

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.345855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:d66fba4c09efff21a20b11070cb44a89964efd9336218b59bf0a9f58b0615856

Observation f7563967-269c-47d5-8f3e-81c7b10add25 · outbound

This paper cites Cnn architectures for large-scale audio classification,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Cnn architectures for large-scale audio classification,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.266149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:c7cbbd0e3d8dfbf4f1e9c47733397b33cae029b07a8fdf0585cc48f3d1a72a78

Observation e1f9f09c-2b45-46e7-8b2e-877fe5b466fa · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.287372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:2026ad0171163490d601b1fbeb6c407ca4764efe3ac574025205dae2973b6dc3

Observation c9f9d3cf-b0fb-4d07-9553-ba8f8c833a22 · outbound

This paper cites Supercharged one-step text-to-image diffusion models with negative prompts,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Supercharged one-step text-to-image diffusion models with negative prompts,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.300261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:53fbd1f64b43994e8c67cf4f2f3f80abee25e16944831b333dde1b77205a33e3

Observation a9eb15af-54cc-46b1-bf8e-393e7099ffdf · outbound

This paper cites Prompt-to-prompt image editing with cross-attention control,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Prompt-to-prompt image editing with cross-attention control,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.282667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:0add72cdaa8ae5914ac11767b4525a315382b64bc76e6076c0326d79b06f3e93

Pith citing papers

No inbound Pith citation observations are available.