Pith. sign in

Paper Citation Record · LEDGER

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation

As of 13 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2606.31259.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.31259 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-01T03:45:40.636290Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy51
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5ea8592d-e260-4e9d-846d-c89534dc9c97 · outbound

This paper cites Make-an-audio: Text-to-audio generation with prompt- enhanced diffusion models,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Make-an-audio: Text-to-audio generation with prompt- enhanced diffusion models,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.254303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:733c456633590f8baefda6ab4f6ecf3dfd38e3eb2998f9ac81a9c7c9f3048d13

Observation e832f22f-5a78-44ff-b0b8-063574399ea2 · outbound

This paper cites Audiogen: Textually guided audio generation,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Audiogen: Textually guided audio generation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.256235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:8377113878b09f142b5739297340016961fdfa5af6cc1f44cfa1ba6049c378ca

Observation 4a620aa2-6883-4dea-aee5-5764793845de · outbound

This paper cites AudioLDM: Text-to-audio generation with latent diffusion models,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation AudioLDM: Text-to-audio generation with latent diffusion models,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.249336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:4e2bbd7ab402555d8ab3f712bc717dd653c9a9f08d7766b11bcb0eba367ad8a2

Observation 5b884f80-2ba9-48b0-abf5-378989279e80 · outbound

This paper cites Audioldm 2: Learning holistic audio generation with self-supervised pretraining.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Audioldm 2: Learning holistic audio generation with self-supervised pretraining

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.245047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:08cb54a58978080a9380f1351ab34f4476bf094b50f58f77d12e76a0e7034439

Observation 414dd74d-f4dd-452a-8876-22be12576b5d · outbound

This paper cites Any-to-any generation via composable diffusion,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Any-to-any generation via composable diffusion,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.263868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:ff00c528a56f6e997e1737b04740ea2cef72ac221a7cf331428e4e1451614463

Observation 49ad1c5a-e797-4ad2-92b8-5d4e199f14f4 · outbound

This paper cites Auffusion: Leveraging the power of diffusion and large language models for text-to-audio generation,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Auffusion: Leveraging the power of diffusion and large language models for text-to-audio generation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.258221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:9866df17f5518ef7d8a3656b4a0746636dbdc3b79f780b88b9e268b2bc29a3e8

Observation 44652a9a-ebf0-4e02-a6cd-f58d8276cf61 · outbound

This paper cites Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.332480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:097b2da6beb05065f3e73f4050ac3eba6fc0ee43a3cd181fa1e22adf32b2ce1c

Observation 28a7277b-f8bf-4cfc-8ef4-3f017db3587e · outbound

This paper cites Denoising diffusion probabilistic models.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Denoising diffusion probabilistic models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.279657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:f20b251d41cb186746692edf547e284808a823a5703fd19fe71b7c24dabff2aa

Observation 4d8ee700-e503-4295-9a56-83ddbf834e01 · outbound

This paper cites Denoising diffusion implicit models.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Denoising diffusion implicit models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.283418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:35f27f032e38a4e3e8ae50897c37aeabf73223cb712c4e0bb8cef6a09975f2a8

Observation ea950e4a-e0df-44d3-b54e-605c408b205c · outbound

This paper cites Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.330382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:6a6f6bbf83311a8a5bb4ca1c31302b41abb3c4b3c83429ecd92aa85b51e114a5

Observation e81d83d3-b117-46d2-8dc2-ed5211fa980c · outbound

This paper cites Dpm-solver++: Fast solver for guided sampling of diffusion probabilistic models,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Dpm-solver++: Fast solver for guided sampling of diffusion probabilistic models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.334663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:79726737d60a08d4236d33ef9dc82d8d04ef958bae23d74cefa41480bbdbbb9a

Observation 1c8c73f1-8741-4dfb-8095-ac27e46cd834 · outbound

This paper cites Elucidating the design space of diffusion-based generative models,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Elucidating the design space of diffusion-based generative models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.338922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:f366db5507e579481658b79d91abbf4c2c92c08ff763316b99967a5df84eebb4

Observation b85ed6f7-fc2e-45d7-b3a3-a855b02a826e · outbound

This paper cites Analytic-dpm: an analytic estimate of the optimal reverse variance in diffusion probabilistic models,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Analytic-dpm: an analytic estimate of the optimal reverse variance in diffusion probabilistic models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.323764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:27f5b361bf9b34cef0ed3f18c18a0ebd4e120b3529c4cff1e24ee6736113405b

Observation 8a6807da-9f65-4cc3-b049-c459aae6b750 · outbound

This paper cites Consistency models,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Consistency models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.353838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:2b3a54d294c36a780ca1799e3087ed232f27ce54ac5bcab794b8be5e07b1c753

Observation 100b490a-db43-47b8-a455-8d17d3293fc5 · outbound

This paper cites Consistencytta: Accelerating diffusion-based text-to-audio generation with consistency distillation,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Consistencytta: Accelerating diffusion-based text-to-audio generation with consistency distillation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.325851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:af0164f051f38dbd60cf7616bdd340091ed2a30528e880555676e5749102f1de

Observation 7e86ca38-588f-42c1-b013-dbf301d508dd · outbound

This paper cites Audiolcm: Efficient and high-quality text-to-audio generation with minimal inference steps,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Audiolcm: Efficient and high-quality text-to-audio generation with minimal inference steps,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.348917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:ac8e87bef87a07871186c3070234720ad47f4cc474f3d159b1fa788892bfd7c9

Observation 29755e60-187a-4df3-8d9b-4a82e9429a5b · outbound

This paper cites Instructpix2pix: Learning to follow image editing instructions.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Instructpix2pix: Learning to follow image editing instructions

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.310263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:1a7bf7e63c0a14998debfcafd0d94f341242f1bcffc33769fd11d0efc57fda7e

Observation 0ce01166-648a-45b2-9f0d-a3d5aa2231c5 · outbound

This paper cites Ts2f: Text-assisted speech-to-face generation,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Ts2f: Text-assisted speech-to-face generation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.321940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:f9a8271dfb432e5ce9cc780b8317faee5e9db5bdb9133a190e8a7795a971339e

Observation 147d604e-9d83-49ac-b59a-af2cdb335604 · outbound

This paper cites AudioCaps: Generating captions for audios in the wild,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation AudioCaps: Generating captions for audios in the wild,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.319977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:e30b8ae9b7f2c274ddc73627fe69ef1ce1df8527bab7a83d5c95156f0e3cbd6d

Observation 4c4083aa-ab37-4cd7-9df1-2c2dc7bbe03d · outbound

This paper cites WavCaps: A ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation WavCaps: A ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.315584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:9396d7ea188c6889064303e5181f3d144caacdd14b1e1a18ef482d114e0133df

Observation d91e1fe3-7281-4a62-bf7f-fdb63386e5b0 · outbound

This paper cites Audiosetcaps: Enriched audio captioning dataset generation using large audio language models,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Audiosetcaps: Enriched audio captioning dataset generation using large audio language models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.317779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:682fe04f7de6cd1d3358c5c0a18f1ae4e6c9c0c2ba6e18e2e97706eab1cfd543

Observation adc3b1f5-3410-4d24-9607-0b7b30dd42f7 · outbound

This paper cites Journeydb: A benchmark for generative image understanding,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Journeydb: A benchmark for generative image understanding,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.328463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:edcaa57c924fe20f01d980c28956048414e53bb7fb46fbc053506479914443d7

Observation 94db4798-f07f-480b-98ea-bb6f2d572426 · outbound

This paper cites Laion- 5b: An open large-scale dataset for training next generation image-text models.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Laion- 5b: An open large-scale dataset for training next generation image-text models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.336702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:390fccf3e0c3d6b5f1a83090461f445230177cacae27caafaaaa43cdbe70208c

Observation 026bd4cc-cd01-408d-a812-8b1271a30ada · outbound

This paper cites Prolific- dreamer: High-fidelity and diverse text-to-3d generation with variational score distillation,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Prolific- dreamer: High-fidelity and diverse text-to-3d generation with variational score distillation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.341108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:c0e2b4bd82a396c75c79f9872b2628ea7820cb9ec9a5d94b880f4ff3a43a2069

Observation a234b00f-cd1a-41e5-87d0-0ac3ff12296f · outbound

This paper cites Swiftbrush: One-step text-to-image diffu- sion model with variational score distillation,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Swiftbrush: One-step text-to-image diffu- sion model with variational score distillation,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.304868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:8fdd60b6149307b8dcc55bdf5e55bed19ade59e496a5bc99f9c3dcb8c06202eb

Observation aaa167fc-e9b2-4255-bd6e-f0f614ed01c2 · outbound

This paper cites Swiftbrush v2: Make your one-step diffusion model better than its teacher,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Swiftbrush v2: Make your one-step diffusion model better than its teacher,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.304672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:ac67ec5929806f833bff5ec668d6261e43fa1c55d76de85c6e3e5b697efc0935

Observation 0c131e16-0f6b-41d6-b6c7-5afd49369a9a · outbound

This paper cites Clotho: An audio captioning dataset,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Clotho: An audio captioning dataset,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.309270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:5a9c9aeb7518851b44b657426137864e5ff0609c2cb1ef31093ec92e7a05fd25

Observation bc36acda-8e77-44d1-a3f6-0dcf8547c4e7 · outbound

This paper cites Audiolm: A lan- guage modeling approach to audio generation,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Audiolm: A lan- guage modeling approach to audio generation,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.343219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:5ecdca33d9a5dae743d4670e297436b8375d0eb6af14167fc2c834da539b16c3

Observation 1d1c1db6-e338-40be-8b43-b8c7fdd30f06 · outbound

This paper cites Diffsound: Discrete diffusion model for text-to-sound generation,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Diffsound: Discrete diffusion model for text-to-sound generation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.301107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:654939e3d49908e910ed07e1aef3f6d003af4e9a05e60c86025d66f5fd355640

Observation 3ab76f2b-28f4-4b3c-9d41-7c62a4abd398 · outbound

This paper cites Improved techniques for training score- based generative models,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Improved techniques for training score- based generative models,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.306635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:8e999a13298260d777d4bdd28745871e8feaad685d3f98c0e86cc472e2b460dd

Observation ce9a2467-275f-41c7-8a8c-ceb53551769b · outbound

This paper cites Text-to-audio gen- eration using instruction guided latent diffusion model,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Text-to-audio gen- eration using instruction guided latent diffusion model,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.294478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:a79fea5b543ca412fbc0dc3290769081469cf129b4fcf4dc0696439456252767

Observation da459659-1138-4f96-9d95-2aca7d617869 · outbound

This paper cites Diffwave: A versatile diffusion model for audio synthesis,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Diffwave: A versatile diffusion model for audio synthesis,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.302321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:b15cd7e125f9ddd941be7d5af1bd4414c4c319b99dadc85760f8ca50bea374b9

Observation 789d25a4-d230-4375-8818-5a5c360ce065 · outbound

This paper cites Wave- grad: Estimating gradients for waveform generation,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Wave- grad: Estimating gradients for waveform generation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.307006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:6574d5fd9bf00fd0a259af21776db6328114936d75e9da466070e38cd9462301

Observation 87601a5d-2d35-48db-ab41-d9713bb0306b · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation High- resolution image synthesis with latent diffusion models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.296572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:e1a891bee0f2af6e74393f45ef46691eaae2531f96e0493ffbf24c75560c5032

Observation be6dbe49-c11d-4e59-b529-b57578e5adfd · outbound

This paper cites Pseudo numerical methods for diffusion models on manifolds.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Pseudo numerical methods for diffusion models on manifolds

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.308408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:8d5c6bb5f80d0197a414a2007551741b3cea45604e3aa74c4d5790801c40964c

Observation a0c3e79c-78bb-4cf8-9702-3bcc8732884e · outbound

This paper cites Progressive distillation for fast sampling of diffusion models,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Progressive distillation for fast sampling of diffusion models,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.284468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:99df85179bf83a4296c43905c352c6e33b2c36e39e5df92b44b2946023f87991

Observation 223603cc-ab08-4380-889e-67d9ff11200d · outbound

This paper cites Dreamfusion: Text-to-3d using 2d diffusion,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Dreamfusion: Text-to-3d using 2d diffusion,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.286544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:8d5196b1568206e71906359c6bd5edc1b869b234f077db5f11ae68878a8210f0

Observation 0d4044a5-481a-4642-a0e2-76a719ae5961 · outbound

This paper cites Score jacobian chaining: Lifting pretrained 2d diffusion models for 3d generation,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Score jacobian chaining: Lifting pretrained 2d diffusion models for 3d generation,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.288395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:c467b0256267f3c4bcf877d57c155d6df59e5df36cc818719ccc4662e7a806ce

Observation bcfc807b-7570-4dd2-ae27-26f39dc238c4 · outbound

This paper cites Magic3d: High-resolution text-to- 3d content creation,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Magic3d: High-resolution text-to- 3d content creation,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.278698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:71114037e0fb21282d04501c1130608e6ab887d8e65cd6bdce8036c37aecb17c

Observation 77f13f2e-3f87-4f52-82fe-41b860d7d387 · outbound

This paper cites Latent-nerf for shape-guided generation of 3d shapes and textures,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Latent-nerf for shape-guided generation of 3d shapes and textures,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.259937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:7cb75e2ac6bbd9f36f0457d64b57b7bff23191bd1e57638289ca758a70e3ac2e

Observation dd4afe92-b547-4ade-b482-e9efa7288a82 · outbound

This paper cites Fantasia3d: Disentangling geometry and appearance for high-quality text-to-3d content creation,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Fantasia3d: Disentangling geometry and appearance for high-quality text-to-3d content creation,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.280578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:570de2faec070909aa5daa3bd2f11528bba21c27f4df89e0a5bf2f37f81dad2e

Observation 99fd5030-4c2e-4868-8fcb-2642dc5ad66b · outbound

This paper cites LoRA: Low-rank adaptation of large language models,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation LoRA: Low-rank adaptation of large language models,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.274475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:39ee67749d34acf0319ec22cb03b9bcd88938a8c649950aa9f697a3cecdf7fba

Observation a4f16c29-6ad6-4037-9475-9f61d0ec565a · outbound

This paper cites Nonlinear total variation based noise removal algorithms.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Nonlinear total variation based noise removal algorithms

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.261798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:9f06eb1b8d48b1ccae5177f360958ff9f79ba16c5bc95be00dd8354d2ec9388f

Observation 96f26093-8399-48bd-8371-45709339df8a · outbound

This paper cites A duality based approach for realtime tv-l 1 optical flow.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation A duality based approach for realtime tv-l 1 optical flow

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.351427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:3ef6e9df590180578fd9837599f8d474033d7ca9dc750eef9b478ba5f2054e3e

Observation 0203c610-71f9-4a7f-a7ed-88409025fc71 · outbound

This paper cites Classifier-free diffusion guidance.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Classifier-free diffusion guidance

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.272520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:0b4f46a7e5ea7913e4167747bb80c39223987bbd598f6c0ee2c29ed5ce8295e8

Observation 8fffec2a-fe23-46f6-8929-0262afc4f7b1 · outbound

This paper cites Decoupled weight decay regularization,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Decoupled weight decay regularization,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.268441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:49e973ad8be0babd6241a8cfe8621c6cb9bb0c1baa56f4bb9dc7fd091ff0ecec

Observation 01c32da1-7a78-4783-a7e8-5acf733b18d1 · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Panns: Large-scale pretrained audio neural networks for audio pattern recognition

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.345855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:c18291935647bd2890cd13590827e15ca360672a5765aafa2cec838284c66c44

Observation f7563967-269c-47d5-8f3e-81c7b10add25 · outbound

This paper cites Cnn architectures for large-scale audio classification,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Cnn architectures for large-scale audio classification,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.266149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:e0045fa9defd16b755774a7fc7ec25b3ec11260eef5d452f0a97a4a6c7372994

Observation e1f9f09c-2b45-46e7-8b2e-877fe5b466fa · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.287372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:e1865a592671fddf85b7191918fad9ecd93b1f8700f5e77ea23d41170abcdbba

Observation c9f9d3cf-b0fb-4d07-9553-ba8f8c833a22 · outbound

This paper cites Supercharged one-step text-to-image diffusion models with negative prompts,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Supercharged one-step text-to-image diffusion models with negative prompts,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.300261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:d0c4298d1dbac7830d9ff016c5cdc76c0375b103607ee845e87a11b87db75e8f

Observation a9eb15af-54cc-46b1-bf8e-393e7099ffdf · outbound

This paper cites Prompt-to-prompt image editing with cross-attention control,.

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation Prompt-to-prompt image editing with cross-attention control,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T01:23:07.282667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T03:45:40.636290Z digest=sha256:8277f1f958c12fc8cc5249932d272aa0bb6d218555201fa6b59e8d81e9e8e1af

Pith citing papers

No inbound Pith citation observations are available.