Pith. sign in

Paper Citation Record · LEDGER

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

As of 12 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 26 inbound Pith citation observations for arXiv:2412.21037.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.21037 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:21:34.742952Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:12:09.446928Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.484714Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact2
  • verified fuzzy1
  • unresolved61
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dfbe94c7-e846-4d66-a21b-4e04b05ebcc5 · outbound

This paper cites write newline.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.457797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.457797Z digest=sha256:02ff133a83b92078781a0491e77f693a1b03c071ed419329ccb2a4410daddbcb

Observation 714224b0-ccce-4531-8ba9-b03a719a33a4 · outbound

This paper cites Building Normalizing Flows with Stochastic Interpolants.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Building Normalizing Flows with Stochastic Interpolants

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.463848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.463848Z digest=sha256:fd4411e656a522b76702e1c1bc820764adeb8a9b2ad3907de4db758fd5164c8b

Observation df247e50-6c54-4caa-aa95-a138a2e0471f · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.469104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.469104Z digest=sha256:d2ccf83961e29750f9a25c476120d1f47074d7f3fe9557c6adde2e4bf717c67d

Observation 5d977bf6-f9ce-49e1-b0b9-353689b69edf · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.473968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.473968Z digest=sha256:2dd79352b2804515a3477e048b050308c9e86122f1a22ae0cb8377eb27e74477

Observation e84b5fc1-60f1-473f-9258-c9cc3ef23958 · outbound

This paper cites Qwen2-Audio Technical Report.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Qwen2-Audio Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.478441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.478441Z digest=sha256:ddad5054db10bb24159017b8fcfc958716cd9cb306dfd8dbcc2614f75cdf93b3

Observation 1f9b7d22-8a77-45a1-aa56-62b4f50e61d2 · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Scaling Instruction-Finetuned Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.484425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.484425Z digest=sha256:7c7d88996d7e4252ef3d0d2f5a7d049ab602dbd71bc1a5ed7d1b80dc30386b07

Observation 17d9e8d6-fb67-4da7-b3a6-bc6ace5d818f · outbound

This paper cites Simple and Controllable Music Generation.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Simple and Controllable Music Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.489417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.489417Z digest=sha256:742859dc37a4ff6a65cfd1d061519821dea1cf152180b07fd7ab07d7bf6e5b72

Observation 94c062f7-c4eb-4548-b789-3bb438cb5475 · outbound

This paper cites Look, listen, and learn more: Design choices for deep audio embeddings.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Look, listen, and learn more: Design choices for deep audio embeddings

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.494946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.494946Z digest=sha256:22739e3b280b14872afa44d3f23b80d05feaf9cde95f43e2a6a34c370d060774

Observation 08d2f6ec-be76-42ff-b0b7-0079eadb4f91 · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.499424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.499424Z digest=sha256:3b64f91a9a077ad2a9fed8c2af772117dfb0d13089868d1af06f5dbe66192936

Observation cf778c2c-57b9-452a-b9f4-639f4e1eb360 · outbound

This paper cites Fast Timing-Conditioned Latent Audio Diffusion.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Fast Timing-Conditioned Latent Audio Diffusion

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.503571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.503571Z digest=sha256:413b182e5b3d06547024fcb05bb3d0ba86def8087b1fa5ac735c4e20305433ab

Observation bd53f9b4-b7ed-4ea9-ac89-ca6ec49087b2 · outbound

This paper cites Long-form music generation with latent diffusion.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Long-form music generation with latent diffusion

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.508244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.508244Z digest=sha256:03c1851b2d272e99112754a72d8e86855f354eab495e264e84fef69df4be59c3

Observation 841081a2-dbfe-4ff0-8364-87da6a98e644 · outbound

This paper cites Stable Audio Open.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Stable Audio Open

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.512529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.512529Z digest=sha256:2488cba94df720e43d52108366187fd3504c30d73ff4c79aabe71fc632b8bbe2

Observation 7fb785a6-bb01-44e5-b7b6-86b9edf7af67 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Scaling Laws for Reward Model Overoptimization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.516577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.516577Z digest=sha256:217acd65e3d8a261463eb354a99fd2350fb33cf3bb06bb301b19299602463002

Observation fee7105a-96f9-46eb-bcba-338e0d6ac279 · outbound

This paper cites Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.520734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.520734Z digest=sha256:95498de90288f76135bfdad2de5d7af5f5c85bc9902835dce7de3b4c0a7d9988

Observation c3475995-b8cf-4317-a40d-374602111266 · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Reinforced Self-Training (ReST) for Language Modeling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.525539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.525539Z digest=sha256:73f27aa956023a4249eebc27a78ee3111846584766f94a9d09c2c4e351c65dfa

Observation 269f0661-ebb8-42a7-a46b-bb353b8d3ac2 · outbound

This paper cites Efficient Diffusion Training via Min-SNR Weighting Strategy.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Efficient Diffusion Training via Min-SNR Weighting Strategy

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.529354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.529354Z digest=sha256:d1984c269c1a316b38c7bcf407e6257f8a1626d46052014120c1c5e884bf3a66

Observation ef2c5bb6-dc29-4bed-b36b-22463c2671dc · outbound

This paper cites Classifier-Free Diffusion Guidance.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Classifier-Free Diffusion Guidance

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.533363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.533363Z digest=sha256:269c1fc5b73475e0be5eb1c5f0e1c40cfed2d3ff4397f0651eded70c814bc95f

Observation 5a84fea2-0bc2-4a98-ad1c-554f2e3bc54b · outbound

This paper cites Denoising Diffusion Probabilistic Models.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Denoising Diffusion Probabilistic Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.537595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.537595Z digest=sha256:bad6d5ca08a86fdd5ba7d397796c2134c14ee1b367b1449bd6b1238d8db82dba

Observation 1b5ff383-615e-4c2d-b58c-a1f4a87c0976 · outbound

This paper cites Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.542147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.542147Z digest=sha256:57bd9629e9eb65f56d4fc9e4b8c1969575dff2a59769266b7657a2c57657f099

Observation 5f2e0827-c153-4474-98f3-ecf7546a84f9 · outbound

This paper cites Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.546697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.546697Z digest=sha256:45c32531a358341f4a6db194b27dab8bad861fee3983cac57c6fe47d81297514

Observation d70bc48e-facc-4e64-b3f0-07b86554d65b · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.551154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.551154Z digest=sha256:61b1f47badc483bd9c82e825473410b504dd3dae042e540994826baf86f1f637

Observation 3406b416-db76-49db-bef0-613d1fef4426 · outbound

This paper cites Elucidating the Design Space of Diffusion-Based Generative Models.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Elucidating the Design Space of Diffusion-Based Generative Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.555413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.555413Z digest=sha256:72f39168366d9d649e3e7dea587915963b640a8948728d5c4cc54109ecca823e

Observation b74a1729-c083-437a-9882-6e724ec345fc · outbound

This paper cites A udio C aps: Generating captions for audios in the wild.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization A udio C aps: Generating captions for audios in the wild

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.559672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.559672Z digest=sha256:6aa099a5b6858de7c697963b86adae41a9e04073a2bfc3fe89b92098546ac3fa

Observation 7f4c4502-1538-4d70-9098-679f50dcd79b · outbound

This paper cites sDPO: Don't Use Your Data All at Once.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization sDPO: Don't Use Your Data All at Once

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.563825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.563825Z digest=sha256:93d5ec1d0c22d19a9ad08fe76786db73b3eaf65a0e87b8dc278387d6cbe26547

Observation 64cb2faa-aacb-4b6f-950b-8942914e7bf4 · outbound

This paper cites Adaptive non-uniform timestep sampling for diffusion model training, 2024 b.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Adaptive non-uniform timestep sampling for diffusion model training, 2024 b

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.568233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.568233Z digest=sha256:6ebc1485a27088518b2ecb2296d2e2786545c93808dfdee47c09dc8a92a1a393

Observation 7fff4a72-c2a1-43a8-a19b-7126519d6feb · outbound

This paper cites Auto-Encoding Variational Bayes.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Auto-Encoding Variational Bayes

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.574308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.574308Z digest=sha256:56270b882efe987726324f83123919fb9c3ca2b1390dae27acd42cb1c21d1f27

Observation f156cb9d-4ed3-4cd2-85d6-228143149360 · outbound

This paper cites Improving Text-To-Audio Models with Synthetic Captions.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Improving Text-To-Audio Models with Synthetic Captions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.578346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.578346Z digest=sha256:5cd19cdd4f4b09e0209f0f65542d197283beaffeeeb213df7dac1173f63dae39

Observation 701f1071-6946-4ab3-98b6-02b5ecff497b · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.582426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.582426Z digest=sha256:d093f52f4c44a03d4adc59548be4606912ebe6c85c13abf4309aa79b792677ba

Observation f15dbce5-7bd9-4a40-bc3d-77dd6e8c322d · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization RewardBench: Evaluating Reward Models for Language Modeling

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.586528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.586528Z digest=sha256:1e579c4a1f9bbaf80c49dc31fc3237dc4fb4d7e209cc044ed5a42826403c131f

Observation b360166c-1f97-4adf-93ca-1ce8e1519132 · outbound

This paper cites Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.590797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.590797Z digest=sha256:90426cedf2a77bc9472837f34bb17b066ce0fe55b0c363ce9c07506951aec4b9

Observation 2ef549bb-797d-4afa-b2bf-e04e44bff316 · outbound

This paper cites BATON: Aligning Text-to-Audio Model with Human Preference Feedback.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization BATON: Aligning Text-to-Audio Model with Human Preference Feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.596114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.596114Z digest=sha256:b08d74ce9ab4636094226affa92db59dbbdef5e95217543c7a0789c7a047d04b

Observation 42ce1007-43b2-4804-8fd2-be2f7163b7fb · outbound

This paper cites Flow Matching for Generative Modeling.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Flow Matching for Generative Modeling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.601091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.601091Z digest=sha256:72f96aae83a7fe2b9709e3f818489904357bb24f3fa06746f918530fe8586a73

Observation 84334678-d58e-4b68-a0fa-7ada18b2f96e · outbound

This paper cites Generative Pre-training for Speech with Flow Matching.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Generative Pre-training for Speech with Flow Matching

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.605049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.605049Z digest=sha256:1156364731dd01dee83ec875f09ec83cc1350b63011cfecfe158fa48439e7d09

Observation 767aad89-3212-4799-8cce-c1d3d53b323c · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.609758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.609758Z digest=sha256:11f790fcf32f3e6d19e1524faf06738abac7b5ff6bdf093a878f02640b1e8a07

Observation 4d751b30-f5da-43c1-a422-3fc34a613848 · outbound

This paper cites AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.613652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.613652Z digest=sha256:6e625f74384f476f063035461047d63ae208a11155e919dbda97af83a4b36297

Observation c71edb85-6817-4f72-9c8a-436701d8accc · outbound

This paper cites FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:21:35.061744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-10T23:21:34.618608Z digest=sha256:0daa7ee133daa640245a61e0ed2e9ffde457397a5134a9db2098961a58f349cd

Observation fe0d98a4-e94f-4339-8982-dbed3a07f311 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.623344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.623344Z digest=sha256:127c0e3f066a37bb03249e9260d4ed5cf37b3d6105615f78cb91704d46f6b8eb

Observation 1f20d4be-a857-4da3-96b6-70a59344e744 · outbound

This paper cites Decoupled Weight Decay Regularization.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Decoupled Weight Decay Regularization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.627496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.627496Z digest=sha256:65ecc764a449ab211c962eb44d335c2014717f673c01f9332b21f90b5cb47ba5

Observation ef0aff9c-f639-460f-a2aa-732ccc90f8ec · outbound

This paper cites Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.631417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.631417Z digest=sha256:a08c1d6e3e106a2fc39565577cbfca05ded426e461d6feb9c214f67b5dc9d16b

Observation cfa23ff7-b5ec-4705-9689-06d6f6473482 · outbound

This paper cites Plumbley, Yuexian Zou, and Wenwu Wang.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Plumbley, Yuexian Zou, and Wenwu Wang

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:21:35.586953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-10T23:21:34.635465Z digest=sha256:32ecf5f2768a4c4bbae52de6cee12d01ffa44124924a0a884d69ea9d4617d2e6

Observation 4490b601-2476-42cd-88e9-4543666b1aeb · outbound

This paper cites OT-Flow: Fast and Accurate Continuous Normalizing Flows via Optimal Transport.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization OT-Flow: Fast and Accurate Continuous Normalizing Flows via Optimal Transport

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:21:34.999252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-10T23:21:34.639882Z digest=sha256:d01aab829163e31aaaf3955d09e04b72952ac7f95b844309f077a1ec009a03bd

Observation 551e976d-da4d-4e16-bb2a-a38aec00af94 · outbound

This paper cites Training language models to follow instructions with human feedback.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Training language models to follow instructions with human feedback

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.644929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.644929Z digest=sha256:8c3107177820bf08dc9e5b34f4497528ad0796717605fdae2cb3db440db28bf6

Observation bcd7b2ad-fd7d-478e-9c4c-b14d9e67b33a · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.648810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.648810Z digest=sha256:23d101242d8de996f544cb4e49e40884197213044ba0ee173e0d416bc787af3a

Observation 04befc61-98dc-4740-9bff-2cc8530727dc · outbound

This paper cites Iterative Reasoning Preference Optimization.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Iterative Reasoning Preference Optimization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.652625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.652625Z digest=sha256:75e8addfa3111dd8a6915c2657d644d168455b523250e59929861459ca99000a

Observation 62450e1f-c527-4a7d-b71c-2069d03824ee · outbound

This paper cites Scalable Diffusion Models with Transformers.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Scalable Diffusion Models with Transformers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.657081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.657081Z digest=sha256:ad8c46e572178de7e3515bf94dd0353fbf48a0fe6f52791965c82dd78df0af9f

Observation b9ef3da2-eead-474a-8b86-200a858c4dc4 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.662433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.662433Z digest=sha256:973d3e9927b2725de9154813aaf3972d1d5725ba370a5e5f4bba5f4ef5b1a0c2

Observation 59acef81-8414-467b-9b42-da40ddef3b22 · outbound

This paper cites From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.666487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.666487Z digest=sha256:26854ce34dea8e041ba5df11653254c58ea0287f8a26e6405ff8d2c93e5df580

Observation f9127c2c-1a6b-4df6-b7a3-3b9c7e20e8fb · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.670941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.670941Z digest=sha256:a7f0173649921cf9003aa57d0f77d6925208eaef91270bb4bc29be9964615ba2

Observation 8b0373eb-58c6-4b89-9b4e-3d54e3ea2033 · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.675200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.675200Z digest=sha256:285ed3881fb3344a3c3d5b4f493a5926c6bf6c3e8d2e8a4af3ce433a30c015fd

Observation 7ab6fe7a-80c9-4823-9087-910b88eb3cbe · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization High-Resolution Image Synthesis with Latent Diffusion Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.679776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.679776Z digest=sha256:2434af4e8277389ff78adda11484de9f318e716662aaca522ba6428a439d6fd1

Observation bd0f18e1-e25d-4a21-9452-e4890320d749 · outbound

This paper cites Improved Techniques for Training GANs.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Improved Techniques for Training GANs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.684082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.684082Z digest=sha256:c64d2f116e5253ba11a53d9301a6183ca01a6d962b21ba43d6425c5fe2a5e5ec

Observation 00f413f2-7dbe-4219-9356-d4a9674bf326 · outbound

This paper cites Denoising Diffusion Implicit Models.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Denoising Diffusion Implicit Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.688478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.688478Z digest=sha256:31351ad2df6a0f32bc2ff5a3f7465fceaa5d75bdabdff1b1333c9a9c7f4683a4

Observation 45071ce0-1e92-4705-8903-b192bfc6f152 · outbound

This paper cites Generative Modeling by Estimating Gradients of the Data Distribution.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Generative Modeling by Estimating Gradients of the Data Distribution

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.692979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.692979Z digest=sha256:47706a1b90af852d43560dd7dd5ded8137c3758ddfd5b9e75aa0bc1d2e0fa4fd

Observation 0043f3bd-eebd-4220-885b-bac5c6b7bd1f · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.697783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.697783Z digest=sha256:3838c4fce6db91344f604ecbaceef17a10459d7b5858b8eadf96968fd051a568

Observation 56e9bb92-efc9-44a4-bbb8-2d5ff679c976 · outbound

This paper cites Attention Is All You Need.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Attention Is All You Need

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.702355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.702355Z digest=sha256:22e0460d38077a62a12eba8607637a3d8721c44fc40ee2a735300ce9b592fc7c

Observation 9ae97aa0-ae99-4f7c-87d9-bcd2b1685718 · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.706605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.706605Z digest=sha256:091e147d01005e46988088d952e83a2f7a6073c6bd5f796d5fa3de46c1aeeac3

Observation fb1276ca-32c7-4723-80ce-0182569c3495 · outbound

This paper cites Diffusion Model Alignment Using Direct Preference Optimization.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Diffusion Model Alignment Using Direct Preference Optimization

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.710431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.710431Z digest=sha256:91250e4bb60057a3ef75212187fc3c7630dd16322a0e4b445d74aed5109e2550

Observation 0965e152-74d5-4b7c-bd7d-67728bb17a1e · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.714901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.714901Z digest=sha256:332a319e3819654e68de2cb407fc349c1c192edb1af53fa3e848dd6b97bf82a7

Observation f5155aad-2141-441c-b2f4-1c807fdde1c4 · outbound

This paper cites Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.719059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.719059Z digest=sha256:4da95685caea739b65ded5bb71b469f9da1945767545460598e6a6054f028940

Observation 0b1e123c-b37e-48b3-ba86-882ec43b24b5 · outbound

This paper cites Self-Rewarding Language Models.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Self-Rewarding Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.723994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.723994Z digest=sha256:b2018b3e3b15cb636c42c197bad3c641495e874eb92ce4e1647ec568ee185a30

Observation 098672e5-dc02-45e7-be75-259d3cc2dafa · outbound

This paper cites STaR: Bootstrapping Reasoning With Reasoning.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization STaR: Bootstrapping Reasoning With Reasoning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.728757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.728757Z digest=sha256:b7612cf5a3e090c2ea7a7fb00cbc0a763255e153a17420a316dd635278854705

Observation 9d46ae32-bd57-4a67-a35a-74489ec233b6 · outbound

This paper cites @esa (Ref.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization @esa (Ref

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.733438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.733438Z digest=sha256:77970ec9508d32aa5d69efa3b65c99a7970385f2b64b6550972f5f2e8c3aee58

Observation de32df46-d16a-4d27-9505-3d1ed74b925e · outbound

This paper cites an unresolved cited work.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.737887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.737887Z digest=sha256:62d78ea57d9b1bbe4e8a0b2431fcd80b16cd2220636580e61c7c1b609097d3f3

Observation 679cdb67-48ae-4f5e-ad85-54c4a6ccfd9f · outbound

This paper cites an unresolved cited work.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.742952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.742952Z digest=sha256:ea1d03b922740202a0a596949c9d6883cba24567ca23535e6b1f9c1d8ab4ddd2

Pith citing papers

Observation b1f054af-4f43-4264-b9f8-10dfb0aa539c · inbound

Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping cites this paper.

Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:41:36.642247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T13:38:17.640841Z digest=sha256:570faadb324e87ef7380e704b143e652d80df2afefd38b4c4ffc1171864287fe

Observation ae659b54-8ae6-4d0f-8e35-7eb036357a6c · inbound

AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation cites this paper.

AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:09.446928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:09.446928Z digest=sha256:107221465c51ed0856c07594cd78088445baa1097d6319c26fbcb9022a1a1f21

Observation d8a1eb31-d758-489d-9bb4-91319a8dc326 · inbound

SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations cites this paper.

SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:04:00.556655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:04:00.556655Z digest=sha256:dc278c7c2af2afda5978bdc812ac485a57275f48542e1df8d69e96183ecb4257

Observation 89187b51-2060-42e8-8272-45b469a396f2 · inbound

JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment cites this paper.

JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T13:14:37.521235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:14:37.521235Z digest=sha256:3c8d8f04a6498d2acaecd031ac3c9c29fec800ae4afaf601b0e47239f3236d8c

Observation a9360261-9ba8-4900-a57c-82210bd60359 · inbound

RFM-Editing: Rectified Flow Matching for Text-guided Audio Editing cites this paper.

RFM-Editing: Rectified Flow Matching for Text-guided Audio Editing TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:36:37.020189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T16:34:17.431134Z digest=sha256:90c68e27645b2468b06b2a1c28061b27ee57513e9482037cc48c3c0f4bf0f947

Observation 7f2626d6-579d-4430-8d7b-a17467c8e3bc · inbound

AudioMoG: Guiding Audio Generation with Mixture-of-Guidance cites this paper.

AudioMoG: Guiding Audio Generation with Mixture-of-Guidance TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:23.820034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-18T13:10:18.700497Z digest=sha256:a0779423b295b7f85711cf5a97da4d3514ffd6695e86b8f2fcf7a466b505bb34

Observation 8393a281-f852-4a31-a326-1d41ad32bbb6 · inbound

SemanticAudio: Audio Generation and Editing in Semantic Space cites this paper.

SemanticAudio: Audio Generation and Editing in Semantic Space TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T07:06:33.216494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:06:33.216494Z digest=sha256:35ddbc3835272ee551a07ab7bcaf4fb66f86a0b2e9b842ce714fed7d0400c4e6

Observation 974ba821-f86c-43ec-aec3-c44a3ab2bc2a · inbound

Woosh: A Sound Effects Foundation Model cites this paper.

Woosh: A Sound Effects Foundation Model TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:53:15.849552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T20:51:08.144573Z digest=sha256:17730c08c277dc3d6115cb00cd9bfc303a8f8df2f7fc2fbcd021bbae2ce66e04

Observation 0064991e-42bb-4efb-b39c-3d94b6b6da73 · inbound

ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling cites this paper.

ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.054519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T09:22:10.483263Z digest=sha256:b747169ca15257828a851452d86324bea1986dfe52c5adb05ce621f2abf1b918

Observation 2a3f17f9-1852-4239-ae25-f01a34248660 · inbound

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation cites this paper.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:46:44.885236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:6836ee90714fffd9044c52b0d20e28af26d7d830c4716e4b21a4a69d88791602

Observation d47cc885-7e9c-4793-8dc5-b325ba92850b · inbound

AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling cites this paper.

AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.274264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T05:06:54.972846Z digest=sha256:358fb95381bf9be6be0f54ebcf9f628dab8056c33b491e791976a8797c21685e

Observation 54e22aad-cf03-451f-a628-d13534dc5ff2 · inbound

AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling cites this paper.

AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:29:52.743864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T08:28:21.790110Z digest=sha256:72397b2ea95b89a192c54f1d45f603401828a45635157a1d332174f086225c55

Observation 669b3f78-a6b5-4039-bc2f-28ca9bcf3bef · inbound

Dasheng AudioGen: A Unified Model for Generating Coherent Audio Scenes from Text cites this paper.

Dasheng AudioGen: A Unified Model for Generating Coherent Audio Scenes from Text TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T11:03:18.758307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T10:56:00.765400Z digest=sha256:29b2f53ca26835ea9f8e4c32f67ec2ed8d50ba44f5d1344bd55717bbd98dcfc8

Observation d4ff063d-a656-4c2b-81a0-0d5a6086c142 · inbound

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment cites this paper.

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:12.629539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T21:12:19.893944Z digest=sha256:19812e5cb215d0e59daaf1b27c746730cd6c1ddb0c0a3620aeb030cf719621e5

Observation eed75720-14f1-41b5-89e4-64dd56b13bc4 · inbound

UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion cites this paper.

UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-06-28T20:52:37.895615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T20:44:20.190064Z digest=sha256:0a1bb2c52664743bb4294d5092d431459407cdfa3ccdf208c6b0dc5d55e26ca0

Observation 7b58e407-6dba-4f8f-8277-b69eda7cb2e4 · inbound

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation cites this paper.

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:28:18.809247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T08:04:48.283908Z digest=sha256:8a25ead5774ed4db5701b9dc9e462c3e0041ed6efd486f360bdb38efc46fd62f

Observation 90e5fdab-695d-4dd1-806f-14f8978ac129 · inbound

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers cites this paper.

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:39:39.325372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T15:51:48.672086Z digest=sha256:d78d9ff9beeef6eba743cd83d712a11422bf51324b23a17e67e0491b37014361

Observation 8592f3b5-52e4-43f6-ba32-29cdba770adc · inbound

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers cites this paper.

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:39:03.009210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-03T23:36:39.368933Z digest=sha256:f207000f3dd5284cce493a495f43504bd872d72d75f6d8ec208abb0403fd7e05

Observation e61d5d58-1987-45ec-9f36-d1c8396bf123 · inbound

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers cites this paper.

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-15T10:41:34.347336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:41:34.347336Z digest=sha256:5651715777b5cdbf922020ebb5f37267486759928cacfddc173365d9f110069e

Observation 4baf9019-5e80-41d4-915c-809308adecbf · inbound

STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation cites this paper.

STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T11:59:50.510846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T07:26:28.527338Z digest=sha256:fc3f6438ce3524c0bec9fa3ccb619b46ea86bcd6be012100f3f1a8e1bd9fc893

Observation 8884b445-2c4e-4aff-a6df-11c1a21fede7 · inbound

AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation cites this paper.

AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:59:50.936746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T07:24:01.244733Z digest=sha256:717bc0b1a4bbd028ab507a2436e610337424db9ca6c2fa3584a6b4ff7428bc22

Observation 0be1038d-40fb-4b97-9e2e-ef62dc9d9c0e · inbound

OTCache: Optimal Transport for Geometry-Aware Caching in Diffusion Models cites this paper.

OTCache: Optimal Transport for Geometry-Aware Caching in Diffusion Models TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:45:39.526005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-01T06:21:22.905093Z digest=sha256:e2086c8e9bd07b5c07dcccf32654b4c8135b60e82e606a05850a390a7f9c5b5c

Observation 8b3a5318-e541-410b-ab9d-f7fa63760899 · inbound

SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation cites this paper.

SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-11T12:34:20.057072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T12:34:20.057072Z digest=sha256:ed3554a80963f0484bf43512f2f6a965a732d5a10d830c059000a865c4ee1365

Observation 36c66311-2b38-4296-a892-9e8c1a3e23f9 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 119

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.486119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:20774459bce0841b0b53d4da76dcbd9afae9804f87ad18587c623402e912dd30

Observation 8241ceab-0a78-49e3-bcd3-0e2e34d842ce · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 119

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:1c9231db9ae6b57a1865ccc2d07920f0e26108f3d413478b093779836d737525

Observation 5225392a-18c1-4d74-bb3a-67134f339fb3 · inbound

FdAudio: MeanFlow-Anchored Fr\'echet-Distance Post-Training for One-Step Text-to-Audio Generation cites this paper.

FdAudio: MeanFlow-Anchored Fr\'echet-Distance Post-Training for One-Step Text-to-Audio Generation TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T11:52:50.598080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:52:50.598080Z digest=sha256:1c56f5b3160bfd329208c9b8f62d72d2c0bb51e88f6fe107ed4e0ce1da61d59b