Pith. sign in

Paper Citation Record · LEDGER

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

As of 20 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 31 inbound Pith citation observations for arXiv:2412.21037.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.21037 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:21:34.742952Z

measured 95 of 95 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:30:28.222670Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.484714Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact2
  • verified fuzzy1
  • unresolved61
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dfbe94c7-e846-4d66-a21b-4e04b05ebcc5 · outbound

This paper cites write newline.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.457797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.457797Z digest=sha256:b4be711bceb15d562ecd85352dc2b6613e0348670c9adb25d1c07b9ae9383753

Observation 714224b0-ccce-4531-8ba9-b03a719a33a4 · outbound

This paper cites Building Normalizing Flows with Stochastic Interpolants.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Building Normalizing Flows with Stochastic Interpolants

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.463848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.463848Z digest=sha256:84ca2d49efc7a0bd6096760047b618a66d287dbc70be111fa9131f4704a77201

Observation df247e50-6c54-4caa-aa95-a138a2e0471f · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.469104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.469104Z digest=sha256:e703325cb1b4268334ae7ce9061ba069f572b358644d8a9bd21407e2e358d911

Observation 5d977bf6-f9ce-49e1-b0b9-353689b69edf · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.473968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.473968Z digest=sha256:b1a5c57050662bda1db258e014d5df560275da25773c2657d8c53860ee446302

Observation e84b5fc1-60f1-473f-9258-c9cc3ef23958 · outbound

This paper cites Qwen2-Audio Technical Report.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Qwen2-Audio Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.478441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.478441Z digest=sha256:fbd4cc455e283f526ff206eeef11eebce2023ec87370f8926120de8cf3b3d433

Observation 1f9b7d22-8a77-45a1-aa56-62b4f50e61d2 · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Scaling Instruction-Finetuned Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.484425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.484425Z digest=sha256:1ddb9f0b9a1805e2bf1b0d6d5f56e37c99d7c270736c8830a0b818b1487130ce

Observation 17d9e8d6-fb67-4da7-b3a6-bc6ace5d818f · outbound

This paper cites Simple and Controllable Music Generation.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Simple and Controllable Music Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.489417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.489417Z digest=sha256:f12e1135661420f94438032150f9a39568acbcfab1222846bf5caf96facd0dae

Observation 94c062f7-c4eb-4548-b789-3bb438cb5475 · outbound

This paper cites Look, listen, and learn more: Design choices for deep audio embeddings.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Look, listen, and learn more: Design choices for deep audio embeddings

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.494946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.494946Z digest=sha256:bd3a063618453d7b14f49aea13a240ded8e99bb651b4479a0e028035fbdfffda

Observation 08d2f6ec-be76-42ff-b0b7-0079eadb4f91 · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.499424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.499424Z digest=sha256:7a84f8c06ac7a636c54bb3f5503b1a8a36ec50a07ed9f895fabda010c32bf402

Observation cf778c2c-57b9-452a-b9f4-639f4e1eb360 · outbound

This paper cites Fast Timing-Conditioned Latent Audio Diffusion.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Fast Timing-Conditioned Latent Audio Diffusion

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.503571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.503571Z digest=sha256:89f6cd101b88b1563ad297d5fcc7fcfd6c2583f457ff8bb9113843936fc2ec01

Observation bd53f9b4-b7ed-4ea9-ac89-ca6ec49087b2 · outbound

This paper cites Long-form music generation with latent diffusion.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Long-form music generation with latent diffusion

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.508244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.508244Z digest=sha256:c3454f182fadd10cfcc17c3e9075ddc5613e42eb339ef4bec126e5f662f91ab6

Observation 841081a2-dbfe-4ff0-8364-87da6a98e644 · outbound

This paper cites Stable Audio Open.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Stable Audio Open

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.512529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.512529Z digest=sha256:034d63d6f575b9db9337b81eaa1b799d900528b8afd0633e316a4a374121b6a2

Observation 7fb785a6-bb01-44e5-b7b6-86b9edf7af67 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Scaling Laws for Reward Model Overoptimization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.516577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.516577Z digest=sha256:da8bab6b7c5da2b7b5da32371a2245efa32f7ea3ecde81ce54f609fc87279c79

Observation fee7105a-96f9-46eb-bcba-338e0d6ac279 · outbound

This paper cites Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.520734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.520734Z digest=sha256:9bcc3901955ce9506e627fb494eaafcbff84fd26273737247def0d536cfcff4c

Observation c3475995-b8cf-4317-a40d-374602111266 · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Reinforced Self-Training (ReST) for Language Modeling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.525539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.525539Z digest=sha256:9c02f2c899ce29010bbc66184e3f60ea4d49c9468cd95d94968f7e4d1f869a5b

Observation 269f0661-ebb8-42a7-a46b-bb353b8d3ac2 · outbound

This paper cites Efficient Diffusion Training via Min-SNR Weighting Strategy.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Efficient Diffusion Training via Min-SNR Weighting Strategy

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.529354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.529354Z digest=sha256:b58b29832569dbfb3180a2bd8d9fcc4f40ab94ce6ddaa76c820b727dd3840cf3

Observation ef2c5bb6-dc29-4bed-b36b-22463c2671dc · outbound

This paper cites Classifier-Free Diffusion Guidance.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Classifier-Free Diffusion Guidance

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.533363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.533363Z digest=sha256:df067e690e6f6e4c451e00c068cabbfe9cfa331ef00dcd11b101ab57772a1073

Observation 5a84fea2-0bc2-4a98-ad1c-554f2e3bc54b · outbound

This paper cites Denoising Diffusion Probabilistic Models.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Denoising Diffusion Probabilistic Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.537595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.537595Z digest=sha256:49efbba07d0f9376839977311b50310d37367a509c41aff314c7c4ee4f9f0070

Observation 1b5ff383-615e-4c2d-b58c-a1f4a87c0976 · outbound

This paper cites Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.542147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.542147Z digest=sha256:36be6a8a83720e2803fdcd02d4b2f7f295b6ae87d30e122ad567aec2cb14e133

Observation 5f2e0827-c153-4474-98f3-ecf7546a84f9 · outbound

This paper cites Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.546697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.546697Z digest=sha256:806df3e9fec003c062dfd13fabf732e4255bae65c2ef70d80db7099bf0c3f43a

Observation d70bc48e-facc-4e64-b3f0-07b86554d65b · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.551154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.551154Z digest=sha256:d30d3485767f9b03e20dc141f72c006007746ce006f4e688d3a518fa5db083e8

Observation 3406b416-db76-49db-bef0-613d1fef4426 · outbound

This paper cites Elucidating the Design Space of Diffusion-Based Generative Models.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Elucidating the Design Space of Diffusion-Based Generative Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.555413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.555413Z digest=sha256:3515189e9ed18c26c67f423ce827335e9face228f0a2c462b224f4c75ef38aa5

Observation b74a1729-c083-437a-9882-6e724ec345fc · outbound

This paper cites A udio C aps: Generating captions for audios in the wild.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization A udio C aps: Generating captions for audios in the wild

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.559672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.559672Z digest=sha256:682c28e7970573509f570f971faa689792ad5332e4068b4eedfa61d21bc5949f

Observation 7f4c4502-1538-4d70-9098-679f50dcd79b · outbound

This paper cites sDPO: Don't Use Your Data All at Once.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization sDPO: Don't Use Your Data All at Once

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.563825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.563825Z digest=sha256:3cad15d0710736dcdb494b7bdf11e5cea97c9b9bb83c75b5c0f6dbc5ccfeb6e4

Observation 64cb2faa-aacb-4b6f-950b-8942914e7bf4 · outbound

This paper cites Adaptive non-uniform timestep sampling for diffusion model training, 2024 b.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Adaptive non-uniform timestep sampling for diffusion model training, 2024 b

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.568233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.568233Z digest=sha256:9a8ce154bde617d8367cd5da03f664c81af994ea2a1073fd8a6a42633af6c8f8

Observation 7fff4a72-c2a1-43a8-a19b-7126519d6feb · outbound

This paper cites Auto-Encoding Variational Bayes.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Auto-Encoding Variational Bayes

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.574308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.574308Z digest=sha256:9069a6b4a8d80ee96b7a6a6e2ebc09fc521a0c588c522fc710c1a77eb8abac7c

Observation f156cb9d-4ed3-4cd2-85d6-228143149360 · outbound

This paper cites Improving Text-To-Audio Models with Synthetic Captions.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Improving Text-To-Audio Models with Synthetic Captions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.578346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.578346Z digest=sha256:f0b5ebc0b9267744d7a9699b23b429b05591401e23391afc40ded33e95aeb9e2

Observation 701f1071-6946-4ab3-98b6-02b5ecff497b · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.582426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.582426Z digest=sha256:d3b91f6ab72baace0d4e0064bf7ce59b3a06c99e87b1ff5a058d1a933ca36195

Observation f15dbce5-7bd9-4a40-bc3d-77dd6e8c322d · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization RewardBench: Evaluating Reward Models for Language Modeling

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.586528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.586528Z digest=sha256:d92305581d7039d315587dd85368b53decddb7889e8499c5fd40fffe0e7d3481

Observation b360166c-1f97-4adf-93ca-1ce8e1519132 · outbound

This paper cites Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.590797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.590797Z digest=sha256:b62018cabf21994ec8a62071d4c859b5e62ed5ddf5eafd43cccef4ab2eb2dd0a

Observation 2ef549bb-797d-4afa-b2bf-e04e44bff316 · outbound

This paper cites BATON: Aligning Text-to-Audio Model with Human Preference Feedback.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization BATON: Aligning Text-to-Audio Model with Human Preference Feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.596114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.596114Z digest=sha256:f74396a8eef2491200d4163362aac35b8dfcab7ed17a1f6504b14c08861d1774

Observation 42ce1007-43b2-4804-8fd2-be2f7163b7fb · outbound

This paper cites Flow Matching for Generative Modeling.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Flow Matching for Generative Modeling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.601091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.601091Z digest=sha256:aa7c8f8fc5f7cf5f6b2edf9f91be17f2aaece4fc10f4a914eceace6bdb366223

Observation 84334678-d58e-4b68-a0fa-7ada18b2f96e · outbound

This paper cites Generative Pre-training for Speech with Flow Matching.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Generative Pre-training for Speech with Flow Matching

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.605049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.605049Z digest=sha256:116acd00667f6a0d9a8191c7623d48c76357adddec236459343db37807545cc1

Observation 767aad89-3212-4799-8cce-c1d3d53b323c · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.609758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.609758Z digest=sha256:baabe66940578abb821303662103099c858ee520c899b436fe343ef84aeddaba

Observation 4d751b30-f5da-43c1-a422-3fc34a613848 · outbound

This paper cites AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.613652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.613652Z digest=sha256:d0a7446385a532398ba162dfedd3802fb557ca414928e31e1f0cce7fbb847442

Observation c71edb85-6817-4f72-9c8a-436701d8accc · outbound

This paper cites FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:21:35.061744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T23:21:34.618608Z digest=sha256:1da621d1a958dedda274bea8ed7c08d7cb1ac6aede0cdf52260cfafd0a233c40

Observation fe0d98a4-e94f-4339-8982-dbed3a07f311 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.623344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.623344Z digest=sha256:afd46bb8e860f7bbab69d79bf8c5df05b63907b6e012a3632b50792e64684ce5

Observation 1f20d4be-a857-4da3-96b6-70a59344e744 · outbound

This paper cites Decoupled Weight Decay Regularization.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Decoupled Weight Decay Regularization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.627496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.627496Z digest=sha256:63020d8d78dca6f719b178574bfeb7130f5ffb9c7a7774cb7de1f440cac0baa4

Observation ef0aff9c-f639-460f-a2aa-732ccc90f8ec · outbound

This paper cites Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.631417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.631417Z digest=sha256:9d330bd3cf162d9964f6d36f87fba6ff20f7b1452b6d8050eb9f35bfce5de77a

Observation cfa23ff7-b5ec-4705-9689-06d6f6473482 · outbound

This paper cites Plumbley, Yuexian Zou, and Wenwu Wang.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Plumbley, Yuexian Zou, and Wenwu Wang

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:21:35.586953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T23:21:34.635465Z digest=sha256:88427a41c01260883b16c31d5860fe078df381698fcb64ef9cf6c7f2fef11576

Observation 4490b601-2476-42cd-88e9-4543666b1aeb · outbound

This paper cites OT-Flow: Fast and Accurate Continuous Normalizing Flows via Optimal Transport.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization OT-Flow: Fast and Accurate Continuous Normalizing Flows via Optimal Transport

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:21:34.999252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T23:21:34.639882Z digest=sha256:07663b15da1727b1e777cf9608426f71300d3543dac8c7644203625bf7ba29ce

Observation 551e976d-da4d-4e16-bb2a-a38aec00af94 · outbound

This paper cites Training language models to follow instructions with human feedback.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Training language models to follow instructions with human feedback

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.644929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.644929Z digest=sha256:a7904c4898bb1b9bf6c789c4accf2f1e4c5f6714c42757d7631187737a73db92

Observation bcd7b2ad-fd7d-478e-9c4c-b14d9e67b33a · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.648810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.648810Z digest=sha256:30ce372003a964bc93e2a452513539ad575f54b3ace5a6962723d8a94a105cdf

Observation 04befc61-98dc-4740-9bff-2cc8530727dc · outbound

This paper cites Iterative Reasoning Preference Optimization.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Iterative Reasoning Preference Optimization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.652625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.652625Z digest=sha256:9cf9748c67bea2004630cd6066a5c00ecd2da6572e6f6d559d2966fecc1b071e

Observation 62450e1f-c527-4a7d-b71c-2069d03824ee · outbound

This paper cites Scalable Diffusion Models with Transformers.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Scalable Diffusion Models with Transformers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.657081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.657081Z digest=sha256:2f5a23e3944740dd1208e5da5c9104531b98085f0ea4b4da37879ea47089be9f

Observation b9ef3da2-eead-474a-8b86-200a858c4dc4 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.662433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.662433Z digest=sha256:38627cd23d28f8d38b632c70c4ebb7baa3212aefc306a34dcafbe1b67dffdf3a

Observation 59acef81-8414-467b-9b42-da40ddef3b22 · outbound

This paper cites From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.666487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.666487Z digest=sha256:4f836e5ae4e5663ebe78533e4f164fd99d257faffd988ea308ade5db323ffe29

Observation f9127c2c-1a6b-4df6-b7a3-3b9c7e20e8fb · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.670941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.670941Z digest=sha256:6a797470854fd46f7f6ddcb02accc63b79d667d8b94001c8a824973b615ee2a7

Observation 8b0373eb-58c6-4b89-9b4e-3d54e3ea2033 · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.675200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.675200Z digest=sha256:8d85431c4681e760fe0c36261c2a769782287b8cb04e3ae364af253d82353341

Observation 7ab6fe7a-80c9-4823-9087-910b88eb3cbe · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization High-Resolution Image Synthesis with Latent Diffusion Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.679776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.679776Z digest=sha256:30efa873f8c9da4e7148e16e4f2d8e6bb89cbfcf187183cef6f151da2c2b840f

Observation bd0f18e1-e25d-4a21-9452-e4890320d749 · outbound

This paper cites Improved Techniques for Training GANs.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Improved Techniques for Training GANs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.684082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.684082Z digest=sha256:0b12e7f819c689d012db9ac44c8155eaf4cdc7f5e2dc7f8375045bbbd13b0996

Observation 00f413f2-7dbe-4219-9356-d4a9674bf326 · outbound

This paper cites Denoising Diffusion Implicit Models.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Denoising Diffusion Implicit Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.688478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.688478Z digest=sha256:3810810be7e54f03650d366846c40d8e15e20fd6097aee781e60a61b1855722c

Observation 45071ce0-1e92-4705-8903-b192bfc6f152 · outbound

This paper cites Generative Modeling by Estimating Gradients of the Data Distribution.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Generative Modeling by Estimating Gradients of the Data Distribution

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.692979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.692979Z digest=sha256:e1f6002dba4f1bed8d3ff284c45a257b86d74093a807425032409412634e0f73

Observation 0043f3bd-eebd-4220-885b-bac5c6b7bd1f · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.697783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.697783Z digest=sha256:626cfe9ecce09c1286d36df1a7b2c9e9a90ef58575287beb726732c805048e47

Observation 56e9bb92-efc9-44a4-bbb8-2d5ff679c976 · outbound

This paper cites Attention Is All You Need.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Attention Is All You Need

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.702355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.702355Z digest=sha256:d77a6c82b96ad0c04bbad28bcc3be520e72ea65a2900e4364b6cd2c19c7a0632

Observation 9ae97aa0-ae99-4f7c-87d9-bcd2b1685718 · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.706605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.706605Z digest=sha256:507e63598c7f3f9ccee0221f9e5c337c1ab4af96a2afadf937caf1f8cf06106d

Observation fb1276ca-32c7-4723-80ce-0182569c3495 · outbound

This paper cites Diffusion Model Alignment Using Direct Preference Optimization.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Diffusion Model Alignment Using Direct Preference Optimization

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.710431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.710431Z digest=sha256:279e3c7b944de1dc0af2c57f4afa7724b1a4d22beda5bb9d82aae6be3e38de41

Observation 0965e152-74d5-4b7c-bd7d-67728bb17a1e · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.714901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.714901Z digest=sha256:997c9c459649b84a3bd0e0bfcb496cec48d4760d29128d6a8e29d177c534e2aa

Observation f5155aad-2141-441c-b2f4-1c807fdde1c4 · outbound

This paper cites Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.719059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.719059Z digest=sha256:1fd6f6d33f1a8b578208870e3641272c833fef269c1e5f86e13474a7d49557fa

Observation 0b1e123c-b37e-48b3-ba86-882ec43b24b5 · outbound

This paper cites Self-Rewarding Language Models.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Self-Rewarding Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.723994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.723994Z digest=sha256:f3f6f010b21c30e1c71f69ac3024be529578d60feac65d43fd2dca7f87dd0060

Observation 098672e5-dc02-45e7-be75-259d3cc2dafa · outbound

This paper cites STaR: Bootstrapping Reasoning With Reasoning.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization STaR: Bootstrapping Reasoning With Reasoning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.728757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.728757Z digest=sha256:3855e6c540410ccb6bd0b6190fce8e6646bcea57483e56c39f256ab05d3a50e1

Observation 9d46ae32-bd57-4a67-a35a-74489ec233b6 · outbound

This paper cites @esa (Ref.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization @esa (Ref

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.733438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.733438Z digest=sha256:1b3cfa9c72accae50c148814ba79147fb9d732ed5a2b37f3419942fc94096bb6

Observation de32df46-d16a-4d27-9505-3d1ed74b925e · outbound

This paper cites an unresolved cited work.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.737887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.737887Z digest=sha256:ee570ecfe1d4327e91236e3f769e266355188a56b654543984f36035b4d2a72d

Observation 679cdb67-48ae-4f5e-ad85-54c4a6ccfd9f · outbound

This paper cites an unresolved cited work.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.742952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.742952Z digest=sha256:cce94065c176209f1a6c3e6d6628ecc7c30aa80735164254954e5e1488b4c992

Pith citing papers

Observation 6a39e9de-aed9-4235-83d5-8ba3fefcf736 · inbound

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder cites this paper.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T22:16:17.260074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:16:17.260074Z digest=sha256:1f3b46bd4f597eda40a7381e16422e9063405165e8f4be2c00ccb66c5ade2f06

Observation b1f054af-4f43-4264-b9f8-10dfb0aa539c · inbound

Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping cites this paper.

Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:41:36.642247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T13:38:17.640841Z digest=sha256:494e2ee9abadaa1cf50d9c047bd1daf2709cce72a008cb9aaf1e6253a11721e2

Observation ae659b54-8ae6-4d0f-8e35-7eb036357a6c · inbound

AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation cites this paper.

AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:09.446928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:09.446928Z digest=sha256:5c847b6baa47631895534f01ca28d5b8a1231284197fb792b3c92622bccc6573

Observation d8a1eb31-d758-489d-9bb4-91319a8dc326 · inbound

SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations cites this paper.

SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:04:00.556655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:04:00.556655Z digest=sha256:7369816f0aad7b5fd4e9aa50cbce38080d118f96348893d764cf32539a193fd2

Observation 89187b51-2060-42e8-8272-45b469a396f2 · inbound

JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment cites this paper.

JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T13:14:37.521235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:14:37.521235Z digest=sha256:d8f860abaee30c8591e1bf7889cc550ef3fc8a39d8964eefec876ebc76852fd2

Observation a9360261-9ba8-4900-a57c-82210bd60359 · inbound

RFM-Editing: Rectified Flow Matching for Text-guided Audio Editing cites this paper.

RFM-Editing: Rectified Flow Matching for Text-guided Audio Editing TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:36:37.020189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T16:34:17.431134Z digest=sha256:b9d6c9292cf38823d657374009bfac429da0d7a41b972af176cbfaa5faec459f

Observation 7f2626d6-579d-4430-8d7b-a17467c8e3bc · inbound

AudioMoG: Guiding Audio Generation with Mixture-of-Guidance cites this paper.

AudioMoG: Guiding Audio Generation with Mixture-of-Guidance TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:23.820034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-18T13:10:18.700497Z digest=sha256:0e0e8edd2cc74373226f8cfdec0a1cc898a6983697b2ac2d6cbec2b95d2062c6

Observation 8393a281-f852-4a31-a326-1d41ad32bbb6 · inbound

SemanticAudio: Audio Generation and Editing in Semantic Space cites this paper.

SemanticAudio: Audio Generation and Editing in Semantic Space TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T07:06:33.216494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:06:33.216494Z digest=sha256:0cbd028e746113311902ab72135f030332c28d8c30a60d4344f1018fb3fb732f

Observation 974ba821-f86c-43ec-aec3-c44a3ab2bc2a · inbound

Woosh: A Sound Effects Foundation Model cites this paper.

Woosh: A Sound Effects Foundation Model TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:53:15.849552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T20:51:08.144573Z digest=sha256:9ddff5844c0aae29ece7cb69291eb9d43e7afe490a35d8b6033162d77e2b2aef

Observation 0064991e-42bb-4efb-b39c-3d94b6b6da73 · inbound

ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling cites this paper.

ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.054519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T09:22:10.483263Z digest=sha256:bb0f59c09b6b1ff60b4445bcf8066e1d1ba2a619c7997473f5a2619dd0724258

Observation 2a3f17f9-1852-4239-ae25-f01a34248660 · inbound

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation cites this paper.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:46:44.885236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:21e77a59ac0b3530f54b2df22a40f36915ac3a17cbf71cb11be5e12e7b2c9cd6

Observation d47cc885-7e9c-4793-8dc5-b325ba92850b · inbound

AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling cites this paper.

AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.274264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T05:06:54.972846Z digest=sha256:9fc5af9f4ed47bbaf15534bfbe87342f50a5dce1d2b9e551d0d53e634be11a51

Observation 54e22aad-cf03-451f-a628-d13534dc5ff2 · inbound

AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling cites this paper.

AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:29:52.743864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T08:28:21.790110Z digest=sha256:997df6ec793979ea9c57183f62852d18a3c9b10e263c803a44c291d6de7b0b86

Observation 669b3f78-a6b5-4039-bc2f-28ca9bcf3bef · inbound

Dasheng AudioGen: A Unified Model for Generating Coherent Audio Scenes from Text cites this paper.

Dasheng AudioGen: A Unified Model for Generating Coherent Audio Scenes from Text TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T11:03:18.758307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T10:56:00.765400Z digest=sha256:fff18a1d9ea1d06130cdfe9cf3c4631d95db188a76a6a5b5783475c269b4d58e

Observation d4ff063d-a656-4c2b-81a0-0d5a6086c142 · inbound

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment cites this paper.

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:12.629539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T21:12:19.893944Z digest=sha256:10959045afcc985b7554191308f9a7cd07c8e48e2f16cbb60eddc03600497c20

Observation eed75720-14f1-41b5-89e4-64dd56b13bc4 · inbound

UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion cites this paper.

UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-06-28T20:52:37.895615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T20:44:20.190064Z digest=sha256:a714270557f0bfa7cbd1e3eb128ce30adfc9bb5073231e1a69f28c2d15231dc9

Observation 7b58e407-6dba-4f8f-8277-b69eda7cb2e4 · inbound

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation cites this paper.

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:28:18.809247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T08:04:48.283908Z digest=sha256:587cc0d99bb67375c2bb5628199b5df83c9bc04f4ad17222ff061419acd237f6

Observation 90e5fdab-695d-4dd1-806f-14f8978ac129 · inbound

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers cites this paper.

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:39:39.325372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T15:51:48.672086Z digest=sha256:ee6ab01c2e55396106aabfb51bfc6fecd65eab4942dd4ce208f505850684926a

Observation 8592f3b5-52e4-43f6-ba32-29cdba770adc · inbound

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers cites this paper.

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:39:03.009210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T23:36:39.368933Z digest=sha256:9fdf66c1cce6f23dfe18ae350bb2c632e82b523e626c5c6f4fe22980cba867cc

Observation e61d5d58-1987-45ec-9f36-d1c8396bf123 · inbound

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers cites this paper.

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-15T10:41:34.347336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:41:34.347336Z digest=sha256:44b843f43830e8517521cf624deba319ec67737c9dfc2c0656cfee8d7134ba9a

Observation 4baf9019-5e80-41d4-915c-809308adecbf · inbound

STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation cites this paper.

STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T11:59:50.510846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T07:26:28.527338Z digest=sha256:f7acb01ae2c169f5fc0e729dc676ae43ff9b0e8a390622666cf15af0309d03b8

Observation 8884b445-2c4e-4aff-a6df-11c1a21fede7 · inbound

AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation cites this paper.

AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:59:50.936746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T07:24:01.244733Z digest=sha256:9b8ce289d2e9037ca34c61ab93fc236f117c2e4f694fb1ffa7c2a0aa3593633b

Observation 0be1038d-40fb-4b97-9e2e-ef62dc9d9c0e · inbound

OTCache: Optimal Transport for Geometry-Aware Caching in Diffusion Models cites this paper.

OTCache: Optimal Transport for Geometry-Aware Caching in Diffusion Models TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:45:39.526005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T06:21:22.905093Z digest=sha256:97429e9ee66fee179c6b62580f8b6732bea3ec3cce84d5381abb1bea9e72584b

Observation 8b3a5318-e541-410b-ab9d-f7fa63760899 · inbound

SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation cites this paper.

SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-11T12:34:20.057072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T12:34:20.057072Z digest=sha256:b0871c1a723cb625a58b85ee54158e175453484a53294071c48634adae3caa40

Observation 36c66311-2b38-4296-a892-9e8c1a3e23f9 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 119

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.486119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:6a44536b1f93d81ca19fd7295ff0db24227376b4f1858e9d7fb1817dc2d3d81e

Observation 8241ceab-0a78-49e3-bcd3-0e2e34d842ce · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 119

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:f731ae23dfc3e4b4e2d0daa28f409e2d5fc24ab6d7f63d1f24fa559d283400d7

Observation 5225392a-18c1-4d74-bb3a-67134f339fb3 · inbound

FdAudio: MeanFlow-Anchored Fr\'echet-Distance Post-Training for One-Step Text-to-Audio Generation cites this paper.

FdAudio: MeanFlow-Anchored Fr\'echet-Distance Post-Training for One-Step Text-to-Audio Generation TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T11:52:50.598080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:52:50.598080Z digest=sha256:a16b18333f52992c6b9aa9f49fad31483dbd5d15b59c9f3e6f8258df32cd33f2

Observation bd374cdc-3ac6-41b9-acb6-42e80f99e63d · inbound

AudioScape-TTA: A Structured Soundscape Benchmark for Fine-Grained Text-to-Audio Evaluation cites this paper.

AudioScape-TTA: A Structured Soundscape Benchmark for Fine-Grained Text-to-Audio Evaluation TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T14:42:59.761744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:42:59.761744Z digest=sha256:10eb6645cf676feae15987678061023e601b9f30eb5b9b15f7043c6e0d6fe084

Observation bd754d77-51d1-4097-8c07-6c7d782b2fed · inbound

RAG-Audio: Retrieval-Augmented Generation for Faithful Brain-to-Audio Reconstruction cites this paper.

RAG-Audio: Retrieval-Augmented Generation for Faithful Brain-to-Audio Reconstruction TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T14:31:33.072606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:31:33.072606Z digest=sha256:160b28d2a660621b486caf869c34cdca181b7c7c8b23350c8e9bf5aec1952bd6

Observation 57b22c66-ff37-4d62-80bf-81d474c6a82c · inbound

MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching cites this paper.

MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T00:30:28.222670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:30:28.222670Z digest=sha256:c6df33adc436f215ed4163736283478e2554a6e57f7de0a6977abd99ed37982c

Observation 4105aa1f-213f-48c2-8a20-1390f1db9ce3 · inbound

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching cites this paper.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:28.042751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:28.042751Z digest=sha256:4963d2ed1584f440733e5988b40edd587ea14b5f427fb8e13706e4e215fa1d2b