Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T20:23:36.774359Z
Paper Citation Record · LEDGER
As of 22 July 2026, this Paper Citation Record lists 50 of 50 outbound references and 1 inbound Pith citation observation for arXiv:2604.04348.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T20:23:36.774359Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-07-21T06:31:05.380196+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T23:39:22.070629Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-06-30T23:45:08.232528Z
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d5871b25-db4c-4d8f-bd79-cdf3399182e8 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text LRS3-TED: a large-scale dataset for visual speech recognition
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 96e47384-3e35-472c-abac-a350b4d9b12b · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation b050bbc0-4049-4f9b-b088-79ca37476313 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Com- mon voice: A massively-multilingual speech corpus
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 8aa19022-5f78-464a-a6bd-4c127539ca8a · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Vggsound: A large-scale audio-visual dataset
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 6eca2bf2-3824-4c21-af5b-47977ae1bb15 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Video-guided foley sound generation with multimodal con- trols
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation d283032e-c15b-4255-b48c-e83ca5528e47 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Mmaudio: Taming multimodal joint training for high-quality video-to- audio synthesis
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 2cd59036-3345-4036-9612-ada8d7506fef · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Scaling rec- tified flow transformers for high-resolution image synthesis
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 1cfe1090-5b0c-4c39-9d47-c81fadba63e7 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Text-to-audio generation using instruc- tion guided latent diffusion model
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation e4793e25-34b8-4a4d-aca4-f520c8c9e2a9 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Classifier-Free Diffusion Guidance
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 7088374f-c118-4150-a1cb-da42a3d21160 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 943b4866-f1cf-4a92-9882-25f341233f4b · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Imagen Video: High Definition Video Generation with Diffusion Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 9c0a7983-c51b-44f6-a298-474dac8d5b04 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Video dif- fusion models.Advances in neural information processing systems, 35:8633–8646
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation d41a35c0-fe4c-4622-9a37-6521f67dab12 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Make-an-audio: Text-to-audio gen- eration with prompt-enhanced diffusion models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 5b40324d-9df9-4461-a02c-04ef2c347ade · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Taming visually guided sound generation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 42c24fd8-13db-4852-8e03-fbc0d474231c · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Synchformer: Efficient synchronization from sparse cues
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 900a516b-0d04-4db6-aa16-b1618722252f · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text V oicedit: Dual-condition diffusion transformer for environment-aware speech synthesis
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 4d5e1765-f75b-413f-ba56-231900503142 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 8a79be29-c6d0-47ec-8fb7-d2ec2d161185 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Guided- tts: A diffusion model for text-to-speech via classifier guid- ance
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation b5c1bd87-780d-4943-92b7-260f673eccf8 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Kingma and Max Welling
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 71d2c3fd-9785-49e7-9466-e673f878638a · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Hifi-gan: Generative adversarial networks for efficient and high fi- delity speech synthesis.Advances in neural information pro- cessing systems, 33:17022–17033
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 1a8ed80d-aa17-4abe-abab-26924a3a0a05 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Audiogen: Textually guided audio gen- eration
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 8eb1b796-b397-44d3-99f6-7e577d08ff61 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Vintage: Joint video and text conditioning for holistic audio generation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 49933331-aeec-4329-8fe2-24fabd908b14 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 0401f434-9b66-4ffe-b986-a327a8d41182 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text V oiceldm: Text-to-speech with environmental con- text
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 2bc6ba92-01c8-47d1-8afd-b53d37b3e81b · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Tri-ergon: Fine-grained video-to-audio generation with multi-modal conditions and lufs control
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation b6349fbb-f2bc-4e58-baa9-000c3196dfe9 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 6bebdc63-228a-448e-8733-65480917d072 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Au- dioldm: Text-to-audio generation with latent diffusion mod- els
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 4c3e5207-7ce4-4266-9df5-bc8887d6c3f2 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Audioldm 2: Learning holistic audio gen- eration with self-supervised pretraining.IEEE/ACM Trans- actions on Audio, Speech, and Language Processing, 32: 2871–2883
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 85394fce-e6b3-4583-8371-1cd84c11e6bd · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Flow straight and fast: Learning to generate and transfer data with rectified flow
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 17f8027b-7823-4084-8e0e-4cbfee6274ed · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Decoupled weight de- cay regularization
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation b388d655-bc4e-4863-a8fe-90016b242897 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Diff-foley: Synchronized video-to-audio synthesis with la- tent diffusion models.Advances in Neural Information Pro- cessing Systems, 36:48855–48876
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 5cf8fa71-ef8f-40a9-ab8a-45d733574964 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 67070e8e-2bf4-4f95-aa9e-d6c6a8ecda85 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Pytorch: An im- perative style, high-performance deep learning library.Ad- vances in neural information processing systems, 32
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation d9a3efc5-8c9e-4cd0-88fc-331eda484d60 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Scalable diffusion models with transformers
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation db27bd8a-e00f-4aae-b0d5-e6578a1b8a81 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Learning transferable visual models from natural language supervision
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation b8a5b66b-b1ce-4a39-95ee-c8ddca8467af · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Robust speech recognition via large-scale weak supervision
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation e8695b06-01ef-4ff9-89ae-83c5f70cc6d6 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 0b3f25cf-3ae0-4bfb-8a04-bf26a5b63826 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text High-resolution image synthesis with latent diffusion models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 11a95843-eead-4cc0-a314-db33cee3c359 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation e3f2adc0-e94b-406a-b0aa-db1d6ff79c72 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text I hear your true colors: Im- age guided audio generation
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation dbe4e294-7924-4f73-b6e8-a00a45ce0ea5 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Make-a-video: Text-to-video generation without text-video data
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation a4d74983-a2cb-4444-aa8c-d3c32bececf2 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Score-based generative modeling through stochastic differential equa- tions
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation c946163e-c7ca-45fa-8784-d5988bcf4ad9 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 2f52d7cf-c424-4d50-b040-5e629ace856a · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Naturalspeech: End-to-end text-to-speech synthesis with human-level quality.IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(6):4234–4245
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation ab2a469d-4513-4a1d-b1ce-c12323d55708 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foun- dation models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 175a2779-c67e-4365-b733-3e4900bcdfd9 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Frieren: Efficient video-to-audio generation network with rectified flow matching.Advances in neural information pro- cessing systems, 37:128118–128138
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 9ff5c10b-551f-4cf5-9ba3-4911237e2de8 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Wav2clip: Learning robust audio repre- sentations from clip
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 5b64560b-b850-4f77-a50f-801584deec55 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Son- icvisionlm: Playing sound with vision language models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 652fc853-fccd-41c3-a7cc-f63e44889002 · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation b6f7e65f-b241-4309-b285-fd1a76bfcdfd · outbound
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text condition–unconditional
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation f703365e-a69d-494b-bbf3-5a12f52a8deb · inbound
Do Joint Audio-Video Generation Models Understand Physics? OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.