Pith. sign in

Paper Citation Record · LEDGER

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text

As of 22 July 2026, this Paper Citation Record lists 50 of 50 outbound references and 1 inbound Pith citation observation for arXiv:2604.04348.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.04348 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T20:23:36.774359Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-21T06:31:05.380196+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T23:39:22.070629Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-06-30T23:45:08.232528Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact5
  • verified fuzzy44
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d5871b25-db4c-4d8f-bd79-cdf3399182e8 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text LRS3-TED: a large-scale dataset for visual speech recognition

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:47.511047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:94399df78b82da4741b7760871750e9d00187e0ca2fa8dec80f4b297638f08bd

Observation 96e47384-3e35-472c-abac-a350b4d9b12b · outbound

This paper cites Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.369239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:5ac8c51ad465ec9dcbabbb69c1018ced39a31508ba1048e53a7b35cbf731e00f

Observation b050bbc0-4049-4f9b-b088-79ca37476313 · outbound

This paper cites Com- mon voice: A massively-multilingual speech corpus.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Com- mon voice: A massively-multilingual speech corpus

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.466475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:5fad60c1edb15ebdc29425e395ad45d1dc0ac468f1a0331dfcca81caa4642e8a

Observation 8aa19022-5f78-464a-a6bd-4c127539ca8a · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Vggsound: A large-scale audio-visual dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.453117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:592bf6cf94336caab7337019860e806eb46087e6064f53b8bbcf11a78e985fe0

Observation 6eca2bf2-3824-4c21-af5b-47977ae1bb15 · outbound

This paper cites Video-guided foley sound generation with multimodal con- trols.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Video-guided foley sound generation with multimodal con- trols

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.457265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:dc3c056b5c603bc8eb94425c9cabefdb161d728858e5056dd2727d08471a7cc3

Observation d283032e-c15b-4255-b48c-e83ca5528e47 · outbound

This paper cites Mmaudio: Taming multimodal joint training for high-quality video-to- audio synthesis.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Mmaudio: Taming multimodal joint training for high-quality video-to- audio synthesis

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.461838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:b315105e6ab2e2ec0c38aca7f053b94088a423230e16329ad064e6e7c76fe24c

Observation 2cd59036-3345-4036-9612-ada8d7506fef · outbound

This paper cites Scaling rec- tified flow transformers for high-resolution image synthesis.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Scaling rec- tified flow transformers for high-resolution image synthesis

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.470968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:b76721f1b91b80e083980069b41b3224bf48d94f6324cacba7fc076203f44d1e

Observation 1cfe1090-5b0c-4c39-9d47-c81fadba63e7 · outbound

This paper cites Text-to-audio generation using instruc- tion guided latent diffusion model.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Text-to-audio generation using instruc- tion guided latent diffusion model

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.475168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:530b6de6ed3080636a0613e058e7a97720ac299b7d9be8c5d5d77051e778b074

Observation e4793e25-34b8-4a4d-aca4-f520c8c9e2a9 · outbound

This paper cites Classifier-Free Diffusion Guidance.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Classifier-Free Diffusion Guidance

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:00:47.536659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:836be7da2cde8bbcfd7d9aae6998b526372894b9fa19c82aef46ebe5f7717b05

Observation 7088374f-c118-4150-a1cb-da42a3d21160 · outbound

This paper cites Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.426814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:52f09c5cead78484c3a6ddc731f559c2b40d033444ed75d5e7ca6c04cbaf3835

Observation 943b4866-f1cf-4a92-9882-25f341233f4b · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Imagen Video: High Definition Video Generation with Diffusion Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:31:08.340948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:63c121563aa363ce7de5a89c095263971f81bc2bb10ec20997fb4ed4f283ea12

Observation 9c0a7983-c51b-44f6-a298-474dac8d5b04 · outbound

This paper cites Video dif- fusion models.Advances in neural information processing systems, 35:8633–8646.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Video dif- fusion models.Advances in neural information processing systems, 35:8633–8646

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.431154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:2efa65b294356401e1adb62f9cdd64b6dd3f93054660e6307a1ff39acf811d34

Observation d41a35c0-fe4c-4622-9a37-6521f67dab12 · outbound

This paper cites Make-an-audio: Text-to-audio gen- eration with prompt-enhanced diffusion models.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Make-an-audio: Text-to-audio gen- eration with prompt-enhanced diffusion models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.417646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:538980b443c32c4b90ba5a4bc0584e85513c4476f82b5213d99dc3b041b62e81

Observation 5b40324d-9df9-4461-a02c-04ef2c347ade · outbound

This paper cites Taming visually guided sound generation.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Taming visually guided sound generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.413435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:5653801e75b0c513ac969bbaca3b9fbeefb11a47d51734e9a5f9c8624d529c75

Observation 42c24fd8-13db-4852-8e03-fbc0d474231c · outbound

This paper cites Synchformer: Efficient synchronization from sparse cues.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Synchformer: Efficient synchronization from sparse cues

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.421687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:45bc4112f2051de254731e491a48ae93daf6984e91303eee23d196670b9a7a6a

Observation 900a516b-0d04-4db6-aa16-b1618722252f · outbound

This paper cites V oicedit: Dual-condition diffusion transformer for environment-aware speech synthesis.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text V oicedit: Dual-condition diffusion transformer for environment-aware speech synthesis

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.435228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:ddfa8c3039d708d1102dab846a711cb8bb80e9a53db0faabac560296c52884e4

Observation 4d5e1765-f75b-413f-ba56-231900503142 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:47.502823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:3f11f63307c54abc1656cded06193fb065904bc556427e2d0f8124dcb1596315

Observation 8a79be29-c6d0-47ec-8fb7-d2ec2d161185 · outbound

This paper cites Guided- tts: A diffusion model for text-to-speech via classifier guid- ance.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Guided- tts: A diffusion model for text-to-speech via classifier guid- ance

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.439435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:970cda59352e0d16987d73adfe427197d3948ccff68cb30dbd7bdfc5f3092a1b

Observation b5c1bd87-780d-4943-92b7-260f673eccf8 · outbound

This paper cites Kingma and Max Welling.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Kingma and Max Welling

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.409325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:fb89f38f27b6fe0e0620ad1ceb38076e6f09fa1ebcbc91d9577896c9a13e3900

Observation 71d2c3fd-9785-49e7-9466-e673f878638a · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fi- delity speech synthesis.Advances in neural information pro- cessing systems, 33:17022–17033.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Hifi-gan: Generative adversarial networks for efficient and high fi- delity speech synthesis.Advances in neural information pro- cessing systems, 33:17022–17033

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.444272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:b40239522104f026aa2270799451bdd7e0528549a6e0d758794d9035d8fd0ce4

Observation 1a8ed80d-aa17-4abe-abab-26924a3a0a05 · outbound

This paper cites Audiogen: Textually guided audio gen- eration.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Audiogen: Textually guided audio gen- eration

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.448850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:f696b75f5f54c7d96b7300eafa7f841e6497693c5f60433a94e0917d1b5da574

Observation 8eb1b796-b397-44d3-99f6-7e577d08ff61 · outbound

This paper cites Vintage: Joint video and text conditioning for holistic audio generation.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Vintage: Joint video and text conditioning for holistic audio generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.479262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:fff42f408a708695498c6bb9e4c268997167d29eaa114de59c342d3bb290ca27

Observation 49933331-aeec-4329-8fe2-24fabd908b14 · outbound

This paper cites an unresolved cited work.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Unresolved cited work

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.396482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:f1d9ed02d908debeab083b23d1f8aff52643051ad758af954c05ea0b73dc62a0

Observation 0401f434-9b66-4ffe-b986-a327a8d41182 · outbound

This paper cites V oiceldm: Text-to-speech with environmental con- text.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text V oiceldm: Text-to-speech with environmental con- text

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.387377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:0876beb889d58580aa0bcdd8d9cdce77794b07145e9b1c4c29e81ebcb454aea6

Observation 2bc6ba92-01c8-47d1-8afd-b53d37b3e81b · outbound

This paper cites Tri-ergon: Fine-grained video-to-audio generation with multi-modal conditions and lufs control.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Tri-ergon: Fine-grained video-to-audio generation with multi-modal conditions and lufs control

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.400734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:63e79b54552ab0797250fc4a52c107a798ad02e90026ffff15932eec637c8b36

Observation b6349fbb-f2bc-4e58-baa9-000c3196dfe9 · outbound

This paper cites an unresolved cited work.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Unresolved cited work

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.340924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:641179354923cd76caa69f409bde880299b8c0a3a10f43f4fe37482de4b78978

Observation 6bebdc63-228a-448e-8733-65480917d072 · outbound

This paper cites Au- dioldm: Text-to-audio generation with latent diffusion mod- els.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Au- dioldm: Text-to-audio generation with latent diffusion mod- els

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.345125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:9609cbe997db69713c2525299164ce6b5aa7cbf17ff5b38c836bab29fca6aafa

Observation 4c3e5207-7ce4-4266-9df5-bc8887d6c3f2 · outbound

This paper cites Audioldm 2: Learning holistic audio gen- eration with self-supervised pretraining.IEEE/ACM Trans- actions on Audio, Speech, and Language Processing, 32: 2871–2883.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Audioldm 2: Learning holistic audio gen- eration with self-supervised pretraining.IEEE/ACM Trans- actions on Audio, Speech, and Language Processing, 32: 2871–2883

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.373949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:c97056cdd44c6cbf28c0407d5f433095eef4ead44c25e3834306ea00660dfff3

Observation 85394fce-e6b3-4583-8371-1cd84c11e6bd · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Flow straight and fast: Learning to generate and transfer data with rectified flow

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.325402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:d48b5be60f11bfadc46b7d07ab197a17728a253558d03c96cad2f32839e0b7e0

Observation 17f8027b-7823-4084-8e0e-4cbfee6274ed · outbound

This paper cites Decoupled weight de- cay regularization.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Decoupled weight de- cay regularization

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.333215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:ea74b3d21c96d83188f0b43283e5fca15e6bec99b2d6bd6236a9c4c290f57f2d

Observation b388d655-bc4e-4863-a8fe-90016b242897 · outbound

This paper cites Diff-foley: Synchronized video-to-audio synthesis with la- tent diffusion models.Advances in Neural Information Pro- cessing Systems, 36:48855–48876.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Diff-foley: Synchronized video-to-audio synthesis with la- tent diffusion models.Advances in Neural Information Pro- cessing Systems, 36:48855–48876

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.321582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:06b987c1a8a613bb021ce1f62f0cd1014404976302cd7e57a9f6671b25bf6f63

Observation 5cf8fa71-ef8f-40a9-ab8a-45d733574964 · outbound

This paper cites Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.329158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:e491298983c228167581501b8bd24b4723234f7d1db574e3ad923e0b0110197d

Observation 67070e8e-2bf4-4f95-aa9e-d6c6a8ecda85 · outbound

This paper cites Pytorch: An im- perative style, high-performance deep learning library.Ad- vances in neural information processing systems, 32.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Pytorch: An im- perative style, high-performance deep learning library.Ad- vances in neural information processing systems, 32

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.317122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:d29dd08853bc477b378252d40a5d0563ca2f25b30817c7c46ff669807c64c248

Observation d9a3efc5-8c9e-4cd0-88fc-331eda484d60 · outbound

This paper cites Scalable diffusion models with transformers.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Scalable diffusion models with transformers

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.312768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:ea1e0d4cdea825fe455357e28249f4cc1d3b690a368bb691462e66b934a96224

Observation db27bd8a-e00f-4aae-b0d5-e6578a1b8a81 · outbound

This paper cites Learning transferable visual models from natural language supervision.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Learning transferable visual models from natural language supervision

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.299666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:541ef70698cca9aa6603cb72767ecfba798ef9555ad3b1f022af644b4a4c3c85

Observation b8a5b66b-b1ce-4a39-95ee-c8ddca8467af · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Robust speech recognition via large-scale weak supervision

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.303958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:6a628f48018620705d2808baa8593beb6b61e5053babaafbaa6bb64123fac86c

Observation e8695b06-01ef-4ff9-89ae-83c5f70cc6d6 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.308183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:1fd6f3676fc316fa19d442336768c47431f98684ea57bccf9f1788bf3585b90d

Observation 0b3f25cf-3ae0-4bfb-8a04-bf26a5b63826 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text High-resolution image synthesis with latent diffusion models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.337089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:5644311b1dfdb18c5c26ab3221e95b7624cd9682e279a50be9dc215a959cb15d

Observation 11a95843-eead-4cc0-a314-db33cee3c359 · outbound

This paper cites HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:47.496754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:bca4689ed87c269b94a91157518387e4b96386783afdfeb5d3054e91463ce7d7

Observation e3f2adc0-e94b-406a-b0aa-db1d6ff79c72 · outbound

This paper cites I hear your true colors: Im- age guided audio generation.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text I hear your true colors: Im- age guided audio generation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.349434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:795ac99aeddf9c35298f506243df83b2ac1aafd1398b8b0c9588ff46b91e3844

Observation dbe4e294-7924-4f73-b6e8-a00a45ce0ea5 · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Make-a-video: Text-to-video generation without text-video data

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.483737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:e6524dabf00a6c36bbcb94573bb43bf74e9e39ea8741833dcbc05886243c1b00

Observation a4d74983-a2cb-4444-aa8c-d3c32bececf2 · outbound

This paper cites Score-based generative modeling through stochastic differential equa- tions.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Score-based generative modeling through stochastic differential equa- tions

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.378333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:6d5e3ffd6afefa8f422d2ec829045df4304864d0527c6c44fe938678baf7aaa9

Observation c946163e-c7ca-45fa-8784-d5988bcf4ad9 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.383158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:ddf83092090a5af0eb0aae47e7a7ab0005355c9d5fdbbc96cd412ea698842913

Observation 2f52d7cf-c424-4d50-b040-5e629ace856a · outbound

This paper cites Naturalspeech: End-to-end text-to-speech synthesis with human-level quality.IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(6):4234–4245.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Naturalspeech: End-to-end text-to-speech synthesis with human-level quality.IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(6):4234–4245

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.392337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:01fa8ea08287efd661863cbdc0fe53c80048747ff95ecc80705434d7e486d6ce

Observation ab2a469d-4513-4a1d-b1ce-c12323d55708 · outbound

This paper cites V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foun- dation models.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foun- dation models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.405129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:c8bd68a54db5d5c4681ee5105353037e57475a769e4f5ad3e58d85371ff5efb3

Observation 175a2779-c67e-4365-b733-3e4900bcdfd9 · outbound

This paper cites Frieren: Efficient video-to-audio generation network with rectified flow matching.Advances in neural information pro- cessing systems, 37:128118–128138.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Frieren: Efficient video-to-audio generation network with rectified flow matching.Advances in neural information pro- cessing systems, 37:128118–128138

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.353573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:74459e483448c1365bc5e0326110cdfaceef528243ce033836b18420492b9b0c

Observation 9ff5c10b-551f-4cf5-9ba3-4911237e2de8 · outbound

This paper cites Wav2clip: Learning robust audio repre- sentations from clip.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Wav2clip: Learning robust audio repre- sentations from clip

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.357445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:aa5efeefba19191a9a03b3d42f36b8e8eebb005726b2bf86097d457efd3e16e8

Observation 5b64560b-b850-4f77-a50f-801584deec55 · outbound

This paper cites Son- icvisionlm: Playing sound with vision language models.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Son- icvisionlm: Playing sound with vision language models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.361354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:6d1cd0b979e60137fd0fab4eb56bcfbdde3af704a19c715e2f7bcea5437b9d12

Observation 652fc853-fccd-41c3-a7cc-f63e44889002 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:47.523966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:3e673a960f6aa3f3c30f3b15f3171f360936325a9acf54006267a2754e1de880

Observation b6f7e65f-b241-4309-b285-fd1a76bfcdfd · outbound

This paper cites condition–unconditional.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text condition–unconditional

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.365145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:c04c5470f5241653dec130e3e6313766e972e40435ad340eed7f82507892bd50

Pith citing papers

Observation f703365e-a69d-494b-bbf3-5a12f52a8deb · inbound

Do Joint Audio-Video Generation Models Understand Physics? cites this paper.

Do Joint Audio-Video Generation Models Understand Physics? OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.233900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:d4275ed607e70c78d3d07567ae1cc1b99d07d245d4f5df4cc625ab64e3d7f513