Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:09:58.262045Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 1 inbound Pith citation observation for arXiv:2505.16195.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:09:58.262045Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-18T08:20:02.986562Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-18T08:21:06.814521Z
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation abc09efc-c13d-4be5-8260-8f0b94acdb03 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet AudioGen: Textually Guided Audio Generation
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7953642e-47fc-455d-a2b1-c9fc8739fcc3 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10e5f49b-3c03-4f1f-b5df-9f047e5e0bfd · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66c067d3-bdd2-4644-98ba-2b6fbf9c6141 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound Generation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe976c38-6e02-4d78-a072-a320e56aaf45 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Stable audio open,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4ba353cf-9ed8-4907-825a-6f6705aa76b4 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b22797a-76b4-46ae-886b-f4261d8e063c · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 23d57b05-ab77-49f4-88f7-9f99a863cba2 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8818c41-138b-4f53-ab6e-af61e8b28b03 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cfd683ff-f0a3-4a25-baab-cead609353e7 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9fef9217-ca06-447b-8cd2-15c042f266c9 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ddf2160-429f-49b1-b0ad-bcdd0b423e1f · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Synchformer: Efficient synchronization from sparse cues,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8994e6ca-0abb-4e83-b1c4-652cdca8e598 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Adding conditional control to text- to-image diffusion models,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3cc6f8c6-e56a-4b03-b059-a741ae33fc09 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Uni-controlnet: All-in-one control to text-to-image diffusion models,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 110c6b42-6d76-4411-a425-12ee6748e1f7 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Read, Watch and Scream! Sound Generation from Text and Video
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72a83755-a8c4-4920-b509-daf337c91156 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dd4b443-d5f0-4fa6-a51c-394681685a84 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Tell What You Hear From What You See -- Video to Audio Generation Through Text
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6278bfb9-8f1b-4e1a-8085-cc24982c664e · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Temporally aligned audio for video with autoregression,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f2a7234f-eedb-4911-a3eb-28b36912076e · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Frieren: Efficient video-to-audio generation network with rectified flow matching,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e6501fb0-d017-4ff1-9d9d-438c83298075 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Mavil: Masked audio-video learners,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3cc382ef-08bf-40c2-a2e7-7cbbe25e8cc4 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Diff-foley: Synchronized video- to-audio synthesis with latent diffusion models,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6cb4d445-1e69-4e49-ba26-f9ebff2361a3 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1153f749-b0bd-4fd4-b5e1-da4c0392e7da · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Smooth-foley: Creating continuous sound for video-to-audio generation under semantic guidance,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 18b0951f-4fe7-4acb-a6e8-fc7fc4972a5f · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Learning transferable visual models from natural language supervision,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 34ad2f69-31e0-421e-aa52-a94edb43a1cc · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Maskgit: Masked generative image transformer,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 389e82fd-b576-42d8-b9dc-c7eb68d08266 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1648a738-ef21-43fe-8484-e30a2be13b05 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Music Foundation Model as Generic Booster for Music Downstream Tasks
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ba595d1-8f83-41b5-b393-85b15143268d · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet High Fidelity Neural Audio Compression
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc0b6b04-276c-4456-8638-737ada8b65fd · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet High- fidelity audio compression with improved rvqgan,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 52abba22-8b04-45dd-aacc-3e6d083918f9 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Taming Visually Guided Sound Generation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea1a425d-b872-4299-bf93-254e01388d49 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Masked autoencoders are scalable vision learners,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3e56c4ab-42ed-407a-8092-60bd86077502 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Extending audio masked autoencoders toward audio restoration,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 87d13178-79e3-43d7-b824-dede84d37966 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Mage: Masked generative encoder to unify representation learning and image synthesis,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ead1e612-6841-4a2b-817b-ab9faebc6d7d · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet AudioLDM: Text-to-Audio Generation with Latent Diffusion Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 158aff55-058b-43f9-a4a9-4fc26dd9b46b · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet PIXART-{\delta}: Fast and Controllable Image Generation with Latent Consistency Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4db9e3c-35ea-45c6-aa2f-f3f918488ffb · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet COCOLA: Coherence-Oriented Contrastive Learning of Musical Audio Representations
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d64c8f29-6bdc-49fe-a011-3afb21821ef2 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Editing music with melody and text: Using controlnet for diffusion transformer,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5e669fe6-ece2-4d8e-88cf-6ced1feec5be · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Classifier-Free Diffusion Guidance
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e77c68c5-45da-461b-aeba-416f1234010b · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Muse: Text-To-Image Generation via Masked Generative Transformers
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1ce6187-a2e0-41d4-842d-2858f15c4994 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Stemgen: A music generation model that listens,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1c375ab3-c629-4e01-8ec7-e5d4b76227e4 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Imagebind: One embedding space to bind them all,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 899ae236-eb3b-43c4-83cd-44ea45f0aa47 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6054d9e-adc8-4eb1-aca5-4fb921a2b459 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Audio set: An ontology and human-labeled dataset for audio events,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a066d024-477c-4e68-9145-818d1990d3b7 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Efficient Training of Audio Transformers with Patchout
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d1e8033-5093-4836-8fc2-544e8816659a · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Vggsound: A large- scale audio-visual dataset,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 330a7253-e81a-4013-98aa-49fd7a4ff704 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ecda695c-59e4-4d15-b432-4efb0dcb829f · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Cnn architectures for large-scale audio classification,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9efe0af6-4ce8-454b-8bf9-b24274ca4411 · outbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Panns: Large-scale pretrained audio neural networks for audio pattern recognition,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2678bf73-aeaa-4275-b270-963fb3ea0938 · inbound
MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.