Pith. sign in

Paper Citation Record · LEDGER

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet

As of 8 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 1 inbound Pith citation observation for arXiv:2505.16195.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16195 v2

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:09:58.262045Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T08:20:02.986562Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T08:21:06.814521Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation abc09efc-c13d-4be5-8260-8f0b94acdb03 · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet AudioGen: Textually Guided Audio Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.545341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.545341Z digest=sha256:6579114b5a8b631b9596b776c8e9b1748a5e48a4f4ae6cb014377974008ee726

Observation 7953642e-47fc-455d-a2b1-c9fc8739fcc3 · outbound

This paper cites Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.619538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.619538Z digest=sha256:d2c7dd686bf35325bab9aced9f584b46392b5536a4a0c4f96fee7fb3f9032813

Observation 10e5f49b-3c03-4f1f-b5df-9f047e5e0bfd · outbound

This paper cites SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.712485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.712485Z digest=sha256:51c090996702913971649996a5e986d13583cb7bbd2bd575118373ca0547ef51

Observation 66c067d3-bdd2-4644-98ba-2b6fbf9c6141 · outbound

This paper cites SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound Generation.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.810995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.810995Z digest=sha256:228d8896891aec0198b4c996d6e24d7c4e876602a7f971ec32b6f5deb9bc77b8

Observation fe976c38-6e02-4d78-a072-a320e56aaf45 · outbound

This paper cites Stable audio open,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Stable audio open,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:02.801403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:54.911275Z digest=sha256:01bbe409f7e46931a6cc69be065e7f080c3c707b8e612768c276ed7b63b56ca2

Observation 4ba353cf-9ed8-4907-825a-6f6705aa76b4 · outbound

This paper cites MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.999597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.999597Z digest=sha256:2b098e8aa6158d02f7a8ad165d5ee2cda0f0705206457e4f54ee93b5272f1aeb

Observation 9b22797a-76b4-46ae-886b-f4261d8e063c · outbound

This paper cites Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:02.647919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:55.139356Z digest=sha256:c648ec1a4178655b1c1254dc788130255d7938514fb707d404e2b4d3b93cd3db

Observation 23d57b05-ab77-49f4-88f7-9f99a863cba2 · outbound

This paper cites Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.278233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.278233Z digest=sha256:f61643d4d605308c242e2389fec10807697fe0843f37e781b617472ae684c146

Observation e8818c41-138b-4f53-ab6e-af61e8b28b03 · outbound

This paper cites V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:02.496790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:55.387681Z digest=sha256:190e1166c900d51b9f517e7324730a1665a781231df2f09b3848578c9569cf44

Observation cfd683ff-f0a3-4a25-baab-cead609353e7 · outbound

This paper cites Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:02.352055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:55.487341Z digest=sha256:51ed391c8c35188bdff14a7a66c2992c188b24b1db998845fa98ecceea3236d8

Observation 9fef9217-ca06-447b-8cd2-15c042f266c9 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.547205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.547205Z digest=sha256:c705322dac4cfd154699cc9a246ae867fdf48b43edbb4ddea708244d16d6767d

Observation 0ddf2160-429f-49b1-b0ad-bcdd0b423e1f · outbound

This paper cites Synchformer: Efficient synchronization from sparse cues,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Synchformer: Efficient synchronization from sparse cues,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:02.207841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:55.606504Z digest=sha256:c547a6c5e1c0823aa7ee5fe4677cbfa7f34cf7f12fb816326eff0d3824028f52

Observation 8994e6ca-0abb-4e83-b1c4-652cdca8e598 · outbound

This paper cites Adding conditional control to text- to-image diffusion models,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Adding conditional control to text- to-image diffusion models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:02.020964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:55.683396Z digest=sha256:27beba668805c7558d081d46a89c75ae50248ea4e2bd8aba269d9a227e0a6df6

Observation 3cc6f8c6-e56a-4b03-b059-a741ae33fc09 · outbound

This paper cites Uni-controlnet: All-in-one control to text-to-image diffusion models,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Uni-controlnet: All-in-one control to text-to-image diffusion models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:01.857838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:55.753711Z digest=sha256:762a853c4caacad08a1c0ed45d096469a1d91d806a601f7a7a4a780a88486e46

Observation 110c6b42-6d76-4411-a425-12ee6748e1f7 · outbound

This paper cites Read, Watch and Scream! Sound Generation from Text and Video.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Read, Watch and Scream! Sound Generation from Text and Video

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.831770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.831770Z digest=sha256:dab28ba7f6e0d6c92377fb97b594feb68beadadc2ed285db545ab5af1110cf9e

Observation 72a83755-a8c4-4920-b509-daf337c91156 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.910310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.910310Z digest=sha256:5b7c42f8c6ea1eaf1b83ebf62a9a18bec1b44854435188341ee6d6a819a6a172

Observation 0dd4b443-d5f0-4fa6-a51c-394681685a84 · outbound

This paper cites Tell What You Hear From What You See -- Video to Audio Generation Through Text.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Tell What You Hear From What You See -- Video to Audio Generation Through Text

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.988743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.988743Z digest=sha256:9ac3c699b25dfe02e097a48d06039494e8111fb899c2acfb450d0d6599ac5e8b

Observation 6278bfb9-8f1b-4e1a-8085-cc24982c664e · outbound

This paper cites Temporally aligned audio for video with autoregression,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Temporally aligned audio for video with autoregression,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:01.708710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:56.045325Z digest=sha256:03f97f6638e15ff0833b980f9729e472caac30f19f5f4e10409c1bbcfbc7c353

Observation f2a7234f-eedb-4911-a3eb-28b36912076e · outbound

This paper cites Frieren: Efficient video-to-audio generation network with rectified flow matching,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Frieren: Efficient video-to-audio generation network with rectified flow matching,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:01.422859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:56.119327Z digest=sha256:701c8251f110d1662b802d9f1bcc016aea326497d2d87727fb289dc7f2b22098

Observation e6501fb0-d017-4ff1-9d9d-438c83298075 · outbound

This paper cites Mavil: Masked audio-video learners,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Mavil: Masked audio-video learners,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:01.261287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:56.195962Z digest=sha256:b1cdbce1666bd169911b90a317d0b1fe3e78400dcedae0366d122a6948f371c2

Observation 3cc382ef-08bf-40c2-a2e7-7cbbe25e8cc4 · outbound

This paper cites Diff-foley: Synchronized video- to-audio synthesis with latent diffusion models,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Diff-foley: Synchronized video- to-audio synthesis with latent diffusion models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:01.089076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:56.264683Z digest=sha256:520df06af652e8ca1c917456a5081b2590a4c3e9030bf1cc47d063d29936dd06

Observation 6cb4d445-1e69-4e49-ba26-f9ebff2361a3 · outbound

This paper cites FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.308460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.308460Z digest=sha256:25ceebee088361ccaf22451d719bbcecce331b20ddaa3d8f6f7b3a67941b6cab

Observation 1153f749-b0bd-4fd4-b5e1-da4c0392e7da · outbound

This paper cites Smooth-foley: Creating continuous sound for video-to-audio generation under semantic guidance,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Smooth-foley: Creating continuous sound for video-to-audio generation under semantic guidance,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.923321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:56.376868Z digest=sha256:e4a0a28e1ff7981656d82868904ef982823ba3a057529ffb3fc0e363bdc842b9

Observation 18b0951f-4fe7-4acb-a6e8-fc7fc4972a5f · outbound

This paper cites Learning transferable visual models from natural language supervision,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Learning transferable visual models from natural language supervision,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.765043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:56.424504Z digest=sha256:3868b2a14e18f4b966c83156d0df79431565610ca7852cba8b457b2a4383e735

Observation 34ad2f69-31e0-421e-aa52-a94edb43a1cc · outbound

This paper cites Maskgit: Masked generative image transformer,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Maskgit: Masked generative image transformer,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.596551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:56.497923Z digest=sha256:7f5cc8b3f38c655b87626d37bdcff5abffaa0f6b834dff7c411cc556d93f2304

Observation 389e82fd-b576-42d8-b9dc-c7eb68d08266 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.595415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.595415Z digest=sha256:857dae3f2874406f889749c8cfc4a48ab9460b7b279b56a3c5c0f874ae9a902e

Observation 1648a738-ef21-43fe-8484-e30a2be13b05 · outbound

This paper cites Music Foundation Model as Generic Booster for Music Downstream Tasks.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Music Foundation Model as Generic Booster for Music Downstream Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.680530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.680530Z digest=sha256:dc5e660896dfd24f77d1863f1aa13c7d1d22d5d5bdda992d5bfa4f7a6bc3f010

Observation 4ba595d1-8f83-41b5-b393-85b15143268d · outbound

This paper cites High Fidelity Neural Audio Compression.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet High Fidelity Neural Audio Compression

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.761212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.761212Z digest=sha256:118a4f8db7b99950dfc92d5a9e41c92dc368f168816999ac4d4a697c9825b3ec

Observation bc0b6b04-276c-4456-8638-737ada8b65fd · outbound

This paper cites High- fidelity audio compression with improved rvqgan,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet High- fidelity audio compression with improved rvqgan,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.473120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:56.840072Z digest=sha256:d3dc40ac3bce7a73729abdc26ef6b8190b175ee2916c0a66ed979ed24fbcd5e8

Observation 52abba22-8b04-45dd-aacc-3e6d083918f9 · outbound

This paper cites Taming Visually Guided Sound Generation.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Taming Visually Guided Sound Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.889056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.889056Z digest=sha256:6e9e786df3c630223dc7eee2d3b32ccfaa84a6dc892f82f9c9381b3ab5de0d3c

Observation ea1a425d-b872-4299-bf93-254e01388d49 · outbound

This paper cites Masked autoencoders are scalable vision learners,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Masked autoencoders are scalable vision learners,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.382535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:56.939962Z digest=sha256:5e5bfecf6a8a89bbb46e416bcea357f3aec8d4b9f9551cfe17f9cb1e9bf185e5

Observation 3e56c4ab-42ed-407a-8092-60bd86077502 · outbound

This paper cites Extending audio masked autoencoders toward audio restoration,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Extending audio masked autoencoders toward audio restoration,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.237928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:57.006139Z digest=sha256:b94c2ecfe2e952c317559874feeaabf4c8d157b6e517df7d016ba706c2762bdd

Observation 87d13178-79e3-43d7-b824-dede84d37966 · outbound

This paper cites Mage: Masked generative encoder to unify representation learning and image synthesis,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Mage: Masked generative encoder to unify representation learning and image synthesis,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.021600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:57.058562Z digest=sha256:56d72433441cdeb667d276446e818d34b1bb60c884beeab2b9e9e923b07b089d

Observation ead1e612-6841-4a2b-817b-ab9faebc6d7d · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:57.138047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:57.138047Z digest=sha256:c25ddf8f114fae8b16ba5b7ed6857a3f914d1fb3f62c0f34071570269546eb67

Observation 158aff55-058b-43f9-a4a9-4fc26dd9b46b · outbound

This paper cites PIXART-{\delta}: Fast and Controllable Image Generation with Latent Consistency Models.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet PIXART-{\delta}: Fast and Controllable Image Generation with Latent Consistency Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:57.179030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:57.179030Z digest=sha256:7ca82fa09122c3f85f7dbfcdd21820e6f0e8b264a9b0885f07d2574d30b69e5b

Observation c4db9e3c-35ea-45c6-aa2f-f3f918488ffb · outbound

This paper cites COCOLA: Coherence-Oriented Contrastive Learning of Musical Audio Representations.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet COCOLA: Coherence-Oriented Contrastive Learning of Musical Audio Representations

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:09:58.450395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:57.248961Z digest=sha256:8bde5c88cf9608d528a924e1a5e634c573e2efe2a57d2e2dac998dcae3456c00

Observation d64c8f29-6bdc-49fe-a011-3afb21821ef2 · outbound

This paper cites Editing music with melody and text: Using controlnet for diffusion transformer,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Editing music with melody and text: Using controlnet for diffusion transformer,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.778127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:57.354123Z digest=sha256:e20aedf9b05aa67d3408f8b20d19a2c13ba599bdccf88068dbbe01c804f8ed9b

Observation 5e669fe6-ece2-4d8e-88cf-6ced1feec5be · outbound

This paper cites Classifier-Free Diffusion Guidance.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Classifier-Free Diffusion Guidance

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:57.420079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:57.420079Z digest=sha256:6049bf255d10ea997ae8349d5d1f88f45ca43f179ac669f61bd93fce72656955

Observation e77c68c5-45da-461b-aeba-416f1234010b · outbound

This paper cites Muse: Text-To-Image Generation via Masked Generative Transformers.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:57.486525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:57.486525Z digest=sha256:4fa5f01915e56c35baa8c92601d0a4677fdc73a449f2c4d66af0287b9d10c551

Observation b1ce6187-a2e0-41d4-842d-2858f15c4994 · outbound

This paper cites Stemgen: A music generation model that listens,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Stemgen: A music generation model that listens,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.611073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:57.567373Z digest=sha256:21c6ce2b4a430846414eae88065d263a483e4838293c1dfae21dbf2b98d78bcd

Observation 1c375ab3-c629-4e01-8ec7-e5d4b76227e4 · outbound

This paper cites Imagebind: One embedding space to bind them all,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Imagebind: One embedding space to bind them all,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.447837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:57.668416Z digest=sha256:254ecdf353809cec7772aaa9b5dfe70e2426e20a6f8aba1e62580d8eebad37f4

Observation 899ae236-eb3b-43c4-83cd-44ea45f0aa47 · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.318967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:57.772268Z digest=sha256:1f68376174620448758384ec192950750b2bb0e6b94ab252da73f494d033c9aa

Observation f6054d9e-adc8-4eb1-aca5-4fb921a2b459 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Audio set: An ontology and human-labeled dataset for audio events,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.207305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:57.862288Z digest=sha256:122a7c96c9aedb4d5acbc38eefc766b81c0b6848f70d030cf4a8f972227b4642

Observation a066d024-477c-4e68-9145-818d1990d3b7 · outbound

This paper cites Efficient Training of Audio Transformers with Patchout.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Efficient Training of Audio Transformers with Patchout

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:57.929534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:57.929534Z digest=sha256:f53a4902538b84b0cb619ed366d0a8589943fa7bb32bd86289a7a70ec65fcbab

Observation 9d1e8033-5093-4836-8fc2-544e8816659a · outbound

This paper cites Vggsound: A large- scale audio-visual dataset,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Vggsound: A large- scale audio-visual dataset,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.139077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:58.028545Z digest=sha256:f8733cfa91297493887507ce574f8b45e7ac32a003243a466ed06037f3e895ad

Observation 330a7253-e81a-4013-98aa-49fd7a4ff704 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.019857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:58.119664Z digest=sha256:efdbdf9a5d1fc1dde4da7b1b24ed2d1c296a44a78aad9549aea913480b4fc2d8

Observation ecda695c-59e4-4d15-b432-4efb0dcb829f · outbound

This paper cites Cnn architectures for large-scale audio classification,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Cnn architectures for large-scale audio classification,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:58.866089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:58.194238Z digest=sha256:ceb79774fb8bf7636d6f9dc2e8f38e945a3d7461f97515880ea1feda74801510

Observation 9efe0af6-4ce8-454b-8bf9-b24274ca4411 · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Panns: Large-scale pretrained audio neural networks for audio pattern recognition,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:58.751950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:58.262045Z digest=sha256:5839c1b9fef8e9f0bcdf2c005359fd125c41339a4697a3647ab5e04fe4a6bc96

Pith citing papers

Observation 2678bf73-aeaa-4275-b270-963fb3ea0938 · inbound

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation cites this paper.

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:21:06.817478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T08:20:02.986562Z digest=sha256:fa3cb88c292d0aef6dbb74049eabb9c1d5e5bce91669f80b65a4a348892f854b