Pith. sign in

Paper Citation Record · LEDGER

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet

As of 8 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 1 inbound Pith citation observation for arXiv:2505.16195.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16195 v2

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:09:58.262045Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T08:20:02.986562Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T08:21:06.814521Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation abc09efc-c13d-4be5-8260-8f0b94acdb03 · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet AudioGen: Textually Guided Audio Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.545341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.545341Z digest=sha256:6579114b5a8b631b9596b776c8e9b1748a5e48a4f4ae6cb014377974008ee726

Observation 7953642e-47fc-455d-a2b1-c9fc8739fcc3 · outbound

This paper cites Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.619538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.619538Z digest=sha256:6f1e3067f32559c99264ab5055cc80be147ce38d1039c2f458f41851e3aae2c6

Observation 10e5f49b-3c03-4f1f-b5df-9f047e5e0bfd · outbound

This paper cites SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.712485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.712485Z digest=sha256:51c090996702913971649996a5e986d13583cb7bbd2bd575118373ca0547ef51

Observation 66c067d3-bdd2-4644-98ba-2b6fbf9c6141 · outbound

This paper cites SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound Generation.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.810995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.810995Z digest=sha256:228d8896891aec0198b4c996d6e24d7c4e876602a7f971ec32b6f5deb9bc77b8

Observation fe976c38-6e02-4d78-a072-a320e56aaf45 · outbound

This paper cites Stable audio open,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Stable audio open,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:02.801403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:54.911275Z digest=sha256:2f1ec1a9bc5be07c26bb3e01780753fbc3adfda8ade1e9e07b4e4f296bf785dd

Observation 4ba353cf-9ed8-4907-825a-6f6705aa76b4 · outbound

This paper cites MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.999597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.999597Z digest=sha256:2b098e8aa6158d02f7a8ad165d5ee2cda0f0705206457e4f54ee93b5272f1aeb

Observation 9b22797a-76b4-46ae-886b-f4261d8e063c · outbound

This paper cites Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:02.647919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:55.139356Z digest=sha256:6df933834e5158127ee6e9a6d4c40ddca6723f2938ec067151d794be905efc16

Observation 23d57b05-ab77-49f4-88f7-9f99a863cba2 · outbound

This paper cites Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.278233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.278233Z digest=sha256:f61643d4d605308c242e2389fec10807697fe0843f37e781b617472ae684c146

Observation e8818c41-138b-4f53-ab6e-af61e8b28b03 · outbound

This paper cites V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:02.496790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:55.387681Z digest=sha256:20f8e2deb57757d9a75910583fc7ae7cd5d75d33c12260e5a3dc75d0625587e8

Observation cfd683ff-f0a3-4a25-baab-cead609353e7 · outbound

This paper cites Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:02.352055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:55.487341Z digest=sha256:e49e358aa10b3b30263b97a4ae015b920a6c83cae655877f58c04b222800ed90

Observation 9fef9217-ca06-447b-8cd2-15c042f266c9 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.547205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.547205Z digest=sha256:c705322dac4cfd154699cc9a246ae867fdf48b43edbb4ddea708244d16d6767d

Observation 0ddf2160-429f-49b1-b0ad-bcdd0b423e1f · outbound

This paper cites Synchformer: Efficient synchronization from sparse cues,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Synchformer: Efficient synchronization from sparse cues,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:02.207841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:55.606504Z digest=sha256:9d0d20186fa05941dab5d72e32bf444fa6f457f8c651cccd22531393e94c0e45

Observation 8994e6ca-0abb-4e83-b1c4-652cdca8e598 · outbound

This paper cites Adding conditional control to text- to-image diffusion models,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Adding conditional control to text- to-image diffusion models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:02.020964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:55.683396Z digest=sha256:bcf35a5472f7a574f002668a4a8816a08ffb1f75e9933089a024395b071aa5f5

Observation 3cc6f8c6-e56a-4b03-b059-a741ae33fc09 · outbound

This paper cites Uni-controlnet: All-in-one control to text-to-image diffusion models,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Uni-controlnet: All-in-one control to text-to-image diffusion models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:01.857838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:55.753711Z digest=sha256:839f31f1b5c33f309d21e2d78e93ec8f5d91681a6ae3b21b3f303ec242d34c76

Observation 110c6b42-6d76-4411-a425-12ee6748e1f7 · outbound

This paper cites Read, Watch and Scream! Sound Generation from Text and Video.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Read, Watch and Scream! Sound Generation from Text and Video

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.831770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.831770Z digest=sha256:dab28ba7f6e0d6c92377fb97b594feb68beadadc2ed285db545ab5af1110cf9e

Observation 72a83755-a8c4-4920-b509-daf337c91156 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.910310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.910310Z digest=sha256:5b7c42f8c6ea1eaf1b83ebf62a9a18bec1b44854435188341ee6d6a819a6a172

Observation 0dd4b443-d5f0-4fa6-a51c-394681685a84 · outbound

This paper cites Tell What You Hear From What You See -- Video to Audio Generation Through Text.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Tell What You Hear From What You See -- Video to Audio Generation Through Text

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.988743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.988743Z digest=sha256:9ac3c699b25dfe02e097a48d06039494e8111fb899c2acfb450d0d6599ac5e8b

Observation 6278bfb9-8f1b-4e1a-8085-cc24982c664e · outbound

This paper cites Temporally aligned audio for video with autoregression,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Temporally aligned audio for video with autoregression,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:01.708710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:56.045325Z digest=sha256:ace7626856c1b71de9b67ddfd16d97472e6134515e576a1c67f8ebc9de6fc41f

Observation f2a7234f-eedb-4911-a3eb-28b36912076e · outbound

This paper cites Frieren: Efficient video-to-audio generation network with rectified flow matching,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Frieren: Efficient video-to-audio generation network with rectified flow matching,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:01.422859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:56.119327Z digest=sha256:542e158c138580a8c52f3320daf727d7bb569dccf7c6ddad083ad52389e8ab59

Observation e6501fb0-d017-4ff1-9d9d-438c83298075 · outbound

This paper cites Mavil: Masked audio-video learners,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Mavil: Masked audio-video learners,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:01.261287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:56.195962Z digest=sha256:118eee4b39d732105e64cb25707629f1bee4836052ac521eba7fb967393dfe18

Observation 3cc382ef-08bf-40c2-a2e7-7cbbe25e8cc4 · outbound

This paper cites Diff-foley: Synchronized video- to-audio synthesis with latent diffusion models,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Diff-foley: Synchronized video- to-audio synthesis with latent diffusion models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:01.089076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:56.264683Z digest=sha256:09706d1aabd64a929c337a0c4d7593ec077d1ba9e1ec90c4beb0daeb70727c70

Observation 6cb4d445-1e69-4e49-ba26-f9ebff2361a3 · outbound

This paper cites FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.308460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.308460Z digest=sha256:25ceebee088361ccaf22451d719bbcecce331b20ddaa3d8f6f7b3a67941b6cab

Observation 1153f749-b0bd-4fd4-b5e1-da4c0392e7da · outbound

This paper cites Smooth-foley: Creating continuous sound for video-to-audio generation under semantic guidance,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Smooth-foley: Creating continuous sound for video-to-audio generation under semantic guidance,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.923321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:56.376868Z digest=sha256:bb9e523ea8cc5af4184fe4ed6dc41f4a113d5dba194bef2ed2a2fc295f19dce2

Observation 18b0951f-4fe7-4acb-a6e8-fc7fc4972a5f · outbound

This paper cites Learning transferable visual models from natural language supervision,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Learning transferable visual models from natural language supervision,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.765043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:56.424504Z digest=sha256:c38641081cbaeeb5623f2dedc12058830e50806a806cc0186699f165aba5b538

Observation 34ad2f69-31e0-421e-aa52-a94edb43a1cc · outbound

This paper cites Maskgit: Masked generative image transformer,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Maskgit: Masked generative image transformer,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.596551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:56.497923Z digest=sha256:3412211127dffe4a497763b73d2c0242fd47fe72a35791d5b93835dd4217d920

Observation 389e82fd-b576-42d8-b9dc-c7eb68d08266 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.595415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.595415Z digest=sha256:857dae3f2874406f889749c8cfc4a48ab9460b7b279b56a3c5c0f874ae9a902e

Observation 1648a738-ef21-43fe-8484-e30a2be13b05 · outbound

This paper cites Music Foundation Model as Generic Booster for Music Downstream Tasks.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Music Foundation Model as Generic Booster for Music Downstream Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.680530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.680530Z digest=sha256:dc5e660896dfd24f77d1863f1aa13c7d1d22d5d5bdda992d5bfa4f7a6bc3f010

Observation 4ba595d1-8f83-41b5-b393-85b15143268d · outbound

This paper cites High Fidelity Neural Audio Compression.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet High Fidelity Neural Audio Compression

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.761212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.761212Z digest=sha256:118a4f8db7b99950dfc92d5a9e41c92dc368f168816999ac4d4a697c9825b3ec

Observation bc0b6b04-276c-4456-8638-737ada8b65fd · outbound

This paper cites High- fidelity audio compression with improved rvqgan,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet High- fidelity audio compression with improved rvqgan,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.473120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:56.840072Z digest=sha256:a3c1ef0f62d4230285a180cc72959004fa6a4f9eafd7f1acb150ea8999a370f9

Observation 52abba22-8b04-45dd-aacc-3e6d083918f9 · outbound

This paper cites Taming Visually Guided Sound Generation.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Taming Visually Guided Sound Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.889056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.889056Z digest=sha256:6e9e786df3c630223dc7eee2d3b32ccfaa84a6dc892f82f9c9381b3ab5de0d3c

Observation ea1a425d-b872-4299-bf93-254e01388d49 · outbound

This paper cites Masked autoencoders are scalable vision learners,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Masked autoencoders are scalable vision learners,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.382535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:56.939962Z digest=sha256:8b41dc07d179821bdf66cabf9bbd75d4d5bbd9d94d87b3111aa5f2c0e6bf0478

Observation 3e56c4ab-42ed-407a-8092-60bd86077502 · outbound

This paper cites Extending audio masked autoencoders toward audio restoration,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Extending audio masked autoencoders toward audio restoration,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.237928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:57.006139Z digest=sha256:b3f131ccf29b594329a081edb8b98e4b52309672db622c005e0d90c805d49606

Observation 87d13178-79e3-43d7-b824-dede84d37966 · outbound

This paper cites Mage: Masked generative encoder to unify representation learning and image synthesis,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Mage: Masked generative encoder to unify representation learning and image synthesis,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.021600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:57.058562Z digest=sha256:c22b47a412d0f6237ad29e1457ff07c1cbb9132ee65a05969236569be1c0a0d1

Observation ead1e612-6841-4a2b-817b-ab9faebc6d7d · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:57.138047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:57.138047Z digest=sha256:c25ddf8f114fae8b16ba5b7ed6857a3f914d1fb3f62c0f34071570269546eb67

Observation 158aff55-058b-43f9-a4a9-4fc26dd9b46b · outbound

This paper cites PIXART-{\delta}: Fast and Controllable Image Generation with Latent Consistency Models.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet PIXART-{\delta}: Fast and Controllable Image Generation with Latent Consistency Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:57.179030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:57.179030Z digest=sha256:7ca82fa09122c3f85f7dbfcdd21820e6f0e8b264a9b0885f07d2574d30b69e5b

Observation c4db9e3c-35ea-45c6-aa2f-f3f918488ffb · outbound

This paper cites COCOLA: Coherence-Oriented Contrastive Learning of Musical Audio Representations.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet COCOLA: Coherence-Oriented Contrastive Learning of Musical Audio Representations

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:09:58.450395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:57.248961Z digest=sha256:6e02d93809e364f60b5aa8cef429731fdc2985a6675c7f2daa12dd94e5ac998b

Observation d64c8f29-6bdc-49fe-a011-3afb21821ef2 · outbound

This paper cites Editing music with melody and text: Using controlnet for diffusion transformer,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Editing music with melody and text: Using controlnet for diffusion transformer,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.778127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:57.354123Z digest=sha256:69b71ce0232e418a0197a475375a2599b7cff5129711f518ed5c9bdbd4218c82

Observation 5e669fe6-ece2-4d8e-88cf-6ced1feec5be · outbound

This paper cites Classifier-Free Diffusion Guidance.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Classifier-Free Diffusion Guidance

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:57.420079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:57.420079Z digest=sha256:6049bf255d10ea997ae8349d5d1f88f45ca43f179ac669f61bd93fce72656955

Observation e77c68c5-45da-461b-aeba-416f1234010b · outbound

This paper cites Muse: Text-To-Image Generation via Masked Generative Transformers.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:57.486525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:57.486525Z digest=sha256:4fa5f01915e56c35baa8c92601d0a4677fdc73a449f2c4d66af0287b9d10c551

Observation b1ce6187-a2e0-41d4-842d-2858f15c4994 · outbound

This paper cites Stemgen: A music generation model that listens,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Stemgen: A music generation model that listens,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.611073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:57.567373Z digest=sha256:8b402d880cbce999f71a51202f94f872fe7d6619070e36dbcb3e031b5898d37a

Observation 1c375ab3-c629-4e01-8ec7-e5d4b76227e4 · outbound

This paper cites Imagebind: One embedding space to bind them all,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Imagebind: One embedding space to bind them all,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.447837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:57.668416Z digest=sha256:55bb20e44394587fe51f2ab0d7c250ce4d6102e3fa96874a3c39f7896a5e73d3

Observation 899ae236-eb3b-43c4-83cd-44ea45f0aa47 · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.318967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:57.772268Z digest=sha256:7de027bdd91d650559c5bd8f1a842ce2c0edb4647af594180212585d0fd7addf

Observation f6054d9e-adc8-4eb1-aca5-4fb921a2b459 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Audio set: An ontology and human-labeled dataset for audio events,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.207305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:57.862288Z digest=sha256:39a7a11bb3c159694654d04f1ccd27c3b02b2f2b55434f939de3ebe58c896a50

Observation a066d024-477c-4e68-9145-818d1990d3b7 · outbound

This paper cites Efficient Training of Audio Transformers with Patchout.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Efficient Training of Audio Transformers with Patchout

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:57.929534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:57.929534Z digest=sha256:f53a4902538b84b0cb619ed366d0a8589943fa7bb32bd86289a7a70ec65fcbab

Observation 9d1e8033-5093-4836-8fc2-544e8816659a · outbound

This paper cites Vggsound: A large- scale audio-visual dataset,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Vggsound: A large- scale audio-visual dataset,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.139077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:58.028545Z digest=sha256:ed88b8830b64a91337c64e90ff297505ccd350e4a61fff36ebafe8f6db533909

Observation 330a7253-e81a-4013-98aa-49fd7a4ff704 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.019857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:58.119664Z digest=sha256:83c51d37fddcfbdac4862f079d24b41e557434822ee7a007f7bb7ca22ced010c

Observation ecda695c-59e4-4d15-b432-4efb0dcb829f · outbound

This paper cites Cnn architectures for large-scale audio classification,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Cnn architectures for large-scale audio classification,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:58.866089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:58.194238Z digest=sha256:3c2c565d6bccbf9f0f17a2d9416a9ca50aa635a340fd41b309fd91b75a78db90

Observation 9efe0af6-4ce8-454b-8bf9-b24274ca4411 · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Panns: Large-scale pretrained audio neural networks for audio pattern recognition,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:58.751950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:09:58.262045Z digest=sha256:e97a7779b0371ec967c76d38c515bd2d706b76cf7310aab50b896ed0496e0422

Pith citing papers

Observation 2678bf73-aeaa-4275-b270-963fb3ea0938 · inbound

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation cites this paper.

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:21:06.817478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T08:20:02.986562Z digest=sha256:eb1256928609c34a10d0237844f489583ab9cb5b829ea7b97469a398619c02b4