Pith. sign in

Paper Citation Record · LEDGER

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet

As of 18 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 1 inbound Pith citation observation for arXiv:2505.16195.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16195 v2

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:09:58.262045Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T08:20:02.986562Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T08:21:06.814521Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation abc09efc-c13d-4be5-8260-8f0b94acdb03 · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet AudioGen: Textually Guided Audio Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.545341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.545341Z digest=sha256:dd19b19217fb37b1ae04456a22240c094efb9a264b5a09ce7094b45a415b8471

Observation 7953642e-47fc-455d-a2b1-c9fc8739fcc3 · outbound

This paper cites Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.619538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.619538Z digest=sha256:3b46034621d287e46594b8316676a1d854cb7af6e06b5c1677cd56caf4cd6939

Observation 10e5f49b-3c03-4f1f-b5df-9f047e5e0bfd · outbound

This paper cites SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.712485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.712485Z digest=sha256:0d45909a6377f1a5488f00f7ce913b20d4b07590ee927be7500ccebb9ec8111c

Observation 66c067d3-bdd2-4644-98ba-2b6fbf9c6141 · outbound

This paper cites SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound Generation.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.810995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.810995Z digest=sha256:79c409e8f80121da7f7b0d0d5e04851727b1114e997736c27e1014ac7c07b278

Observation fe976c38-6e02-4d78-a072-a320e56aaf45 · outbound

This paper cites Stable audio open,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Stable audio open,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:02.801403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:54.911275Z digest=sha256:cab3662808281b3ef8d33c20df5ba103bbecdeaa19cc24e7391bd29605697111

Observation 4ba353cf-9ed8-4907-825a-6f6705aa76b4 · outbound

This paper cites MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.999597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.999597Z digest=sha256:a490cd0cfc7d1c40321999034d80cc8815bee91f37fed750c870474d284688b7

Observation 9b22797a-76b4-46ae-886b-f4261d8e063c · outbound

This paper cites Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:02.647919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:55.139356Z digest=sha256:6e966c5d0afc1a59321304f061cec9671a61ad2915aedd4f75a5c143fdc40f38

Observation 23d57b05-ab77-49f4-88f7-9f99a863cba2 · outbound

This paper cites Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.278233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.278233Z digest=sha256:a4d516e7da04f329443953be328aa9f1e342d8b6d6cad8e492122ecaef10165b

Observation e8818c41-138b-4f53-ab6e-af61e8b28b03 · outbound

This paper cites V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:02.496790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:55.387681Z digest=sha256:17051318c91ecbd79d0fb0bd7e42c6514f304ffca77b69c82a35430ba9bffd59

Observation cfd683ff-f0a3-4a25-baab-cead609353e7 · outbound

This paper cites Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:02.352055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:55.487341Z digest=sha256:04b2a8c111b9e79086d27ea7757e6895f4a1e59e016e7ddee60f08cfd95e133a

Observation 9fef9217-ca06-447b-8cd2-15c042f266c9 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.547205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.547205Z digest=sha256:f0aac916bb5805f64b74e9c30cae5debf9975c7468f266206ef3f6e04a5f499a

Observation 0ddf2160-429f-49b1-b0ad-bcdd0b423e1f · outbound

This paper cites Synchformer: Efficient synchronization from sparse cues,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Synchformer: Efficient synchronization from sparse cues,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:02.207841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:55.606504Z digest=sha256:2de12a6291f4f7e6e296f5ffb99f039a45ffc2eea7597d1852d348befdda3ce1

Observation 8994e6ca-0abb-4e83-b1c4-652cdca8e598 · outbound

This paper cites Adding conditional control to text- to-image diffusion models,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Adding conditional control to text- to-image diffusion models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:02.020964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:55.683396Z digest=sha256:d9326a3d9aa6ff951d776e473e5e06aa732ef5220536dcee1c45d758f7191834

Observation 3cc6f8c6-e56a-4b03-b059-a741ae33fc09 · outbound

This paper cites Uni-controlnet: All-in-one control to text-to-image diffusion models,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Uni-controlnet: All-in-one control to text-to-image diffusion models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:01.857838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:55.753711Z digest=sha256:63aa64ab74c312b256bb6c5377d4f8ae2ca542c5c254cf47c92c7ccb1089a2b1

Observation 110c6b42-6d76-4411-a425-12ee6748e1f7 · outbound

This paper cites Read, Watch and Scream! Sound Generation from Text and Video.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Read, Watch and Scream! Sound Generation from Text and Video

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.831770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.831770Z digest=sha256:d3e345b32ed7f3903121fa783964903c7371d6549128bf2492601da1d8d2a866

Observation 72a83755-a8c4-4920-b509-daf337c91156 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.910310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.910310Z digest=sha256:21bafbbf87e3190d7e133ce603b04e87b8d1560876609e449b27b3922a2faabd

Observation 0dd4b443-d5f0-4fa6-a51c-394681685a84 · outbound

This paper cites Tell What You Hear From What You See -- Video to Audio Generation Through Text.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Tell What You Hear From What You See -- Video to Audio Generation Through Text

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.988743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.988743Z digest=sha256:3fbb77387c256f27837cb39c7e08f558ac955d956e0bdbdeca4b48bbbbc93c94

Observation 6278bfb9-8f1b-4e1a-8085-cc24982c664e · outbound

This paper cites Temporally aligned audio for video with autoregression,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Temporally aligned audio for video with autoregression,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:01.708710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:56.045325Z digest=sha256:8ed89b1882ee5ed2bbfa55cc3fe06bfd6f7be1c8d6fffa73fa568c502281fa38

Observation f2a7234f-eedb-4911-a3eb-28b36912076e · outbound

This paper cites Frieren: Efficient video-to-audio generation network with rectified flow matching,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Frieren: Efficient video-to-audio generation network with rectified flow matching,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:01.422859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:56.119327Z digest=sha256:f122de269c96f8a0dc09815a0c76db604e1132af9c8ce7c467da0d75ed2f220e

Observation e6501fb0-d017-4ff1-9d9d-438c83298075 · outbound

This paper cites Mavil: Masked audio-video learners,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Mavil: Masked audio-video learners,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:01.261287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:56.195962Z digest=sha256:d7a10adcbcfd4bc5550e53986e490c66ec645a3dbb95b1cfa4af2e45473a724c

Observation 3cc382ef-08bf-40c2-a2e7-7cbbe25e8cc4 · outbound

This paper cites Diff-foley: Synchronized video- to-audio synthesis with latent diffusion models,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Diff-foley: Synchronized video- to-audio synthesis with latent diffusion models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:01.089076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:56.264683Z digest=sha256:e065d80d80422586a4256b3e37a47c54f8a7ec11e012aae15ba78c56b4580f26

Observation 6cb4d445-1e69-4e49-ba26-f9ebff2361a3 · outbound

This paper cites FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.308460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.308460Z digest=sha256:d644c01d052c2cc2011d3c7c23486e9b20cd520feae2dca20f7d799f9b01d336

Observation 1153f749-b0bd-4fd4-b5e1-da4c0392e7da · outbound

This paper cites Smooth-foley: Creating continuous sound for video-to-audio generation under semantic guidance,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Smooth-foley: Creating continuous sound for video-to-audio generation under semantic guidance,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.923321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:56.376868Z digest=sha256:07b407c834179ee0facacec57df5b4e99ea72766510c07a80224750baecd4994

Observation 18b0951f-4fe7-4acb-a6e8-fc7fc4972a5f · outbound

This paper cites Learning transferable visual models from natural language supervision,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Learning transferable visual models from natural language supervision,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.765043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:56.424504Z digest=sha256:9f9e65d37dcca4bedb239e308806ddfc974f194024e448b3d2515f74e2c45423

Observation 34ad2f69-31e0-421e-aa52-a94edb43a1cc · outbound

This paper cites Maskgit: Masked generative image transformer,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Maskgit: Masked generative image transformer,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.596551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:56.497923Z digest=sha256:f829ce2a00e889cad9f60d43052782a9d932da6fcad9841baf06372b703f10fe

Observation 389e82fd-b576-42d8-b9dc-c7eb68d08266 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.595415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.595415Z digest=sha256:72ab06be64e6cc0cf07cd2f9bc742281febdf7928b0fadc1533b911c90a41a9b

Observation 1648a738-ef21-43fe-8484-e30a2be13b05 · outbound

This paper cites Music Foundation Model as Generic Booster for Music Downstream Tasks.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Music Foundation Model as Generic Booster for Music Downstream Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.680530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.680530Z digest=sha256:699e2b50c4a39ffde41579fe17748ba872d53b62de7743d75a9282ea12174e6d

Observation 4ba595d1-8f83-41b5-b393-85b15143268d · outbound

This paper cites High Fidelity Neural Audio Compression.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet High Fidelity Neural Audio Compression

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.761212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.761212Z digest=sha256:e925de9e0f5df1dcb017b0e18c3dd2b07a09e7a409a458b3a209dcb6f3e83bf1

Observation bc0b6b04-276c-4456-8638-737ada8b65fd · outbound

This paper cites High- fidelity audio compression with improved rvqgan,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet High- fidelity audio compression with improved rvqgan,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.473120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:56.840072Z digest=sha256:32eadbf0f6b1ff80e50aa1e34abec7f3b214ba2f52b0e1ca5f256878c7788068

Observation 52abba22-8b04-45dd-aacc-3e6d083918f9 · outbound

This paper cites Taming Visually Guided Sound Generation.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Taming Visually Guided Sound Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.889056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.889056Z digest=sha256:93079a834587b86c5d0713d4ee98389b7305de35ee5246946d85b02920ed02d8

Observation ea1a425d-b872-4299-bf93-254e01388d49 · outbound

This paper cites Masked autoencoders are scalable vision learners,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Masked autoencoders are scalable vision learners,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.382535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:56.939962Z digest=sha256:973ca23481a5be72a5c87e11284f00415eabca31d24357317d303d98fd9f03c8

Observation 3e56c4ab-42ed-407a-8092-60bd86077502 · outbound

This paper cites Extending audio masked autoencoders toward audio restoration,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Extending audio masked autoencoders toward audio restoration,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.237928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:57.006139Z digest=sha256:7f2734ed34bcccb61a34c118be3a8614ba7c733beceff1584e9eb1237d704e22

Observation 87d13178-79e3-43d7-b824-dede84d37966 · outbound

This paper cites Mage: Masked generative encoder to unify representation learning and image synthesis,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Mage: Masked generative encoder to unify representation learning and image synthesis,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:00.021600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:57.058562Z digest=sha256:c8c0ac94d767aeec63b641c547fd13d3a76e743c9424c3bb0b5a61251f223e51

Observation ead1e612-6841-4a2b-817b-ab9faebc6d7d · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:57.138047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:57.138047Z digest=sha256:389b8ca56c99a1a1b4f21ca78af32ca3a615b3427bcd716d8f96435b7b183123

Observation 158aff55-058b-43f9-a4a9-4fc26dd9b46b · outbound

This paper cites PIXART-{\delta}: Fast and Controllable Image Generation with Latent Consistency Models.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet PIXART-{\delta}: Fast and Controllable Image Generation with Latent Consistency Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:57.179030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:57.179030Z digest=sha256:e9a1d80be3750c300321784604b729795544557b6cf1da601d39f152cf520b8a

Observation c4db9e3c-35ea-45c6-aa2f-f3f918488ffb · outbound

This paper cites COCOLA: Coherence-Oriented Contrastive Learning of Musical Audio Representations.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet COCOLA: Coherence-Oriented Contrastive Learning of Musical Audio Representations

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:09:58.450395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:57.248961Z digest=sha256:8217366edc4bc60d41e3921194f33547f6454542fb6fec22a4450849f79c9304

Observation d64c8f29-6bdc-49fe-a011-3afb21821ef2 · outbound

This paper cites Editing music with melody and text: Using controlnet for diffusion transformer,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Editing music with melody and text: Using controlnet for diffusion transformer,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.778127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:57.354123Z digest=sha256:253129c56982c858ff8d6d7b8d0acefc0efdb3e2f958e41af75b9f3b85c9d08d

Observation 5e669fe6-ece2-4d8e-88cf-6ced1feec5be · outbound

This paper cites Classifier-Free Diffusion Guidance.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Classifier-Free Diffusion Guidance

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:57.420079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:57.420079Z digest=sha256:a63505c715a59b0dc34e678d0d433d379a25d1b72b417529699d229e27a3a499

Observation e77c68c5-45da-461b-aeba-416f1234010b · outbound

This paper cites Muse: Text-To-Image Generation via Masked Generative Transformers.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:57.486525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:57.486525Z digest=sha256:c423336f7025f056e3550ddf741bccd00d0edeb2169b86fb68cd27af2ce83caa

Observation b1ce6187-a2e0-41d4-842d-2858f15c4994 · outbound

This paper cites Stemgen: A music generation model that listens,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Stemgen: A music generation model that listens,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.611073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:57.567373Z digest=sha256:6f6485e03c820615ab01f9e9294306608bb7776b95b5b587da58390ff358cf44

Observation 1c375ab3-c629-4e01-8ec7-e5d4b76227e4 · outbound

This paper cites Imagebind: One embedding space to bind them all,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Imagebind: One embedding space to bind them all,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.447837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:57.668416Z digest=sha256:ab8234c545ed8829effd5b13666525a1735b132c63b87b3ad3695c787d7ed78d

Observation 899ae236-eb3b-43c4-83cd-44ea45f0aa47 · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.318967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:57.772268Z digest=sha256:63421758a2f3b767f38c36b59f1be0353ffb86c13f0bcd687fe66cecf0ad1feb

Observation f6054d9e-adc8-4eb1-aca5-4fb921a2b459 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Audio set: An ontology and human-labeled dataset for audio events,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.207305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:57.862288Z digest=sha256:e8b5edb9ac4c72aef901ffad5a60ed7d4738fd27c46381c338bc00648316565c

Observation a066d024-477c-4e68-9145-818d1990d3b7 · outbound

This paper cites Efficient Training of Audio Transformers with Patchout.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Efficient Training of Audio Transformers with Patchout

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:57.929534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:57.929534Z digest=sha256:11ed5b80d47b218b71d7c5b979bd4817be67c32b4b51a6215a5d02f84c285358

Observation 9d1e8033-5093-4836-8fc2-544e8816659a · outbound

This paper cites Vggsound: A large- scale audio-visual dataset,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Vggsound: A large- scale audio-visual dataset,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.139077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:58.028545Z digest=sha256:1434bb178aab8362bf61e743cda5b740040eb9dcfab1ae0210b76ab1f7dfc6a4

Observation 330a7253-e81a-4013-98aa-49fd7a4ff704 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.019857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:58.119664Z digest=sha256:3fc9086562ea4e68d2428461b15c9502053ed1b7f22a86947514a7be4be4dddb

Observation ecda695c-59e4-4d15-b432-4efb0dcb829f · outbound

This paper cites Cnn architectures for large-scale audio classification,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Cnn architectures for large-scale audio classification,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:58.866089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:58.194238Z digest=sha256:67f3ef91a989136ca403604a52f8e0ff010546af0d043a21397cfa0d74e7adec

Observation 9efe0af6-4ce8-454b-8bf9-b24274ca4411 · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition,.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Panns: Large-scale pretrained audio neural networks for audio pattern recognition,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:58.751950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:09:58.262045Z digest=sha256:43ef2102ecc775c150191521d6e202f9facf6390009fc803f4dd09c78d77a785

Pith citing papers

Observation 2678bf73-aeaa-4275-b270-963fb3ea0938 · inbound

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation cites this paper.

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:21:06.817478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T08:20:02.986562Z digest=sha256:b2e06d4b8f275bfa4c13e219b91fe743a6ae5cb9e1bb385a2156aaa9ac7f33ac