Pith. sign in

Paper Citation Record · LEDGER

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation

As of 13 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2606.08393.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.08393 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T18:20:24.588765Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact7
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8f266217-a0c5-4a5d-94f7-11b9ca228fac · outbound

This paper cites Diff-Foley: Synchronized video-to-audio synthesis with latent diffusion models,.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation Diff-Foley: Synchronized video-to-audio synthesis with latent diffusion models,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T18:20:24.588765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:e799bb788a73ef35b83cd80a0d0b927cc552e0282551329ca90bc9ffc82112ad

Observation ba41b621-6845-44f8-a357-69f53fb8fe9a · outbound

This paper cites FoleyCrafter: Bring silent videos to life with lifelike and synchronized sounds,.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation FoleyCrafter: Bring silent videos to life with lifelike and synchronized sounds,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T18:20:24.588765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:10ddc42ffb6b51cda673e24fb6e6524716a98486ebadd44c7e028cc15d6ab6a6

Observation 4316867a-93e6-47cf-a5a5-212acc1d0e0a · outbound

This paper cites Tell what you hear from what you see - video to audio generation through text,.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation Tell what you hear from what you see - video to audio generation through text,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-27T18:20:24.588765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:0633ab075427ec930ee963f9e79bc2d062a99966454f88f28b4b4c905760a89c

Observation 56bdf54d-94cb-49d3-ae2d-c3e708ed681e · outbound

This paper cites MMAudio: Taming multimodal joint training for high- quality video-to-audio synthesis,.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation MMAudio: Taming multimodal joint training for high- quality video-to-audio synthesis,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T18:20:24.588765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:c3a5fd40839c5f838f6f1a580f0428c5386be47c0cedc313da86155768573c6e

Observation 4dc22907-9f55-4542-a95b-c71793a589ae · outbound

This paper cites PrismAudio: Decomposed chain-of-thought and multi- dimensional rewards for video-to-audio generation,.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation PrismAudio: Decomposed chain-of-thought and multi- dimensional rewards for video-to-audio generation,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-27T18:20:24.588765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:3c4832ffd7c5c3737e9684dc9526bba360c8c02d42fb3ffe8bee1fed52384cdc

Observation 67f577bb-e257-447d-a2c9-a1ec3d531ab6 · outbound

This paper cites Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-02T23:17:29.636908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:40bbc1bf3eca1c3e54cb19abd0a463e521100aa819d0a910373c6e3542c06ebc

Observation 3200a5c9-a463-450c-934c-f40b3c9b0414 · outbound

This paper cites AC- Foley: Reference-audio-guided video-to-audio synthesis with acoustic transfer,.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation AC- Foley: Reference-audio-guided video-to-audio synthesis with acoustic transfer,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T18:20:24.588765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:15090f9f016551b6223b8c4e44cca762adffaaa29d962efa61987effdfc4c1b4

Observation 2bc0f2e9-60a5-4d7d-babf-52b3b2f3a8c5 · outbound

This paper cites Aligning Text-to-Image Models using Human Feedback.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation Aligning Text-to-Image Models using Human Feedback

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-02T23:17:29.633974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:2e3508a5d185fb6e32f30d1e43523af7a9c8caf291eb4f078dc50124d6bf4312

Observation fb771b12-94b5-46ca-b2d4-66dc7e726e22 · outbound

This paper cites Test-time alignment of diffusion mod- els without reward over-optimization,.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation Test-time alignment of diffusion mod- els without reward over-optimization,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T18:20:24.588765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:29249f2afcfb58ca162e3474256148a1f69901aec5f873f392c8058dc69f7577

Observation 42a071a5-0ffc-4efa-b5db-1d26b50709b2 · outbound

This paper cites Symbolic music generation with non-differentiable rule guided diffusion,.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation Symbolic music generation with non-differentiable rule guided diffusion,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T18:20:24.588765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:16f9de8776d69c8e0d72453ae100ba6f8d1c4502f1229c0250162a40ccb01119

Observation beabb127-e16b-43b9-9517-a91b4dc5089e · outbound

This paper cites Inference-time text- to-video alignment with diffusion latent beam search,.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation Inference-time text- to-video alignment with diffusion latent beam search,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T18:20:24.588765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:1bbd1cadf4c0c966ac06e3ca2bd2cfc4fd1b35983adb785d2112118c86283999

Observation 809d1abe-026f-4810-a99d-d60f97da9c0d · outbound

This paper cites SCORE: Scaling au- dio generation using standardized composite rewards,.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation SCORE: Scaling au- dio generation using standardized composite rewards,

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:17:29.639630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:be8b74139fee53d745c197b0f641ebc562cd4e17af8866be3af1aebec2768c21

Observation 1bce797e-4b2b-487a-9461-f33b56f5bc10 · outbound

This paper cites Sequential monte carlo samplers,.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation Sequential monte carlo samplers,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T18:20:24.588765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:561f74440ea64a65641b8b0b12f56621883eca6c7422c7123252836d61b3c601

Observation 23ed9224-ef5b-4807-b16a-54cfbb489cd7 · outbound

This paper cites Synchformer: Efficient synchronization from sparse cues,.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation Synchformer: Efficient synchronization from sparse cues,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T18:20:24.588765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:4bac452661d6901c4b2e213f908204883cba9c8aa778744ed87fe9267190e652

Observation 0f42320b-bb3f-405b-b9c6-7f95e944b621 · outbound

This paper cites Dynamic chunking for end-to- end hierarchical sequence modeling,.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation Dynamic chunking for end-to- end hierarchical sequence modeling,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T18:20:24.588765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:f9796afa2f6bb6a444e8bf395d37fa1e01c68bfb65487c246dbe6dbe1d184acb

Observation 226fdad3-c376-4638-a542-ff5229e94dde · outbound

This paper cites Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-02T23:17:29.628265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:40082570a22f6d0688e480dc3aac191c956dc88298e98d39f2ef06c01616a681

Observation 0a84f4e7-67f5-41a2-b210-7d7bef3eba38 · outbound

This paper cites A general framework for inference-time scaling and steering of diffusion models,.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation A general framework for inference-time scaling and steering of diffusion models,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T18:20:24.588765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:0ecc5c64d5c9e4a48af52e090e20d32a6b606bafc31d842eef5698d93e5b06ba

Observation 2eb7d162-4dd4-4697-ac0d-2d329f1cf429 · outbound

This paper cites Diffusion models beat GANs on image synthesis,.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation Diffusion models beat GANs on image synthesis,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T18:20:24.588765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:643ce5c41af12d03738497657f728ff8950c08065ada43cc13a24b35073770a9

Observation f3d1149e-21b4-4d86-9cb6-2f45421c339d · outbound

This paper cites Dynamic Search for Inference-Time Alignment in Diffusion Models.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation Dynamic Search for Inference-Time Alignment in Diffusion Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:17:29.631102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:a8804704d9f4d310434e3bb46ae5f2e7ad1f507021218022ab5cc1f24729d940

Observation 5d23bec1-4f79-4fed-bf1e-2ea27c097e34 · outbound

This paper cites Scaling Image and Video Generation via Test-Time Evolutionary Search.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation Scaling Image and Video Generation via Test-Time Evolutionary Search

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:17:29.625536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:7d4469f7a0f2ba033b5f88cc1a5e30eefb6c3b8fb03582acaffb1e3281d189fd

Observation c0344d0a-46e3-4fd6-a1ed-d3175fdc15f6 · outbound

This paper cites Flow matching for generative modeling,.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation Flow matching for generative modeling,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-27T18:20:24.588765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:6c39d6985eb5941342a5d72e3178238192a21088b40dbc273631dbebf2fd441a

Observation f7a47ae1-002e-4348-940c-3d364e175271 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation Score-Based Generative Modeling through Stochastic Differential Equations

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-02T23:17:29.622788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:ba018df7c90f1cb874e2c28f5e72ff2c803eda420366267adc5c97a3f9d7d553

Observation 4932171a-6b38-429d-bba3-6be020466039 · outbound

This paper cites HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-27T18:20:24.588765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:7948b246cd3c581cea31811a4983dc17d565448ba1e1448b8c44610d61bb673e

Observation 2203ed8e-e4b3-464d-9347-fcc5c717c751 · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-27T18:20:24.588765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:9b44d020d43c034af195df3223918ab7308775b189697fea6eea18fd2908adb1

Observation a6b129f5-c755-4ca6-8fb6-e0f418bf6369 · outbound

This paper cites ImageBind: One embedding space to bind them all,.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation ImageBind: One embedding space to bind them all,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T18:20:24.588765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:89cf17c548c95cc45182efd8c809d7eadb1c9a694037a4f24d6111b266661fca

Observation c9edc96f-9bc9-4667-b44c-c34bf8ffd557 · outbound

This paper cites Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-02T23:17:29.620375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:3f6f159c271e00b41a8eebc33b229f4fee01db7403a37987b943fd4b33ccfa25

Observation f5cbcf12-c892-48a7-8048-019c15de33a8 · outbound

This paper cites Improved particle filter for nonlinear problems,.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation Improved particle filter for nonlinear problems,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-27T18:20:24.588765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:dee6246ae6670ed79ea8ec59a8e2543e5590ddb0ee2cc0fd467f2d21ff525d33

Observation a0b15939-bf1a-42d6-82df-cbd363e79325 · outbound

This paper cites Negative association, ordering and convergence of resampling methods,.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation Negative association, ordering and convergence of resampling methods,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-27T18:20:24.588765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:c6ce9b3855ab5d76a8c5e015cb4f1ba04e613627ad2fe8f7becd705b55eabcc5

Observation 64fed236-bce3-4def-8f20-0040893abc33 · outbound

This paper cites VGGSound: A large-scale audio-visual dataset,.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation VGGSound: A large-scale audio-visual dataset,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-27T18:20:24.588765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:a7a01023be256ede1c8ba77096792ff6c3a38631df97270e2cc0ae1b741c1f32

Observation 60f57178-b473-4595-952c-74d889ee6dae · outbound

This paper cites AGA V-Rater: Adapting large multimodal model for AI-generated audio-visual quality assess- ment,.

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation AGA V-Rater: Adapting large multimodal model for AI-generated audio-visual quality assess- ment,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-27T18:20:24.588765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T18:20:24.588765Z digest=sha256:053d7b0672d26e88eb676b87e7aa07569231e39a33d3fd7606c9837013c0d1e6

Pith citing papers

No inbound Pith citation observations are available.