Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2412.15322.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T14:26:46.641742Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-22T23:07:14.501348Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 21929990-4d8e-41c6-992e-2b69a4e41bde · inbound
Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91795adc-1b3d-410e-8887-008a4e0d6bb9 · inbound
Wan: Open and Advanced Large-Scale Video Generative Models MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4ba353cf-9ed8-4907-825a-6f6705aa76b4 · inbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2309e642-f146-464a-a083-36fb2b903a66 · inbound
Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30f33d15-5961-462a-a3c1-2596da605e46 · inbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14e806fc-87e0-48fa-9c31-79ec9b6b96c0 · inbound
Sounding that Object: Interactive Object-Aware Image to Audio Generation MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 713fd661-c934-40de-b955-c4023f195efc · inbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7244bc48-d4f4-41f9-9ec9-caa3a12af7b1 · inbound
FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 19f406e3-594d-419b-8900-c45e54a35f16 · inbound
TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 85d98edd-c600-454c-a9dc-983ee4baa906 · inbound
SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cfa0b25-621d-4ebf-a247-cdecb6287abc · inbound
KVAE: Family of Tokenizers for Multimodal Generative Models MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.