Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T16:26:00.285585Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 7 inbound Pith citation observations for arXiv:2512.13677.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T16:26:00.285585Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-05T12:34:41.758238Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-05T12:41:04.375376Z
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b83231be-d556-495e-bca1-e9f384e77472 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2818ee41-0988-4952-81f6-d2edff52087b · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 935e2106-9a9f-41df-b2d6-650d550ed0ae · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Mmaudio: Taming multimodal joint training for high-quality video-to-audio synthesis
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d57b0d29-2d33-4b37-8203-aba4a80749cc · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing UniMax: Fairer and more Effective Language Sampling for Large-Scale Multilingual Pretraining
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8853a92-b483-46c6-86d2-74c9621b13c9 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Out of time: automated lip sync in the wild
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aff03e9e-1449-4dfd-a312-d23a650ece39 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Wan2.5.https://wan.video/, 2025
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfa242a5-ea45-4b8d-8504-44c2213035b3 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Veo3.https://deepmind.google/models/veo/, 2025
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23c081a3-7633-48cd-be2c-102805c73517 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Clotho: An audio captioning dataset
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83a2e068-c7a7-4f7b-ba91-9281acb8b475 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Scaling rectified flow transformers for high-resolution image synthesis
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10a0ad3d-c642-4bd7-b316-333a7b4bf0a9 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1e3d810-2745-4a83-8ad8-805a95a82067 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Wan-S2V: Audio-Driven Cinematic Video Generation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dedc90e-dc1b-43f1-b792-0ce29ca42ea7 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ae5821c-dec4-4de1-bb54-c93df909c2ee · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing ACE-Step: A Step Towards Music Generation Foundation Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16755c82-6b7f-4480-9a83-682402e864d4 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96a44fdf-00a6-40bc-8e5d-41ecf84ffd9c · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Classifier-Free Diffusion Guidance
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cee6463b-6430-4500-a5af-9c211722155e · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3d0283c-3482-443b-8a3a-a6c00bac2951 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Musiq: Multi-scale image quality transformer
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation accb6232-f55c-403a-b743-a096b1edbb8d · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Audiocaps: Generating captions for audios in the wild
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 834b6cf6-fc56-4714-9886-927c91e9f777 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Panns: Large-scale pretrained audio neural networks for audio pattern recognition.TASLPRO, 2020
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11c20f79-fd10-4c1e-b34b-bba0579a3fc9 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing HunyuanVideo: A Systematic Framework For Large Video Generative Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d88f42e7-2290-47d5-b122-5dc5d3a2c19f · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Efficient Training of Audio Transformers with Patchout
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fd660bf-dc2f-47c7-a7db-6010fe912ec8 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bf3a7e8-b71c-41e3-8bda-8ebd4f354554 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Flow Matching for Generative Modeling
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 285ce9c0-83cd-4b94-97cd-2a3c479351a1 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Javisdit: Joint audio-video diffusion transformer with hierarchical spatio-temporal prior synchronization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d5f3efb-364d-4b6e-b5f0-50642d1d0df1 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79ff5cde-cdb9-4555-b3dc-0d72394d7751 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfa016d6-3512-4940-bc74-a4df94e98d23 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models.NeurIPS, 2023
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6557d60-88be-4a3c-a4e0-f3073bef3e71 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 474e3a89-80a7-420b-a03d-8e6aa4100e9d · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Sora2.https://openai.com/zh-Hans-CN/index/sora-2/, 2025
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eacb55e3-a723-4bd5-8f44-07b12faaa521 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Scalable diffusion models with transformers
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89690cbb-5a30-4057-b2aa-d2b740fcd615 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Robust speech recognition via large-scale weak supervision
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1a0b9b9-30a5-4662-9df1-5e55423b0535 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing U-net: Convolutional networks for biomedical image segmentation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58fe299d-b20e-4aeb-888e-d8472845d91f · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44ecad2f-0840-4635-9386-8dd77c7ae7d4 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da686c8d-b660-4472-81d4-5c950e054d7b · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing DINOv3
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1998f3a2-5d6f-4c6b-943a-f96ca72f0e5b · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Raft: Recurrent all-pairs field transforms for optical flow
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62d3aea7-532b-462f-a0df-bdb4706434f2 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb05a9c8-7e16-421e-bb07-23ee8354420b · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Wan: Open and Advanced Large-Scale Video Generative Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57ea88af-15b7-443b-b837-fa61829a4c83 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing UniVerse-1: Unified Audio-Video Generation via Stitching of Experts
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf7f4057-b622-4a49-ab4a-43c4f13fc4c2 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Kling-foley: Multimodal diffusion transformer for high-quality video-to-audio generation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caab88d7-d3bc-457d-ba24-7955d5933cb6 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Av-dit: Taming image diffusion transformers for efficient joint audio and video generation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5db16c02-bee4-4b43-8764-32f82764cc54 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fbe48f2-9e25-4364-a75e-8b7cebdb85f4 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Fantasytalking: Realistic talking portrait generation via coherent motion synthesis
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ffc17f1-87cf-46b9-890f-4755ca5b67c0 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7b54d83-a12a-48b5-b8e2-e59d7e19407d · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b018707-b816-462a-aac6-27b64fe7b027 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 102b32df-884b-43d1-b960-e2d78ecd357d · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Qwen2.5 Technical Report
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fec9ff2-183b-46a0-8969-c4bbb43d83ab · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Maniqa: Multi-dimension attention network for no-reference image quality assessment
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3a5c374-d6a5-4f20-8374-6f3f60c15a09 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8945482-6b3c-4c05-8c01-c345ae36d1ed · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a7707f2-2d37-44f4-9fee-803663185f88 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Uniavgen: Unified audio and video generation with asymmetric cross-modal interactions.arXiv preprint arXiv:2511.03334, 2025
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69a06d9f-05bd-435d-9b0d-efe3bfd608a5 · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Waver: Wave Your Way to Lifelike Video Generation
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd5b86f8-7e4b-42c5-aa90-1aaa05de5c4d · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83f1c4a2-278e-4948-ad99-b96b99a7765a · outbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Uniform: A unified diffusion transformer for audio-video generation.arXiv e-prints, 2025
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 039d90a1-8358-47fe-a869-deab40f14716 · inbound
Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 106168a0-6b39-4c67-a652-01b39ba9afa4 · inbound
OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c01f1bf5-3a78-4e08-b828-3d6f4c924adb · inbound
OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 81c0ff12-b1ad-4d20-94d9-b90843758d4a · inbound
OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 19ec8ee7-9b08-4f8c-8865-d6376b661e51 · inbound
Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5dc8bbe9-ae49-4ffe-8850-c4b2235f4e8d · inbound
SyncDPO: Enhancing Temporal Synchronization in Video-Audio Joint Generation via Preference Learning JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 41339ad0-1024-4195-90ff-7285d7d83b7c · inbound
From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing
Reference 133
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.