Pith. sign in

Paper Citation Record · LEDGER

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing

As of 5 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 7 inbound Pith citation observations for arXiv:2512.13677.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.13677 v2

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T16:26:00.285585Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-05T12:34:41.758238Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-05T12:41:04.375376Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b83231be-d556-495e-bca1-e9f384e77472 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:53.994887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:53.994887Z digest=sha256:de1383f4dd90123310dd971de33d0152c006c40a790959184c1fef5218aa9751

Observation 2818ee41-0988-4952-81f6-d2edff52087b · outbound

This paper cites HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:54.145826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:54.145826Z digest=sha256:63906c55ca8d62a0addfd6d83e5afe3ede1d430cc59f598988b95fe5c4dfc013

Observation 935e2106-9a9f-41df-b2d6-650d550ed0ae · outbound

This paper cites Mmaudio: Taming multimodal joint training for high-quality video-to-audio synthesis.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Mmaudio: Taming multimodal joint training for high-quality video-to-audio synthesis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:54.310475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:54.310475Z digest=sha256:3d1a381ed9078e25dabb36823a93a54e6ba8bf9570bf34bbec9772ba6fc412f7

Observation d57b0d29-2d33-4b37-8203-aba4a80749cc · outbound

This paper cites UniMax: Fairer and more Effective Language Sampling for Large-Scale Multilingual Pretraining.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing UniMax: Fairer and more Effective Language Sampling for Large-Scale Multilingual Pretraining

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:54.531863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:54.531863Z digest=sha256:f6caebc54f09277a3c719cc8198a0ebc0e2a77a554981e2532c1f07a4de6a6c7

Observation a8853a92-b483-46c6-86d2-74c9621b13c9 · outbound

This paper cites Out of time: automated lip sync in the wild.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Out of time: automated lip sync in the wild

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:54.730459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:54.730459Z digest=sha256:2780b32b03da0e8f50c76e7e4005fd24975e81a1430b1d26a5077fe620aade26

Observation aff03e9e-1449-4dfd-a312-d23a650ece39 · outbound

This paper cites Wan2.5.https://wan.video/, 2025.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Wan2.5.https://wan.video/, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:54.827406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:54.827406Z digest=sha256:3e64fb9b90d85dd626d3cadeab49f5607b3891e17fe2a8ecab8e333ae08dd79b

Observation cfa242a5-ea45-4b8d-8504-44c2213035b3 · outbound

This paper cites Veo3.https://deepmind.google/models/veo/, 2025.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Veo3.https://deepmind.google/models/veo/, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:54.934432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:54.934432Z digest=sha256:111e7ea7d1c300893131b55b76bd34ad4bee5a189a120b5f72f48778f3e35d8f

Observation 23c081a3-7633-48cd-be2c-102805c73517 · outbound

This paper cites Clotho: An audio captioning dataset.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Clotho: An audio captioning dataset

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:55.065359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:55.065359Z digest=sha256:9cea01e5723d9e2d18f4244a23ebd8e2ee43df5e1ee1bee3f263e4eac93bfa86

Observation 83a2e068-c7a7-4f7b-ba91-9281acb8b475 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Scaling rectified flow transformers for high-resolution image synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:55.273666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:55.273666Z digest=sha256:f8aa4522a208b140b8df31bc8154d5c59ed0a87f609c34e2250d108e67cfda8b

Observation 10a0ad3d-c642-4bd7-b316-333a7b4bf0a9 · outbound

This paper cites OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:55.436179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:55.436179Z digest=sha256:c780d66f1b6a7f9fa972f8c551e93fe41c9f37ec769e3d6f982ac391e8499b66

Observation f1e3d810-2745-4a83-8ad8-805a95a82067 · outbound

This paper cites Wan-S2V: Audio-Driven Cinematic Video Generation.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:55.546710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:55.546710Z digest=sha256:6492602fc273830f3e8a418752863d17fd6c7db92f0e31f49d4e4877f7429f95

Observation 0dedc90e-dc1b-43f1-b792-0ce29ca42ea7 · outbound

This paper cites Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:55.712311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:55.712311Z digest=sha256:6f3641fe0b7e9e06dc68de397850dc5755b609fc8dace1175ed4baacaa0e3079

Observation 7ae5821c-dec4-4de1-bb54-c93df909c2ee · outbound

This paper cites ACE-Step: A Step Towards Music Generation Foundation Model.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing ACE-Step: A Step Towards Music Generation Foundation Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:55.887395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:55.887395Z digest=sha256:8316c45f667a19c853b32178b44e5b96079e5dbbfc3d3b9899987fa9bd20c378

Observation 16755c82-6b7f-4480-9a83-682402e864d4 · outbound

This paper cites AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:56.032167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:56.032167Z digest=sha256:074f0c5c3ca789aee565040abd5598e3e92dd2955801b7177da77d257c898660

Observation 96a44fdf-00a6-40bc-8e5d-41ecf84ffd9c · outbound

This paper cites Classifier-Free Diffusion Guidance.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Classifier-Free Diffusion Guidance

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:56.183618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:56.183618Z digest=sha256:1e6b5966e4931ad8fe9c0e4b42f2476c06ce6952a6b057f3cd4a840f1ae86420

Observation cee6463b-6430-4500-a5af-9c211722155e · outbound

This paper cites A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:56.290355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:56.290355Z digest=sha256:850aa540ef0ed353589af5de6c00802b4a9656caa3c8db67f1ad087a1827a682

Observation d3d0283c-3482-443b-8a3a-a6c00bac2951 · outbound

This paper cites Musiq: Multi-scale image quality transformer.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Musiq: Multi-scale image quality transformer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:56.399787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:56.399787Z digest=sha256:a550ee961af42e0b44dd99347d7894a151d8ecd4ed54f151113c0bd09feb7639

Observation accb6232-f55c-403a-b743-a096b1edbb8d · outbound

This paper cites Audiocaps: Generating captions for audios in the wild.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Audiocaps: Generating captions for audios in the wild

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:56.574327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:56.574327Z digest=sha256:847ffe026c5689a63d895038e74dae4b4a3fe7abe7b8dc6a2424cd078cb926b5

Observation 834b6cf6-fc56-4714-9886-927c91e9f777 · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition.TASLPRO, 2020.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Panns: Large-scale pretrained audio neural networks for audio pattern recognition.TASLPRO, 2020

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:56.653546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:56.653546Z digest=sha256:a49d42ffd4cc793ce03170a0040cd764ddc83b1d8c8298c93b2483ee36b91ae7

Observation 11c20f79-fd10-4c1e-b34b-bba0579a3fc9 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:56.736283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:56.736283Z digest=sha256:3a4a39101a664a39093ad8d77ea0e6c3baf87f8e4b821d22a3107e2ce6c11645

Observation d88f42e7-2290-47d5-b122-5dc5d3a2c19f · outbound

This paper cites Efficient Training of Audio Transformers with Patchout.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Efficient Training of Audio Transformers with Patchout

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:56.830276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:56.830276Z digest=sha256:78d252540c7341f8e0fe2c0d0ff7f1bd40cf4a89d4ec66cc3f23e9dbfaa1a21e

Observation 4fd660bf-dc2f-47c7-a7db-6010fe912ec8 · outbound

This paper cites OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:56.942377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:56.942377Z digest=sha256:797867e199a9cdef2acfbe6ac88014c0e74f35af4825569bcb9aeca7b9a647f1

Observation 3bf3a7e8-b71c-41e3-8bda-8ebd4f354554 · outbound

This paper cites Flow Matching for Generative Modeling.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Flow Matching for Generative Modeling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:57.046712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:57.046712Z digest=sha256:f91cde933743ac8d13ec82ae8f58d66912a30e2ef468b5a110a90ff4867d41e0

Observation 285ce9c0-83cd-4b94-97cd-2a3c479351a1 · outbound

This paper cites Javisdit: Joint audio-video diffusion transformer with hierarchical spatio-temporal prior synchronization.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Javisdit: Joint audio-video diffusion transformer with hierarchical spatio-temporal prior synchronization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:57.115445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:57.115445Z digest=sha256:8b698fea14bf5ba4eab981d06cf1a9dd6fabc4dd347d03cbb979ce8d4fe27ea1

Observation 6d5f3efb-364d-4b6e-b5f0-50642d1d0df1 · outbound

This paper cites Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:57.224612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:57.224612Z digest=sha256:c59c6a45e7f7d2f4ce788c914c77c64a8c1a7cdb1c4e7f9da3723d7932812243

Observation 79ff5cde-cdb9-4555-b3dc-0d72394d7751 · outbound

This paper cites Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:57.336256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:57.336256Z digest=sha256:cefb0ad381b3fdb15fbeff5d7563dc2b9417854ca8bf5e049f59a7fade7cc282

Observation dfa016d6-3512-4940-bc74-a4df94e98d23 · outbound

This paper cites Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models.NeurIPS, 2023.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models.NeurIPS, 2023

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:57.443780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:57.443780Z digest=sha256:ec50fc1907931072524673a3176dfd831c1b68bdc97188e688fd827798c2105f

Observation c6557d60-88be-4a3c-a4e0-f3073bef3e71 · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:57.571217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:57.571217Z digest=sha256:515cbedb78b13196d6d1aba171d016c39dc90406f3bbafbc8c4d0973f6c19b2a

Observation 474e3a89-80a7-420b-a03d-8e6aa4100e9d · outbound

This paper cites Sora2.https://openai.com/zh-Hans-CN/index/sora-2/, 2025.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Sora2.https://openai.com/zh-Hans-CN/index/sora-2/, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:57.752438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:57.752438Z digest=sha256:0f9373374c94e4c2c296498ff828f11b2486523714053f8187ceecaf7624b38e

Observation eacb55e3-a723-4bd5-8f44-07b12faaa521 · outbound

This paper cites Scalable diffusion models with transformers.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Scalable diffusion models with transformers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:57.834730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:57.834730Z digest=sha256:87eee8150c13c0f01feda012ed18ba7e814b65157d715c877f36725e3772b154

Observation 89690cbb-5a30-4057-b2aa-d2b740fcd615 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Robust speech recognition via large-scale weak supervision

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:57.897622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:57.897622Z digest=sha256:d313c2e4e3c010e8432c1af85302c7838112f79459b3cf935aa2a79740bb9a57

Observation d1a0b9b9-30a5-4662-9df1-5e55423b0535 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing U-net: Convolutional networks for biomedical image segmentation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:58.028988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:58.028988Z digest=sha256:8ad1623f0cb5f8af503d4c728f6839d93a9e8bb5e31db6404ceaf7b67c0ee742

Observation 58fe299d-b20e-4aeb-888e-d8472845d91f · outbound

This paper cites Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:58.145178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:58.145178Z digest=sha256:2a13585953f5547b47371bd6a4fb84749925d02abe7b3e8f788ec64d438ce7cc

Observation 44ecad2f-0840-4635-9386-8dd77c7ae7d4 · outbound

This paper cites HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:58.270794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:58.270794Z digest=sha256:5f8ec03a00dbfeb59cc3c3c9b0345f31f97214e1a9b9d2fbd80b702f54bcd4c6

Observation da686c8d-b660-4472-81d4-5c950e054d7b · outbound

This paper cites DINOv3.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing DINOv3

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:58.357377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:58.357377Z digest=sha256:28d566abb5030e924064c32fbfaddf8d033606c2125fe23147a4f0fb2f492d14

Observation 1998f3a2-5d6f-4c6b-943a-f96ca72f0e5b · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Raft: Recurrent all-pairs field transforms for optical flow

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:58.480419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:58.480419Z digest=sha256:8ac37d74bcc18efa8198a2348a42a5a7b34ab62ef83ac1ef751a68989c0b4cde

Observation 62d3aea7-532b-462f-a0df-bdb4706434f2 · outbound

This paper cites Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:58.606661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:58.606661Z digest=sha256:147c8c7c07de478bd138092311e8e91b47a39a697b9e991914df148982df63dc

Observation cb05a9c8-7e16-421e-bb07-23ee8354420b · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Wan: Open and Advanced Large-Scale Video Generative Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:58.675273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:58.675273Z digest=sha256:de3f653b2e122bdb669ba87368a68fe69dbdda93c20c8c7763bafbbbda448115

Observation 57ea88af-15b7-443b-b837-fa61829a4c83 · outbound

This paper cites UniVerse-1: Unified Audio-Video Generation via Stitching of Experts.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing UniVerse-1: Unified Audio-Video Generation via Stitching of Experts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:58.726931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:58.726931Z digest=sha256:de684a759929db8f86dd607323987226444df6abca3ca7af2d928033ccfffc04

Observation bf7f4057-b622-4a49-ab4a-43c4f13fc4c2 · outbound

This paper cites Kling-foley: Multimodal diffusion transformer for high-quality video-to-audio generation.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Kling-foley: Multimodal diffusion transformer for high-quality video-to-audio generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:58.763471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:58.763471Z digest=sha256:42529d617c6892f3c182f0b8b23960dc2faaf204c1312aa719eeb71a8923afd7

Observation caab88d7-d3bc-457d-ba24-7955d5933cb6 · outbound

This paper cites Av-dit: Taming image diffusion transformers for efficient joint audio and video generation.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Av-dit: Taming image diffusion transformers for efficient joint audio and video generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:58.776199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:58.776199Z digest=sha256:cdc555b6c3d98203343906478ed16e68a7d7e36e36d27649732b26414230d358

Observation 5db16c02-bee4-4b43-8764-32f82764cc54 · outbound

This paper cites AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:58.872482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:58.872482Z digest=sha256:d08d79d91867d2d4b1d2b1a86ffd34df66202b8af072edd9d1f47d53b9a9d411

Observation 6fbe48f2-9e25-4364-a75e-8b7cebdb85f4 · outbound

This paper cites Fantasytalking: Realistic talking portrait generation via coherent motion synthesis.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Fantasytalking: Realistic talking portrait generation via coherent motion synthesis

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:58.948739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:58.948739Z digest=sha256:27fadade9c67381b75ccf46ab8652a659df7bccac7730f83d475b99714bbac0d

Observation 7ffc17f1-87cf-46b9-890f-4755ca5b67c0 · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:59.060875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:59.060875Z digest=sha256:359593ecbe7c0ace18d9e7522bf0377bb7c7b4da7a6970e17f709f89ef93cf9b

Observation b7b54d83-a12a-48b5-b8e2-e59d7e19407d · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:59.152970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:59.152970Z digest=sha256:d48c86daa2019f2f3ffbec1a48b811e558f6eee71dc260a536f5eee4380cbd14

Observation 6b018707-b816-462a-aac6-27b64fe7b027 · outbound

This paper cites Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:59.307542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:59.307542Z digest=sha256:f24960cf22c1702955bf6e58a0c2eb576904f3c375beb948955f394d1ddc78a7

Observation 102b32df-884b-43d1-b960-e2d78ecd357d · outbound

This paper cites Qwen2.5 Technical Report.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Qwen2.5 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:59.436247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:59.436247Z digest=sha256:8448fc18569ea03b36d2d944f5c692dadf2f43c8276d300aacf3c6436cdf87ff

Observation 7fec9ff2-183b-46a0-8969-c4bbb43d83ab · outbound

This paper cites Maniqa: Multi-dimension attention network for no-reference image quality assessment.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Maniqa: Multi-dimension attention network for no-reference image quality assessment

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:59.576862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:59.576862Z digest=sha256:cace0a4ac43e4174c6e0429014c568cfe849ccf81d64300cc67b76e3058cfc7b

Observation d3a5c374-d6a5-4f20-8374-6f3f60c15a09 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:59.692876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:59.692876Z digest=sha256:c32d16d52aa8124653f44531a8c55a95163db5cd343fa45fdb71223ce695b27f

Observation d8945482-6b3c-4c05-8c01-c345ae36d1ed · outbound

This paper cites Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:59.808955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:59.808955Z digest=sha256:07485b203fc1617d128fbb7f3f581ac7b09b496da08545713547d7b958204301

Observation 1a7707f2-2d37-44f4-9fee-803663185f88 · outbound

This paper cites Uniavgen: Unified audio and video generation with asymmetric cross-modal interactions.arXiv preprint arXiv:2511.03334, 2025.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Uniavgen: Unified audio and video generation with asymmetric cross-modal interactions.arXiv preprint arXiv:2511.03334, 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:59.923575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:59.923575Z digest=sha256:0dddaded35976110c99618d0b095d20d9fdee11b756ba105b381b2e9117218d2

Observation 69a06d9f-05bd-435d-9b0d-efe3bfd608a5 · outbound

This paper cites Waver: Wave Your Way to Lifelike Video Generation.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Waver: Wave Your Way to Lifelike Video Generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T16:26:00.037365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:26:00.037365Z digest=sha256:94a7a5240526c1d154cc3001c31730d9108779a373236b7cc31e0cb44c301d78

Observation fd5b86f8-7e4b-42c5-aa90-1aaa05de5c4d · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T16:26:00.172036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:26:00.172036Z digest=sha256:3880b7b949f4b46aa1430754557bd03f3d4e99d027caef71b82033de8c59194a

Observation 83f1c4a2-278e-4948-ad99-b96b99a7765a · outbound

This paper cites Uniform: A unified diffusion transformer for audio-video generation.arXiv e-prints, 2025.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Uniform: A unified diffusion transformer for audio-video generation.arXiv e-prints, 2025

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T16:26:00.285585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:26:00.285585Z digest=sha256:3ec3626f337149688874d182990a2369968c17f9efbec4319374d0f672eded59

Pith citing papers

Observation 039d90a1-8358-47fe-a869-deab40f14716 · inbound

Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence cites this paper.

Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-08-03T03:15:29.416338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:13:06.005185Z digest=sha256:7f387428adae5ad70675b6cd44afedf1e105c78c8913af21c5936d20dedb0474

Observation 106168a0-6b39-4c67-a652-01b39ba9afa4 · inbound

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation cites this paper.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-08-03T03:15:29.416338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:bc24df15eaba1e83b9b02c41a2a2c5654b4750b9594df75e641b18b60c966a71

Observation c01f1bf5-3a78-4e08-b828-3d6f4c924adb · inbound

OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation cites this paper.

OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-08-03T03:15:29.416338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T05:19:49.671136Z digest=sha256:8b571d800d6e0128a7ae4105de57fe04b1b7174671268c4a245a84960a4b943c

Observation 81c0ff12-b1ad-4d20-94d9-b90843758d4a · inbound

OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation cites this paper.

OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-08-03T03:15:29.416338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-05T12:34:41.758238Z digest=sha256:afd771dc88d84da93fc8c12ba6c4290692aab8399cbef5f526aec7d7cf01c853

Observation 19ec8ee7-9b08-4f8c-8865-d6376b661e51 · inbound

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation cites this paper.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-08-03T03:15:29.416338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:38ae78ea187bb904649596cdddb23d5713d506c50dab6590b5fb94fd15d1027a

Observation 5dc8bbe9-ae49-4ffe-8850-c4b2235f4e8d · inbound

SyncDPO: Enhancing Temporal Synchronization in Video-Audio Joint Generation via Preference Learning cites this paper.

SyncDPO: Enhancing Temporal Synchronization in Video-Audio Joint Generation via Preference Learning JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-08-03T03:15:29.416338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T07:01:40.333448Z digest=sha256:1bf7a1cc5d84e7664d7759478752c51c189de46d6c218a9d4f81375f37c5e98b

Observation 41339ad0-1024-4195-90ff-7285d7d83b7c · inbound

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning cites this paper.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing

Reference 133

Resolution
verified exact
arxiv_id, observed 2026-08-03T03:15:29.416338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:a0cce46ac6881aee36d5b13a2799e36bb7927afd3473780fb50b595e0157bd07