Pith. sign in

Paper Citation Record · LEDGER

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

As of 4 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2607.23855.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.23855 v2

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T01:51:01.976970Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e782bb66-d958-4d08-a0de-535576c1bae4 · outbound

This paper cites Auto-Encoding Variational Bayes.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Auto-Encoding Variational Bayes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:56.838514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:56.838514Z digest=sha256:2018460f78d1fc03efd35922bd97f3de1875d1bf353bac706aa9ba04300b8bdf

Observation f1e7537f-ad84-4ebd-ba0f-abe4e25c184d · outbound

This paper cites Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:56.920896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:56.920896Z digest=sha256:edb3eaea4ff11c1f2e453e5cc70684be68629977c7916c9e0106a4f493adc813

Observation e3a61767-901c-4fab-a645-7f9a4cca5204 · outbound

This paper cites Diffusion Transformers with Representation Autoencoders.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Diffusion Transformers with Representation Autoencoders

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:57.072895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:57.072895Z digest=sha256:4aa3b9de1f9951704459a2fbff08e06723fcffc0d5354474fc6c2e22d21960e9

Observation 184be546-ffba-4adf-86ec-b83c1e6b6403 · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:57.207168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:57.207168Z digest=sha256:395d42aefb9aa79b5844dea733be4ee4bc502a3bf4297cb3a511fd77e7f3ddef

Observation 7a5a578d-8314-48f2-bb12-9320ee25e173 · outbound

This paper cites MOSS-Audio-Tokenizer: Scaling audio tokenizers for future audio foundation models, 2026.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation MOSS-Audio-Tokenizer: Scaling audio tokenizers for future audio foundation models, 2026

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:57.372899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:57.372899Z digest=sha256:69873bf39fcd90752f09307c2c19edd41060855770e12081912d37700e2c14ac

Observation b56f5550-89ea-42d0-86ea-fcdc90528827 · outbound

This paper cites MMAudio: Taming multimodal joint training for high-quality video-to-audio synthesis.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation MMAudio: Taming multimodal joint training for high-quality video-to-audio synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:57.498868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:57.498868Z digest=sha256:d43f210e0e8f57d1e5a1aa49354b484c75be2a762dd8098821b2ff0a3517e554

Observation cd819ac4-c5f8-4256-a072-c63afde3c7ac · outbound

This paper cites HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:57.596815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:57.596815Z digest=sha256:68b916a2367e55274110ce271431facc5b4649c3f93a457b61c2f773c6d58f89

Observation faf7098c-f117-4ad2-9625-efb1f819750f · outbound

This paper cites Synchformer: Efficientsynchronizationfromsparse cues.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Synchformer: Efficientsynchronizationfromsparse cues

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:57.708665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:57.708665Z digest=sha256:758eaa9e9cbb4c12e17ecae4d7747accf36d4b087ff9a85fadf6dbcbe2e92b80

Observation 20bcff55-fcfd-4787-b494-5d7ceb524813 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Representation Learning with Contrastive Predictive Coding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:57.909221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:57.909221Z digest=sha256:f4817b42e5564268ea7e245d6ff5ee6818ef8a2881c0ac117ad445907d73dcd8

Observation 8504279f-e18e-4704-a13f-cb813b8e5905 · outbound

This paper cites Qwen3-Omni Technical Report.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Qwen3-Omni Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:57.917314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:57.917314Z digest=sha256:12011507289c51c7db3129d3ac65952b28fe639bfc2e5f2ec06483c98634104e

Observation e7ac17c1-a950-4760-90fc-950c95dead22 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Learning Transferable Visual Models From Natural Language Supervision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:57.934235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:57.934235Z digest=sha256:b76693ba8a94e5890be4bf76b9ff2de55b160cad52099dcafbab31c6774194b5

Observation 83822ea9-a08a-4d7b-b085-7fa3729229f7 · outbound

This paper cites ImageBind: One Embedding Space To Bind Them All.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation ImageBind: One Embedding Space To Bind Them All

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:57.995004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:57.995004Z digest=sha256:37f8091ebbb3496b53a890b562318a47d262129f2220e8eba7e95aa4d67e1d5e

Observation d29f2f3e-8aef-4907-9929-12debd347a58 · outbound

This paper cites LanguageBind: Extending video-language pretraining to n-modality by language-based semantic alignment.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation LanguageBind: Extending video-language pretraining to n-modality by language-based semantic alignment

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:58.086704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:58.086704Z digest=sha256:46ee4cd23f0ef0869e8d57aa17e41ecf1f3dbe9494f21eec33047bee24ad4bc4

Observation 3b704d31-f3ba-4ae5-8cbe-814ae169467a · outbound

This paper cites Liu, SouYoung Jin, Cheng-I Lai, Andrew Rouditchenko, Aude Oliva, and James Glass.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Liu, SouYoung Jin, Cheng-I Lai, Andrew Rouditchenko, Aude Oliva, and James Glass

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:58.338091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:58.338091Z digest=sha256:4e0b1b941d526ca024094c2c4872d26c8a5859213f4052dc58ad175e0805cf1f

Observation 4ebfe323-4658-46e5-b1e2-e9c93b42f741 · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:58.195894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:58.195894Z digest=sha256:8fd5f3006bbb63d2f05feb10d66f055118d837cd9124ac10d31f89f4d20d63ce

Observation 6b474239-886f-4980-a8fe-b3b86fb3092f · outbound

This paper cites Enhancing multimodal unified representations for cross modal generalization.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Enhancing multimodal unified representations for cross modal generalization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:58.554333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:58.554333Z digest=sha256:70557ddb808b5ac3fb096e2450bd0ee811939b3598b8c423bfc890c1932a3fb3

Observation 89510f85-4493-46d8-8234-162221f21ee9 · outbound

This paper cites Achieving cross modal generalization with multimodal unified representation.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Achieving cross modal generalization with multimodal unified representation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:58.446914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:58.446914Z digest=sha256:9a7f2a5275f921358ad54ce9e295836cc508124a0f6c007c15400d383ce6d477

Observation 8ff6db8f-ccb6-4ee2-9e5e-5aab3c13a336 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:58.816957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:58.816957Z digest=sha256:737fa3ab18aebe478f471ad4ac34867a6e7d5257d9ff314ba060dffc71a9abc6

Observation a9270da1-f028-4722-8008-f5395ea7f785 · outbound

This paper cites Out of time: Automated lip sync in the wild.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Out of time: Automated lip sync in the wild

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:58.703150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:58.703150Z digest=sha256:03256ac85e46a032e4bdb5b1c631d456aecf9810ffd0037be5ff2de6105ea5fa

Observation b693b372-f65d-4222-9903-f5161639ba97 · outbound

This paper cites High-Fidelity Audio Compression with Improved RVQGAN.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation High-Fidelity Audio Compression with Improved RVQGAN

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:59.050808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:59.050808Z digest=sha256:d40a623e811463fbbf0a056e3b70c1faf84c1f199032041fc87afe261c883a28

Observation 2064a707-8cb0-484b-9b56-6f8a5534a63f · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Movie Gen: A Cast of Media Foundation Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:58.947738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:58.947738Z digest=sha256:cbbbf00bc25f48c4707fbe4b03b58adce6c6e5edc3528a4393b1cbe120708ddc

Observation 4b26fcb7-0187-4926-859f-28296984aec5 · outbound

This paper cites Qwen-Audio-VAE Technical Report.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Qwen-Audio-VAE Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:59.197280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:59.197280Z digest=sha256:2f0469c27ec748067ea2a8d2a9bcffc51fa63517c3e79a885af1bad20c378c0e

Observation 78d88f69-da67-4d9b-8c79-b72340a512e1 · outbound

This paper cites Fast Timing-Conditioned Latent Audio Diffusion.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Fast Timing-Conditioned Latent Audio Diffusion

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:59.120923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:59.120923Z digest=sha256:35c58a30000a2b7978a78ac69b3dbe542d5c8cb1661b2c21a9c4b0c59d7f9a9f

Observation 26a46fe7-08e5-46da-85b9-7bb7048e11c0 · outbound

This paper cites UniTok: A unified tokenizer for visual generation and understanding.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation UniTok: A unified tokenizer for visual generation and understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:59.330818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:59.330818Z digest=sha256:75baadda16769642dc72c0723898cc118551a081254ae88c6eccd2df5b95a057

Observation b20d0e2d-08d8-46e7-a443-ecfffa5a2f23 · outbound

This paper cites REPA-E: Unlocking VAEforend-to-endtuningwithlatentdiffusiontransformers.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation REPA-E: Unlocking VAEforend-to-endtuningwithlatentdiffusiontransformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:59.256143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:59.256143Z digest=sha256:7b49f7174de436148dc35bfb732f12dda0eb3d3cd5a7de2dfb9e205ef87199f5

Observation 48b9be9c-e0bc-4d85-ae61-c108aca71a42 · outbound

This paper cites MM-Diffusion: Learning multi-modal diffusion models for joint audio and video generation.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation MM-Diffusion: Learning multi-modal diffusion models for joint audio and video generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:59.481977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:59.481977Z digest=sha256:9581a52782bbac038707d0912426e82cef1492b133aac3d816a6e12a5a6b6892

Observation 74a75c4e-7aac-4022-bbb6-70c029d71a07 · outbound

This paper cites Towards scalable pre-training of visual tokenizers for generation, 2025.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Towards scalable pre-training of visual tokenizers for generation, 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:59.413155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:59.413155Z digest=sha256:cd3cf46e755f5623950e9a73169b226b555637eb4d9efe253b32bd43b36d5a2a

Observation 8bb63cca-50b1-4840-8161-b25366335da4 · outbound

This paper cites Sora 2 system card.https://openai.com/index/sora-2-system-card/, 2025.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Sora 2 system card.https://openai.com/index/sora-2-system-card/, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:59.644940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:59.644940Z digest=sha256:11708af95b65672c9f0102b5bab92a5c5ec7568806ac8a81a9d27be1979d1e5d

Observation 61d446cc-1419-4cc7-9154-3d12603f7e74 · outbound

This paper cites AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:59.576752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:59.576752Z digest=sha256:0be578de475c24a0c59a7c10a0ed72d86063f9c8a487227c1c22b275fada21cb

Observation e996ac06-3573-4d32-b2b9-0bf410c64498 · outbound

This paper cites Seedance 2.0: Advancing Video Generation for World Complexity.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Seedance 2.0: Advancing Video Generation for World Complexity

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:59.779959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:59.779959Z digest=sha256:dbca900bff65abd5c9b82ac11b17d84a32f041226c3d218d30b6389b78bd2a14

Observation e3612b57-5072-443f-8836-3efa3022208d · outbound

This paper cites Veo 3.https://deepmind.google/models/veo/, 2025.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Veo 3.https://deepmind.google/models/veo/, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:59.705460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:59.705460Z digest=sha256:1e5b39f40af3e93da45e6e4e96db96f9c346afa2ebe086aa2325fc9a41e89e07

Observation 38d2269b-b1ec-4fac-a98f-f90a832b2a85 · outbound

This paper cites Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:59.933544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:59.933544Z digest=sha256:25723820d638cb36bb62af263e3fc046cfeb523792b7285b5ea7836276ea5bc4

Observation 418cd17a-d556-40e1-8eae-683d909792ad · outbound

This paper cites JavisDiT: Joint audio-video diffusion transformer with hierarchical spatio-temporal prior synchro- nization, 2025.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation JavisDiT: Joint audio-video diffusion transformer with hierarchical spatio-temporal prior synchro- nization, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:59.876252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:59.876252Z digest=sha256:1317531f2867c5ea6b0ccaf1b18d0c7ad0fa2189268344b881661c5c86d353b7

Observation caa853c6-7d08-49ee-a873-848a57706c12 · outbound

This paper cites MOVA: Towards scalable and synchronized video-audio generation, 2026.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation MOVA: Towards scalable and synchronized video-audio generation, 2026

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T01:51:00.058226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:51:00.058226Z digest=sha256:82424da662a0a50a8ee2a35088848525cea0cd7b8caad33aaa64582c694810a0

Observation 1acf9e11-fef4-4762-864b-c2ba78f7c000 · outbound

This paper cites LTX-2: Efficient joint audio-visual foundation model, 2026.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation LTX-2: Efficient joint audio-visual foundation model, 2026

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T01:51:00.009008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:51:00.009008Z digest=sha256:8cd9f87eb0aa05998719db74427aab43ad60980fdd45403458e305ae43cd22b8

Observation 8a89e8ce-6141-4927-ab1e-00ef378e3e35 · outbound

This paper cites Rep- resentation alignment for generation: Training diffusion transformers is easier than you think.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Rep- resentation alignment for generation: Training diffusion transformers is easier than you think

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T01:51:00.206338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:51:00.206338Z digest=sha256:46564a1bc0ab215a86255c4d2349f9d1e2872ae1f4a514a62d41e276bd8e9ddf

Observation 25f649a4-0a53-4691-8c6d-e7ed0adf4c2b · outbound

This paper cites Wan2.2-T2V-A14B model card.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Wan2.2-T2V-A14B model card

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T01:51:00.124013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:51:00.124013Z digest=sha256:993ae365b1b9308fc6f3b0fda465d022155fa4b0784df1e6b281d68a358859da

Observation d1900e63-3023-4997-8c15-883fd06e9243 · outbound

This paper cites Efros, Eli Shechtman, and Oliver Wang.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Efros, Eli Shechtman, and Oliver Wang

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T01:51:00.360367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:51:00.360367Z digest=sha256:b530a890b1fdb6aff720be5a211096721e3e9f63a30f345df4f0ee809f6f67d1

Observation ca114542-a971-4ab7-9cda-9b701c9b95d1 · outbound

This paper cites What matters for representation alignment: Global information or spatial structure?, 2025.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation What matters for representation alignment: Global information or spatial structure?, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T01:51:00.271601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:51:00.271601Z digest=sha256:f27a346a2d2d7ce12bbdcbb0fd1825b5900eb65bddee08c15e00ade4716603c7

Observation c1be5cbd-a4f5-4885-a934-f673833d93e6 · outbound

This paper cites Gemmeke, Daniel P.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Gemmeke, Daniel P

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T01:51:00.544877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:51:00.544877Z digest=sha256:2531fc442ccdb19e93d5f3a8440145e928d3650d882ff12a461a540dced2c617

Observation 3cdad4a4-601b-452a-afb4-4458397ec441 · outbound

This paper cites BEATs: Audio Pre-Training with Acoustic Tokenizers.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T01:51:00.418732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:51:00.418732Z digest=sha256:0f9104e277c2cb3348b205b659c22aaace9e9f0ebf426466e4341a12d2e29fc3

Observation 8ce5ffb6-5ac8-47f8-80a3-e747aa20f0e6 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T01:51:00.729036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:51:00.729036Z digest=sha256:87bdb145c13455f513c047851f34b118f29e16cea08105ab744eb33d74615704

Observation 49a66139-33f0-48e3-be67-1d7851216126 · outbound

This paper cites Qwen3.5: Towards native multimodal agents, 2026.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Qwen3.5: Towards native multimodal agents, 2026

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T01:51:00.634694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:51:00.634694Z digest=sha256:ee343241b5c8291b63625febd8222977f56fbf41935c3e0122693fc878162298

Observation fae64e3d-4093-4708-bc11-4c0930228803 · outbound

This paper cites LibriSpeech: An ASR corpus based on public domain audio books.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation LibriSpeech: An ASR corpus based on public domain audio books

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T01:51:00.962236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:51:00.962236Z digest=sha256:cc42e28da0440333ab95ba8d24204806d74b108873d3c75ad53161357ce2908f

Observation f8fdcd0a-d5ec-4d00-bd86-7db18e700773 · outbound

This paper cites Panda-70M: Captioning 70m videos with multiple cross-modality teachers.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Panda-70M: Captioning 70m videos with multiple cross-modality teachers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T01:51:00.888437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:51:00.888437Z digest=sha256:6f3dd8a5de0576ea0c61ea1512b7da9cf2d619ac585b0ca29c323b8a73261723

Observation 40bc4501-15fe-44f1-ad95-1c5d689f1cb4 · outbound

This paper cites Parker, Matthew Rice, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Parker, Matthew Rice, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T01:51:01.215735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:51:01.215735Z digest=sha256:303f4f8ec73a81e9ec77ac9cd4865b254ef265807b63ac9743a1aab9d8fb3800

Observation c461c2d0-a264-46cc-9259-24ad685b063e · outbound

This paper cites The MUSDB18 corpus for music separation.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation The MUSDB18 corpus for music separation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T01:51:01.120402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:51:01.120402Z digest=sha256:7ac1968de2eaf406d32d18f8455675b3a08c3de3d1615c94753d7539b8d4dd8a

Observation 21161479-4a3c-482d-ba48-e507aec59e1c · outbound

This paper cites Sparse in Space and Time: Audio-visual Synchronisation with Trainable Selectors.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Sparse in Space and Time: Audio-visual Synchronisation with Trainable Selectors

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T01:51:01.639669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:51:01.639669Z digest=sha256:f21fea43a3eabbf88c4affa4eba28dfc1b4542417e5446e5108326d71de3c6d5

Observation bf7e1cdb-655d-47ae-bc6f-20dc5ede3e92 · outbound

This paper cites Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T01:51:01.799954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:51:01.799954Z digest=sha256:049256af6b3beca531483d26c0fae381d32c5cecb456f6d0b03ad9e10b63062a

Observation ca8cbbe6-f844-4c67-b15f-830ea4881695 · outbound

This paper cites VGGSound: A large-scaleaudio-visual dataset.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation VGGSound: A large-scaleaudio-visual dataset

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T01:51:01.477177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:51:01.477177Z digest=sha256:1ec0b14664744bf5a4ac8ec8f98eac249a9663b685aa549a3f5994b0d0b6f194

Observation 43146570-5131-4033-93ae-b3197ea0369b · outbound

This paper cites Multi-task learning using uncertainty to weigh losses for scene geometry and semantics.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Multi-task learning using uncertainty to weigh losses for scene geometry and semantics

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T01:51:01.976970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:51:01.976970Z digest=sha256:86b1af209a13ae39a74213665e1671dbbcc3152cb9d4a3a81546bd461b8f7cb9

Observation 118dfa17-bd41-4389-ab74-5aae05eaebde · outbound

This paper cites UniVerse-1: Unified Audio-Video Generation via Stitching of Experts.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation UniVerse-1: Unified Audio-Video Generation via Stitching of Experts

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T01:51:01.908095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:51:01.908095Z digest=sha256:2bc637f2c7f2859ee0581653cd2e526c61668792d2184585919f99ebd308cd78

Observation 7dee2a1a-47b2-437d-bb24-c1d1b052b0d3 · outbound

This paper cites an unresolved cited work.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:57.814691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:57.814691Z digest=sha256:7f1e277513870a6f3f309d33f06901d59777c7aa5a64e69f0387711c4f6f6591

Observation 8a027708-046b-4f67-8d43-77234f34fa63 · outbound

This paper cites Stable Audio 3.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Stable Audio 3

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-03T01:51:01.349808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:51:01.349808Z digest=sha256:bb6ba15178f5bbd9dc5e47004ac41aa4464626d4d771e25d893c3b4ddfe328a9

Pith citing papers

No inbound Pith citation observations are available.