Pith. sign in

Paper Citation Record · LEDGER

Vorch-Omni: Multi-Task Orchestration of Sight and Sound

As of 8 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2608.05803.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05803 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:16:52.552901Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact5
  • verified fuzzy2
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 085f7559-aa81-4101-b8b4-4d893902df8a · outbound

This paper cites SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.478464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.478464Z digest=sha256:400471521ea49415fee5cfe864bc393942a35610db5cfd3be91e4f87aa61e6e6

Observation b7705bf0-d6d7-4bb3-a837-1f6f546c3682 · outbound

This paper cites Dreamid-omni: Unified framework for controllable human-centric audio-video generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Dreamid-omni: Unified framework for controllable human-centric audio-video generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.482667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.482667Z digest=sha256:4bb9ad143c1ed9b428a2f9f1079cff31fd3275e54da473ea33ea1e4b5f399a22

Observation efce4dc6-b3b2-4491-864c-6a11ebde837e · outbound

This paper cites Mmdisco: Multi-modal discriminator- guided cooperative diffusion for joint audio and video generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Mmdisco: Multi-modal discriminator- guided cooperative diffusion for joint audio and video generation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:16:53.584280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:16:52.491684Z digest=sha256:435d1bfe12c796583d59978b28d82db9bb5ee21156010a76791cf20f2f5ff59b

Observation c560d014-c578-457c-b133-377d5f55c84e · outbound

This paper cites HunyuanVideo 1.5 Technical Report.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound HunyuanVideo 1.5 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.495601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.495601Z digest=sha256:642f45dc9885660a401de5d7cf2e4ede9873a8337cb2affa4d20ea6782b54860

Observation 3da2db8c-0201-4506-a4fd-aed8acb3510d · outbound

This paper cites Native Audio-Visual Alignment for Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Native Audio-Visual Alignment for Generation

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:16:52.959820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:16:52.499515Z digest=sha256:1ac40837020487e9d9c3d37641624182241478049bff423e50ae6e76428c7ea1

Observation 5b81f74d-1820-4e9d-8c3f-7ca0ab47edb7 · outbound

This paper cites Kling-Omni Technical Report.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Kling-Omni Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.502930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.502930Z digest=sha256:a2729080c33f2464d5a6b12de1c2470974f5941785718680e6d3afeda27ecdc3

Observation 9ccd18d3-a82e-4af8-9f5d-da02fe9e350a · outbound

This paper cites MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:16:52.929357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:16:52.506326Z digest=sha256:84f07c256c1f4b162714f4f8c43a11beaf290e8e8a43deb2756e5fdfc2fe8ec8

Observation 787248e6-93ad-49ca-9c0f-a24ddfdf72b8 · outbound

This paper cites Flow Matching for Generative Modeling.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Flow Matching for Generative Modeling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.509929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.509929Z digest=sha256:3d458b28aef81859f725913d80ce3d33b897380f6cb4fc9bb670ae266ccd83e8

Observation 9468b1d8-23a1-490a-82c1-a3b1ef9b346b · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.513807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.513807Z digest=sha256:833d590b24f7b99b24efe67323caa875421938dac69d440d63b84c87cb7d75c4

Observation 89ef2161-5831-4a26-a182-418b358c793b · outbound

This paper cites Team OpenMOSS.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Team OpenMOSS

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.521689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.521689Z digest=sha256:97ce40ec99f49fa9e6a697d5fa710a35cadfc458f1409fb7977f16af9493ed13

Observation ae50f3a9-5f51-43d7-b4ea-c0c1fc095286 · outbound

This paper cites Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.529134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.529134Z digest=sha256:00494fec4257a7d07b7ac59291104bc8007cb6ed2ec7fee2b0731e41d347ed40

Observation 7d54ed40-72b2-4164-9ea7-829bc0695111 · outbound

This paper cites Baton: Explicit Semantic Blueprints for Joint Video-Audio Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Baton: Explicit Semantic Blueprints for Joint Video-Audio Generation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:16:52.635253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:16:52.532992Z digest=sha256:224d87c768d86e81ed224c1db0b0ce11ae3590d54ed1170b9947129dcceafbb6

Observation b30009e5-84ca-4f87-a714-af268271ee5b · outbound

This paper cites UniVerse-1: Unified Audio-Video Generation via Stitching of Experts.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound UniVerse-1: Unified Audio-Video Generation via Stitching of Experts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.540959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.540959Z digest=sha256:897f777af32b744b248941051f3dbd06fba71917b2c1e3e429f52e07cf21cd85

Observation d2b55d41-71fa-4823-8db6-cac76c6ffe71 · outbound

This paper cites Uniavgen: Unified audio and video generation with asymmetric cross-modal interactions.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Uniavgen: Unified audio and video generation with asymmetric cross-modal interactions

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:16:53.571568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:16:52.544882Z digest=sha256:7414653b7f598fbcd953099d535c6f3f0c6c87959250399b76246849fcac0250

Observation dd650d56-eda4-4d06-80ef-7bbfaa557f56 · outbound

This paper cites UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.548806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.548806Z digest=sha256:2cedf42465c3a58e8cb17ae6d91cf71dacae7f1853c17eb5bccc4480d51fa2ec

Observation d846b8d5-983e-4fe4-b01d-7996cf6fe358 · outbound

This paper cites InstructAV2AV: Instruction-Guided Audio-Video Joint Editing.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound InstructAV2AV: Instruction-Guided Audio-Video Joint Editing

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.552901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.552901Z digest=sha256:35781f7d79611bd613fc711d4d112ce31085b3ad265a9c2f2a62ba3628b6b08c

Observation 088567b2-ec9d-4d18-bd2c-453312f10e2e · outbound

This paper cites Guibin Chen, Dixuan Lin, Jiangping Yang, Youqiang Zhang, Zhengcong Fei, Debang Li, Sheng Chen, Chaofeng Ao, Nuo Pang, Yiming Wang, et al.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Guibin Chen, Dixuan Lin, Jiangping Yang, Youqiang Zhang, Zhengcong Fei, Debang Li, Sheng Chen, Chaofeng Ao, Nuo Pang, Yiming Wang, et al

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.464969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.464969Z digest=sha256:0e273e5b276bdc95b9f5931c22a05679dd9589e0cd343bd99905551de593c55f

Observation 87f681e6-8702-48ec-9471-5b10076424dd · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Wan: Open and Advanced Large-Scale Video Generative Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.536889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.536889Z digest=sha256:ee6c94e4db8764aacae34e778d1501dd234c5ea9351f9f2daae32341990ada77

Observation 8454eb56-0efb-475a-93c2-af9c3b5f0718 · outbound

This paper cites Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.517841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.517841Z digest=sha256:861669cd08f2ed9ce3ad3281f58a04ab5b5f948ca9fdefcbd35ffb0e0bbd62b1

Observation d3063727-59cd-4966-925a-49b47f222a7c · outbound

This paper cites AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation

Reference 2023

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:16:52.664679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:16:52.525251Z digest=sha256:56bf3a112022423dce8ab4a7c64b664fd49d486f7a1fe4dc8e65d9d0e4f180fa

Observation 47bbb723-024c-4a6a-8dc9-451b43a86f7e · outbound

This paper cites CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation

Reference 2024

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:16:53.347134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:16:52.469576Z digest=sha256:5b6ad79278a9903bcc58116c2005181fb4e361646e9ccca6ebeabc0d2c71f755

Observation 0fadbedf-3701-49b7-866b-674715952e60 · outbound

This paper cites Speed by simplicity: A single-stream architecture for fast audio-video generative foundation model.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Speed by simplicity: A single-stream architecture for fast audio-video generative foundation model

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.474289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.474289Z digest=sha256:b693ecc14e4f7d57456a6fccc8dd208db09c77d9b5d0e026a2b5f08540effbd5

Observation ec011e65-83f0-4560-aa54-334be2c026ab · outbound

This paper cites LTX-2: Efficient Joint Audio-Visual Foundation Model.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound LTX-2: Efficient Joint Audio-Visual Foundation Model

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.486493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.486493Z digest=sha256:d70d312524e75a4b49acd894a46ad6f6b6e54f44f5b3c3071401983db2577178

Pith citing papers

No inbound Pith citation observations are available.