Pith. sign in

Paper Citation Record · LEDGER

Vorch-Omni: Multi-Task Orchestration of Sight and Sound

As of 8 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2608.05803.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05803 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:16:52.552901Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact5
  • verified fuzzy2
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 085f7559-aa81-4101-b8b4-4d893902df8a · outbound

This paper cites SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.478464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.478464Z digest=sha256:211b786e271ed74289af3c63c2e2d1b64d79b06dcebfe20bc3031dff16ba3791

Observation b7705bf0-d6d7-4bb3-a837-1f6f546c3682 · outbound

This paper cites Dreamid-omni: Unified framework for controllable human-centric audio-video generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Dreamid-omni: Unified framework for controllable human-centric audio-video generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.482667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.482667Z digest=sha256:c8376a9fdde2ec7abd4f32a9d9c137469c30062e745be2eff9c8017611dd49a1

Observation efce4dc6-b3b2-4491-864c-6a11ebde837e · outbound

This paper cites Mmdisco: Multi-modal discriminator- guided cooperative diffusion for joint audio and video generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Mmdisco: Multi-modal discriminator- guided cooperative diffusion for joint audio and video generation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:16:53.584280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:16:52.491684Z digest=sha256:b55bab9a6f9884b07c5c0f937f86efb474ac062941c040e388379a1edd17aba5

Observation c560d014-c578-457c-b133-377d5f55c84e · outbound

This paper cites HunyuanVideo 1.5 Technical Report.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound HunyuanVideo 1.5 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.495601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.495601Z digest=sha256:bf0321b59f6767f41e808443e39214b98b1fdd12be86735da978b6b4d7a9bd62

Observation 3da2db8c-0201-4506-a4fd-aed8acb3510d · outbound

This paper cites Native Audio-Visual Alignment for Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Native Audio-Visual Alignment for Generation

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:16:52.959820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:16:52.499515Z digest=sha256:f1d2fae3fa0062d26f2c0b17609762e25a255b8f015ccfb2f2abbffe41a0a04c

Observation 5b81f74d-1820-4e9d-8c3f-7ca0ab47edb7 · outbound

This paper cites Kling-Omni Technical Report.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Kling-Omni Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.502930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.502930Z digest=sha256:9b122b825f6d40f608641f1f0f452d7b0d1540a770bf502b8672a00bfd47b794

Observation 9ccd18d3-a82e-4af8-9f5d-da02fe9e350a · outbound

This paper cites MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:16:52.929357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:16:52.506326Z digest=sha256:5bca3f7f720c8d6930debed76e881178f52e3be2c252d553345ca413321b3051

Observation 787248e6-93ad-49ca-9c0f-a24ddfdf72b8 · outbound

This paper cites Flow Matching for Generative Modeling.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Flow Matching for Generative Modeling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.509929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.509929Z digest=sha256:35d6b31e2a5e1c93133435f9a3153ee48056c937d8dbda13003f13c025c726d1

Observation 9468b1d8-23a1-490a-82c1-a3b1ef9b346b · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.513807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.513807Z digest=sha256:ec286f078b53ebef1c41478a3478877f0406453251be21020bf01431eed3598a

Observation 89ef2161-5831-4a26-a182-418b358c793b · outbound

This paper cites Team OpenMOSS.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Team OpenMOSS

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.521689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.521689Z digest=sha256:3d0ba5c9ec03bd7ed475a66db65f91ff0abd77f664c410e2682c5461b7ac0d70

Observation ae50f3a9-5f51-43d7-b4ea-c0c1fc095286 · outbound

This paper cites Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.529134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.529134Z digest=sha256:6638ef367d7e6d06beea7b2fcac9fb9d9e1b1a39c73809d3a07cc04acf493bf6

Observation 7d54ed40-72b2-4164-9ea7-829bc0695111 · outbound

This paper cites Baton: Explicit Semantic Blueprints for Joint Video-Audio Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Baton: Explicit Semantic Blueprints for Joint Video-Audio Generation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:16:52.635253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:16:52.532992Z digest=sha256:52e67b98ebd2fe0e6e3fda36df07f8e8ddad5ac2caa43a798b09443135b93b65

Observation b30009e5-84ca-4f87-a714-af268271ee5b · outbound

This paper cites UniVerse-1: Unified Audio-Video Generation via Stitching of Experts.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound UniVerse-1: Unified Audio-Video Generation via Stitching of Experts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.540959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.540959Z digest=sha256:fd17e22f8ce6b91854fb97c407b7c57c97c02218fd930f40e2b2ece772c47f0c

Observation d2b55d41-71fa-4823-8db6-cac76c6ffe71 · outbound

This paper cites Uniavgen: Unified audio and video generation with asymmetric cross-modal interactions.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Uniavgen: Unified audio and video generation with asymmetric cross-modal interactions

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:16:53.571568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:16:52.544882Z digest=sha256:c912a80ad1b98210b8b436f79dae74f10118fde0fbb9334a264185e59bf0429a

Observation dd650d56-eda4-4d06-80ef-7bbfaa557f56 · outbound

This paper cites UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.548806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.548806Z digest=sha256:246d1bd536aec51ae446cb44ecdec71b9c9d12b1cc042a4780a6d85fb52ce118

Observation d846b8d5-983e-4fe4-b01d-7996cf6fe358 · outbound

This paper cites InstructAV2AV: Instruction-Guided Audio-Video Joint Editing.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound InstructAV2AV: Instruction-Guided Audio-Video Joint Editing

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.552901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.552901Z digest=sha256:2de6839d805b95ec41d3a196301127107cfa2b825ea0252620b571a305db9a29

Observation 088567b2-ec9d-4d18-bd2c-453312f10e2e · outbound

This paper cites Guibin Chen, Dixuan Lin, Jiangping Yang, Youqiang Zhang, Zhengcong Fei, Debang Li, Sheng Chen, Chaofeng Ao, Nuo Pang, Yiming Wang, et al.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Guibin Chen, Dixuan Lin, Jiangping Yang, Youqiang Zhang, Zhengcong Fei, Debang Li, Sheng Chen, Chaofeng Ao, Nuo Pang, Yiming Wang, et al

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.464969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.464969Z digest=sha256:1fca53f7039cce804b5aec4589cce9489be510be59b321ec2dcc3009ff0fd619

Observation 87f681e6-8702-48ec-9471-5b10076424dd · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Wan: Open and Advanced Large-Scale Video Generative Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.536889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.536889Z digest=sha256:e6cb13d26f047c2bf4e9aa0ea55042d1857666d52ed40a60b80e7aaf9814a758

Observation 8454eb56-0efb-475a-93c2-af9c3b5f0718 · outbound

This paper cites Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.517841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.517841Z digest=sha256:7b3e585b3e06423494d19965edeaa4245511af34746c5b94ae4741738d77d20f

Observation d3063727-59cd-4966-925a-49b47f222a7c · outbound

This paper cites AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation

Reference 2023

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:16:52.664679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:16:52.525251Z digest=sha256:1a9cbf68713e12fd8c08e2b63b3695bcab32d1c7c941ffbc157f47689664c25d

Observation 47bbb723-024c-4a6a-8dc9-451b43a86f7e · outbound

This paper cites CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation

Reference 2024

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:16:53.347134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:16:52.469576Z digest=sha256:642ca232fed7fbfd9e417021029b56707c516c7f0eb0b8e8303fe9f10367d2ce

Observation 0fadbedf-3701-49b7-866b-674715952e60 · outbound

This paper cites Speed by simplicity: A single-stream architecture for fast audio-video generative foundation model.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Speed by simplicity: A single-stream architecture for fast audio-video generative foundation model

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.474289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.474289Z digest=sha256:fa8621380cda31454cfda806bec9b711e86a240257a94d769746825340b90b94

Observation ec011e65-83f0-4560-aa54-334be2c026ab · outbound

This paper cites LTX-2: Efficient Joint Audio-Visual Foundation Model.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound LTX-2: Efficient Joint Audio-Visual Foundation Model

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.486493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.486493Z digest=sha256:6be8b3d6751d78bea3485c0e289dad8602b438510d2d0f84e43b30d381c0964f

Pith citing papers

No inbound Pith citation observations are available.