Pith. sign in

Paper Citation Record · LEDGER

Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming

As of 21 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2608.05663.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05663 v2

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:29:35.963805Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eb89b08a-c5a7-4804-9ca3-f2981b74f2f2 · outbound

This paper cites Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation.

Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:35.933949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:35.933949Z digest=sha256:0482681028412fb91b3244b5f11a52429cf09e91d136e177b07863e05fa9ef0b

Observation d61c9683-b47c-459c-9e2d-fd1d50540a1d · outbound

This paper cites Yaofeng Su, Yuming Li, Zeyue Xue, Jie Huang, Siming Fu, Haoran Li, Ying Li, Zezhong Qian, Haoyang Huang, and Nan Duan.

Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming Yaofeng Su, Yuming Li, Zeyue Xue, Jie Huang, Siming Fu, Haoran Li, Ying Li, Zezhong Qian, Haoyang Huang, and Nan Duan

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:35.945967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:35.945967Z digest=sha256:1226178790e4cb99e18efb6e07714e6413eebf04890f4369abb57e53d1714601

Observation 69749f90-9b72-4028-b2bc-1b123ea579f9 · outbound

This paper cites StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration.

Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:35.949368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:35.949368Z digest=sha256:e02e1e888980ebc4220d04e82395d1bbe4b93fcafbad5b0a1768b3a715d96254

Observation f77f7ea1-626b-479a-bf65-0705ae76118e · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:35.952948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:35.952948Z digest=sha256:c4384d17edbe78a607e49eb1359d6a4f082f77bf9accd6bfeb4b0ce12ec53bd3

Observation 687f4def-a464-4abe-8d0f-7873a4c9aa4f · outbound

This paper cites LongLive: Real-time Interactive Long Video Generation.

Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming LongLive: Real-time Interactive Long Video Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:35.956596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:35.956596Z digest=sha256:31ea180b9d8b2616e7a28c5b3ea9aba582060d827e72512db2cffae961399daf

Observation 4e7a7e3d-bf41-4100-84a0-d93a515b1c39 · outbound

This paper cites VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness.

Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:35.960244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:35.960244Z digest=sha256:fcde83e708de5d00442bf8bf8dec79663ee7b3096a08d5856ee18e48f6daa36e

Observation b5c7cb3e-d353-42a9-96f8-8c187d8dfd57 · outbound

This paper cites Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation.

Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:35.963805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:35.963805Z digest=sha256:7b465c5a634448ef121015b15c982b74dd414d1cccc70fb82dd42832f4c16cf9

Observation 610c8caf-2a41-4d73-85b1-d1a24639c7c6 · outbound

This paper cites Qwen3-TTS Technical Report.

Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming Qwen3-TTS Technical Report

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:35.920343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:35.920343Z digest=sha256:ad6882c029c0098f30435bd391d9ba25a579cf6c04e72fdf2a9e4b71ac477a4d

Observation fc6e11c0-209c-4487-8b6f-b08e39271b5c · outbound

This paper cites LTX-2: Efficient Joint Audio-Visual Foundation Model.

Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming LTX-2: Efficient Joint Audio-Visual Foundation Model

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:35.915778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:35.915778Z digest=sha256:b702cad07ed695b6b3808d7a10d6fc342cc174108309ae7de8230947d38d4bf1

Observation 2e437e8c-84f7-4676-9e8f-1d3fc6846179 · outbound

This paper cites Seedance 2.0: Advancing Video Generation for World Complexity.

Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming Seedance 2.0: Advancing Video Generation for World Complexity

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:35.941299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:35.941299Z digest=sha256:f46c330719368a843791f466a231cc9bb5f963df2637a55bfc6e2f847fc98a10

Observation 6933ea50-b74c-4838-9e28-92ee8a085d9f · outbound

This paper cites Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation.

Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:35.929035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:35.929035Z digest=sha256:ddd5cb3a8b46aaf77206a41c4669bb56294e512f660147ffaeec22e6c233f4ed

Observation 2001127f-8aa6-46fc-9f4c-31a90c4e13ea · outbound

This paper cites Build llm-based zero-shot streaming tts system with cosyvoice.

Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming Build llm-based zero-shot streaming tts system with cosyvoice

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:35.937878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:35.937878Z digest=sha256:889e4c290dcd7ec11e253637bc26571b49364bf5e27dc4e097340b99404e9a51

Observation e8b79d37-e3e2-4fea-9eda-c81ff451a5f5 · outbound

This paper cites Wan-Streamer v0.2: Higher Resolution, Same Latency.

Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming Wan-Streamer v0.2: Higher Resolution, Same Latency

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:35.925051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:35.925051Z digest=sha256:6f0f163d1171e9d0e8a38a001a846040b7f7cb5b049d9252128ac88971ec614c

Pith citing papers

No inbound Pith citation observations are available.