Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T13:07:28.484954Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2507.20987.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T13:07:28.484954Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
46 of 46 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9f8eb8eb-c408-4fb3-bb91-927714cba40b · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Better speech synthesis through scaling
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd26a3e6-2448-4370-a2b4-b2a57e88ba75 · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aec1f845-f6cc-4662-82f3-f749f6d011ae · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Caron, H
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9ba859d9-ee81-4060-baff-24242d4ea042 · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46495ded-d986-48ff-b43c-3b6fce0f7863 · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc3a77a5-34c6-4bfa-8977-54a8de14198f · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 DeepMind
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 90e4fa50-9a9a-4a34-b8f1-b5d945e5cd43 · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Eleven Multilingual v2: A foundational multi- lingual text-to-speech model supporting 29 languages
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0d58c0c6-b868-4e86-a62e-335215e975b1 · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8e0641f0-7e13-4a3d-92ce-7959a053128d · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bb966614-8d46-43d4-8541-a0b8d25470cc · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b8e0228-cbc8-4122-ac5a-358bf6dfcf9c · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c62460bc-87fe-4d26-96fc-13a071ce6405 · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Huynh-Thu and M
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d7897926-fa51-4621-b3e6-b33f79123df8 · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Sonic: Shifting Focus to Global Audio Perception in Portrait Animation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 466c15f3-68e5-430a-a523-6f0b68f7208f · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21564f20-7b37-41e5-a88a-01c967b8ac72 · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9aad734f-4aa9-49d6-b058-f8c508a0b8d1 · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 HunyuanVideo: A Systematic Framework For Large Video Generative Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a359396-9b3a-4e95-887c-2bf2ba39097f · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Laion aesthetic predictor.https : / / github.com/LAION-AI/aesthetic-predictor,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00912a17-7a4f-48c1-b26c-328ee3e8e517 · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6c8752d-af4f-44c6-848d-8be9cc31a660 · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d7bc63b-cc26-415a-8776-22e95544bb13 · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Li, Z.-L
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5071510a-7b7c-4f9b-ae7e-6bffc5a1d227 · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 493cbbe6-407d-4264-9cb9-17e1b0a7a46d · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 EmergentTTS-Eval: Evaluating TTS Models on Complex Prosodic, Expressiveness, and Linguistic Challenges Using Model-as-a-Judge
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 249acc42-76b7-431f-9846-9f602c448da1 · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d65ff508-4c87-4cf0-bbb2-a3828f0ab2e5 · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 StyleTalker: One-shot Style-based Audio-driven Talking Head Video Generation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b24eed8-d9d4-4bcd-9af6-f624f2a4b329 · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 GPT-4o-Mini- Audio-Preview: A compact, cost-efficient audio-capable 5 multimodal model
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c36f0941-a34f-487b-9951-f4a20911b540 · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5625fff3-0bf1-40cd-92e9-1fc6ebba0e6f · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Prajwal, R
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 08efa263-aea2-490e-922e-01498341bd55 · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Versatile Multimodal Controls for Expressive Talking Human Animation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 012c4635-443b-4b3d-a756-32c02e46e51c · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c9c5499-7162-41be-8e14-9bfc4b43b05a · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Radford, J
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1d1b5d94-8103-4b62-a707-d8b6ac5e2323 · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Make-A-Video: Text-to-Video Generation without Text-Video Data
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f891e8f8-1f30-426a-85e8-3b82099ddc4f · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Bark: Text-prompted generative audio model
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fafa85cc-e0c0-4ce5-850c-464e45134edd · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Teed and J
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4494dd48-0f1b-4b3c-bb72-5b9b09201c56 · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Google introduces stable gemini 2.5 flash and pro, previews gemini 2.5 flash-lite.The Economic Times (India)
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 17262ae0-cc5e-4735-ba82-327d2cc7e50b · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 StableAnimator: High-Quality Identity-Preserving Human Image Animation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 911f72e8-e992-464b-a06f-98a3a7828c36 · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Towards Accurate Generative Models of Video: A New Metric & Challenges
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b96490e4-aa69-40c6-bacf-c7c7a17f4b8f · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Wan: Open and Advanced Large-Scale Video Generative Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f05d9c9e-8209-40d0-92dd-2439c2c20af3 · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5f57e324-f84c-4947-9944-0fa4a996ba26 · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e2a42bc8-8de1-48c5-86fa-931fa8aea62a · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Qwen2.5-Omni Technical Report
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58d16370-0f4a-40c3-99fb-33a481335cdd · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0e7c99b-5778-404a-9df4-1a142fa9dfc4 · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a65d3b03-373a-4102-baa0-0f2224156e3b · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d3e204d-6399-4076-b7db-fb482bc4b46a · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 49282654-e179-469a-b54b-936a4e4ae31d · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Unresolved cited work
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bbd571e3-6e76-491f-80dc-fdf992197bc7 · outbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 org / abs / 2505
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.