Pith. sign in

Paper Citation Record · LEDGER

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos

As of 9 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2508.08891.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.08891 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:24:08.721929Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact2
  • verified fuzzy10
  • unresolved27
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation de2bf21f-4036-459e-8002-dcefadac5b4b · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:05.359067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:05.359067Z digest=sha256:6b1b66b292a0a6c72bda22207eefa0229477a5c7a717abd1e6c102e95274b283

Observation f01b6fc9-2761-4af7-8c24-c0df7620a36c · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Emerg- ing properties in self-supervised vision transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:05.426808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:05.426808Z digest=sha256:cc1befdc8f9214b19eefe136bbb3b3eb4ad4a4fb9fd2d1b77f3ab895682d8d0d

Observation 66a2cb35-9ef2-4aa1-990f-4f17cadf014f · outbound

This paper cites Dimitra: Audio-driven Diffusion model for Expressive Talking Head Generation.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Dimitra: Audio-driven Diffusion model for Expressive Talking Head Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:05.525128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:05.525128Z digest=sha256:ab6dd232241a5041c9c2da73ba72dbc57bb0acaf8ba8691e71ed5a8db3a7f99c

Observation 9d7248bf-9fae-4233-a7e6-de05a57562a0 · outbound

This paper cites Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:05.595256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:05.595256Z digest=sha256:7165bb1a33d0c56584c97720c03da05b1d2551ddc85b308bb121727b40691396

Observation d160ab63-8d60-43af-b7dd-56f2ed06878c · outbound

This paper cites Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:05.723175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:05.723175Z digest=sha256:5b7aa3a209bfac44ce27432051cf7d552f32df344bb6298a22997088762658b7

Observation d42da7a8-dbf2-4411-9906-0afc46933a68 · outbound

This paper cites Gemini 2.5 pro: Our most intelligent ai model, 2025.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Gemini 2.5 pro: Our most intelligent ai model, 2025

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:24:11.556123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T21:24:05.819589Z digest=sha256:0291d21b3922c71dc987387d8097f4367d9c0938d25a42679e9806a1bed55f52

Observation dd4d3973-3b7d-4bba-98d6-54fc4a4c9f97 · outbound

This paper cites Stylesync: High-fidelity generalized and personalized lip sync in style-based genera- tor.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Stylesync: High-fidelity generalized and personalized lip sync in style-based genera- tor

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:24:11.416785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T21:24:05.908061Z digest=sha256:64994ccc68fdce9614b4375946f1f9a6747a8358ebb9da56b074923e12c3fcc8

Observation 02cab4ec-693a-4221-b10c-8befa4cd10b2 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:06.035484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:06.035484Z digest=sha256:f1751bb07db149372d7007787abcea9e873b82733df579bc080a8b0b56294431

Observation 34596d63-1629-49b9-8b82-950be13908ba · outbound

This paper cites Video dif- fusion models.Advances in Neural Information Processing Systems, 35:8633–8646, 2022.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Video dif- fusion models.Advances in Neural Information Processing Systems, 35:8633–8646, 2022

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:06.127789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:06.127789Z digest=sha256:e3070ed561935cebcaa5e6b913db145a95deb5b39fbfb75a440c118b25cb13a6

Observation 5fa83f4b-3bcc-4a19-bffc-ba006d6dbe76 · outbound

This paper cites Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:06.186484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:06.186484Z digest=sha256:58eeb5365288ab8c0e81fddafb4bd3bedce52889054e059e1f0cb02620f6c039

Observation 6471fbcf-7a4d-4bf6-921e-b3a5270daf23 · outbound

This paper cites The accu- racy of psnr in predicting video quality for different video scenes and frame rates.Telecommunication systems, 49:35– 48, 2012.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos The accu- racy of psnr in predicting video quality for different video scenes and frame rates.Telecommunication systems, 49:35– 48, 2012

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:24:11.226142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T21:24:06.279489Z digest=sha256:6529a0c36c307efff6609fcdbb65b2f01055187f47dfd2f0d7d9d6d372577753

Observation e54293de-0b29-48d9-924f-b7446fb7133d · outbound

This paper cites Sonic: Shifting Focus to Global Audio Perception in Portrait Animation.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Sonic: Shifting Focus to Global Audio Perception in Portrait Animation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:06.340492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:06.340492Z digest=sha256:2cbd2520a666c44343e29977326c130fc91d2ac0fcf0216778e99e58013ff9b0

Observation 80926fae-08ce-4f87-9008-0757494756b5 · outbound

This paper cites Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:06.431313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:06.431313Z digest=sha256:b86253e6bc6fe3e19dec4a4a942d495d8b6065ebe1daccf30e419c35fef2644b

Observation 893ea3b1-3f2f-4b33-880b-bca847ba4a81 · outbound

This paper cites Musiq: Multi-scale image quality transformer.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Musiq: Multi-scale image quality transformer

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:24:11.048008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T21:24:06.551175Z digest=sha256:6b115c88b691c6d8c537f76d0e0b9031c589b63906815dbed391fee0156dc9ff

Observation 994d9fab-d6d1-4b4d-8c98-01ac622ec4b9 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:06.688755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:06.688755Z digest=sha256:a71e3236b1b80cdaba03f37ad6b128912c31daa8df01e2d5eb37d2853af9a9c9

Observation 27d6c694-364a-4966-928b-b35d9c2e7396 · outbound

This paper cites Luma ray 2 video model.https : / / lumalabs.ai/ray.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Luma ray 2 video model.https : / / lumalabs.ai/ray

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:24:10.854778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T21:24:06.807169Z digest=sha256:70841b11745a33d827c043a9980415533828dceeb3ee96a94d738509209d8815

Observation 8bcc75af-195b-48d1-bd86-58844b0ad169 · outbound

This paper cites Laion aesthetic predictor.https : / / github.com/LAION-AI/aesthetic-predictor,.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Laion aesthetic predictor.https : / / github.com/LAION-AI/aesthetic-predictor,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:24:10.700967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T21:24:06.887819Z digest=sha256:d65ad56ef5277fb62f5dd2a71becda47d48932aea8ec81b119cccbaf24e640b5

Observation 98cb3211-2ac2-4a1d-8998-c10475b07b4b · outbound

This paper cites LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:07.088948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:07.088948Z digest=sha256:dfe33f4ae2c468d1b9812d7a23660816f154b873119fb29f4ed4be964859da86

Observation 57901d56-65ea-4a19-9d9f-23e6f2a6ce66 · outbound

This paper cites VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:07.155337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:07.155337Z digest=sha256:71d69ea377121747ee41415e5efaad5613c53e9d3194876440da21986b6bfb9a

Observation 57d2b5d0-0628-4f84-b681-3eafd24e774d · outbound

This paper cites Amt: All-pairs multi-field transforms for efficient frame interpolation.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Amt: All-pairs multi-field transforms for efficient frame interpolation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:24:10.374361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T21:24:07.257087Z digest=sha256:eb4c2e8ad5bf3dc9e6ec8c9f639112bb93c912baad060692dab5550c0c48821e

Observation 8685b4a7-33cc-4bf9-98ef-1e5cffc6b000 · outbound

This paper cites Echomimicv2: Towards striking, simplified, and semi-body human animation.arXiv preprint arXiv:2411.10061, 2024.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Echomimicv2: Towards striking, simplified, and semi-body human animation.arXiv preprint arXiv:2411.10061, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:07.313868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:07.313868Z digest=sha256:15eddd3d83d0daa1a7a525d8f2f12982591c72c5db4e0a187ae79d601890e29c

Observation c88e2b08-679f-473c-a951-557d39ae9b46 · outbound

This paper cites StyleTalker: One-shot Style-based Audio-driven Talking Head Video Generation.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos StyleTalker: One-shot Style-based Audio-driven Talking Head Video Generation

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:24:09.334830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T21:24:07.379667Z digest=sha256:8924e6d790d8d79b8a16d4997761b1f12f84024c8efad8c5a684ea8ff976f46e

Observation 9da70060-4f55-40ab-a22f-0f3d457aee67 · outbound

This paper cites Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:07.462449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:07.462449Z digest=sha256:b5ec75907939cb4ec27e444a20c4fc714e3f61067284676564ca2f88427bce7a

Observation b7948f18-721e-4235-828d-3ad5a26b8791 · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos A lip sync expert is all you need for speech to lip generation in the wild

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:24:10.148523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T21:24:07.587910Z digest=sha256:350e64c21609d70c142211d87b3dcfd09e857fa8e95143015e7b927ae468d509

Observation 67dfffa3-33ec-4b33-a529-05935706537a · outbound

This paper cites Versatile Multimodal Controls for Expressive Talking Human Animation.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Versatile Multimodal Controls for Expressive Talking Human Animation

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:24:09.176859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T21:24:07.672135Z digest=sha256:601d18ba47d4ec056e65c6a94cd8a546b1ec5d853a8670ec2b9fa4057afe41bd

Observation 5101bf6c-1045-4805-b9a0-3fa55900749e · outbound

This paper cites SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:07.747573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:07.747573Z digest=sha256:326b21eb8c2a66dc0e837f6a03da75d776e351f0d9639dd893489b0045b98b90

Observation 8f735ff8-6cea-431c-bd36-ad67c6eb923c · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Learning transferable visual models from natural language supervi- sion

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:07.813333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:07.813333Z digest=sha256:a94c23730bbee6318803814f73c85d737add785b4837e42b5309b641929be9c5

Observation aeca28fe-9986-42bb-a0a5-5681bdb3fc2b · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:07.908609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:07.908609Z digest=sha256:bb8e9e1578583e1952eeecfe1cfef8b4cd1ee5f6e9cde140abe6fb5f5eafbbf9

Observation cd24ca58-02cc-4f59-afc6-e10a8b08301d · outbound

This paper cites Diffused heads: Diffusion models beat gans on talking-face genera- tion.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Diffused heads: Diffusion models beat gans on talking-face genera- tion

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:24:09.922883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T21:24:07.986728Z digest=sha256:f86520e567f7e31b2e3685c04c52b42e3608dc34f985ddcd056d11ca7cdb75e9

Observation 5a0d1fb3-327d-4405-8d2e-8cd94feae4db · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Raft: Recurrent all-pairs field transforms for optical flow

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:08.082383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:08.082383Z digest=sha256:9a1d103eb66c0f318891275ef7741944687172bbab8f8bc0c35c3382424edcc4

Observation cf2d0577-8566-4b21-bb82-73ba931ad1ae · outbound

This paper cites StableAnimator: High-Quality Identity-Preserving Human Image Animation.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos StableAnimator: High-Quality Identity-Preserving Human Image Animation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:08.163993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:08.163993Z digest=sha256:02fe24a0adefa74f7dfde93c2bf32594a75da8874c94396414455086656a07a5

Observation 42b3d4fe-6134-4f02-b366-e5a92e2e859a · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:08.229848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:08.229848Z digest=sha256:201b89299eca256ca8fa1bb2bf82957c607ad38467b5aee309fa3e197d43eccb

Observation fc65f202-0682-40f7-ace3-c363d5aed292 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Wan: Open and Advanced Large-Scale Video Generative Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:08.292449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:08.292449Z digest=sha256:b37b0244627806e284b69cd6911d0e25b0b8980429dbcce923282a791a08f275

Observation 3182e70d-7162-457e-a16b-6345e95beead · outbound

This paper cites Lavie: High-quality video generation with cascaded latent diffusion models.International Journal of Computer Vision, pages 1–20, 2024.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Lavie: High-quality video generation with cascaded latent diffusion models.International Journal of Computer Vision, pages 1–20, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:24:09.761627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T21:24:08.373562Z digest=sha256:2ce3ad026c16691104c5aa4b183f9d71e1c0a9ef3fc23f32abade4a0313df877

Observation f3bcd8ed-ba27-4c22-85a8-309c428e58a6 · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4):600–612, 2004.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4):600–612, 2004

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:08.460168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:08.460168Z digest=sha256:2e2fa638a3645f2d23e00b7521bca37d6e212225997ba6911aec419190f0230a

Observation 6e2d9b1d-8293-4304-a815-28494b7920de · outbound

This paper cites Easyanimate: A high-performance long video generation method based on transformer architecture.arXiv preprint arXiv:2405.18991,.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Easyanimate: A high-performance long video generation method based on transformer architecture.arXiv preprint arXiv:2405.18991,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:08.543292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:08.543292Z digest=sha256:395d797827518c7ebc7dbf4c5959c9d5c6a59a02a2a0baf840e222d79036a595

Observation d6d1aa7b-b701-498d-aa61-9fe116bdff55 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:08.608636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:08.608636Z digest=sha256:c32c0f78f17088fe23aa72c69fb3cafea9cca9cb28e4b27b1f8b9e9c017bda2c

Observation adcbc9e8-b530-4c38-8aea-e0c81a9ad333 · outbound

This paper cites I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:08.669726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:08.669726Z digest=sha256:0daee0b382be447569ece3d14931f2373d3bb81fdba9e596d1ea6fe68fc8dd65

Observation d0689d79-1e84-4c1b-b2cf-a75876eef380 · outbound

This paper cites Pose-controllable talking face generation by implicitly modularized audio-visual rep- resentation.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Pose-controllable talking face generation by implicitly modularized audio-visual rep- resentation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:08.721929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:08.721929Z digest=sha256:a37c41441dbd64afd2f2e3d808183099ab0288906b7e28c93e133207d236045d

Observation 22c0ad9a-3c42-4827-90bf-7a15d8360f3d · outbound

This paper cites an unresolved cited work.

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos Unresolved cited work

Reference 2022

Resolution
parse uncertain
raw_fallback, observed 2026-08-05T21:24:10.530104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T21:24:06.990235Z digest=sha256:b4b125046b058efa04e071a767a0332e5492d5f73123c28c70df1964ea4a20db

Pith citing papers

No inbound Pith citation observations are available.