Pith. sign in

Paper Citation Record · LEDGER

Co$^{3}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion

As of 19 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2505.01746.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.01746 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:14:50.968563Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1d1185b4-7b7e-4cc5-a354-db54cf3f9f87 · outbound

This paper cites Style transfer for co-speech gesture animation: A multi-speaker conditional-mixture approach.

Co$^{3}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion Style transfer for co-speech gesture animation: A multi-speaker conditional-mixture approach

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:14:51.182659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:14:50.910118Z digest=sha256:27825649efacb3cf3eeb9eb4054ea4ab13b61a5ed157d8fb8a0bdc4b8a55fb88

Observation fadd2ae3-1a55-4516-9fe1-b002f5832ef3 · outbound

This paper cites The genea challenge 2023: A large-scale evaluation of gesture generation models in monadic and dyadic settings.

Co$^{3}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion The genea challenge 2023: A large-scale evaluation of gesture generation models in monadic and dyadic settings

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:14:51.143260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:14:50.925954Z digest=sha256:c674bd9b68608c2444ffbfed8d9cd80fe959700c09b040be0046492cd9e36f0c

Observation 63b94378-b1f3-4303-816f-ab0c8e74c46f · outbound

This paper cites Gesture controllers.

Co$^{3}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion Gesture controllers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:14:51.131367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:14:50.930150Z digest=sha256:f52c6dfcce3a44ad227526c15628165a1dbb6d76a898d97b625fa34005b84967

Observation e5492f55-4775-429d-8903-8bcc30ad9b70 · outbound

This paper cites Montreal forced aligner: Trainable text-speech alignment using kaldi.

Co$^{3}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion Montreal forced aligner: Trainable text-speech alignment using kaldi

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:14:51.109386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:14:50.940078Z digest=sha256:bfc7c798181491ee667ad75ec2704fdd6482581e320986005623670ccff57c7c

Observation 59c0ae9e-3f28-41c9-b293-eafeeec51142 · outbound

This paper cites CoCoGesture: Toward Coherent Co-speech 3D Gesture Generation in the Wild.

Co$^{3}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion CoCoGesture: Toward Coherent Co-speech 3D Gesture Generation in the Wild

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T04:14:50.945893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:14:50.945893Z digest=sha256:be8aa0e30e5ff1d915363fdb0cef30f90ad23fb2f126db212c536e97a27216ee

Observation 542bc9f3-d96f-4de1-85d4-d7965cf0482a · outbound

This paper cites MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model.

Co$^{3}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T04:14:50.955939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:14:50.955939Z digest=sha256:0bae8916336099c375f5c5b80d2cdd19988a8f93179fad71ec8e845c2d9ec2ae

Observation 437d6244-2501-4eb7-b2ac-17a5805e3073 · outbound

This paper cites an unresolved cited work.

Co$^{3}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:14:51.072370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:14:50.959341Z digest=sha256:68e2cf7c51d0001b5d79aabd48f7097cd65c1e9738570ed85a234801e2d39671

Observation 7a4c3f1c-a0fd-441d-90f4-7cb345104221 · outbound

This paper cites They are then processed using automated methods to extract both audio and motion information.

Co$^{3}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion They are then processed using automated methods to extract both audio and motion information

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:14:51.061522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:14:50.962185Z digest=sha256:e7ea80483dfb05c45927baf6277d8eba3c0c5a3c94a317cf5de2b8c7127e1c73

Observation e78d77d6-f7bd-42dd-bced-d0d1d69dd853 · outbound

This paper cites Due to the strict keyword selection in raw video crawling, our dataset rarely contains two speakers standing or walking around.

Co$^{3}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion Due to the strict keyword selection in raw video crawling, our dataset rarely contains two speakers standing or walking around

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:14:51.050724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:14:50.965433Z digest=sha256:b19a30275cdceada6041c554da89452f6c8f33244d61908e1b4c78e3dfa46165

Observation 77cfc049-5c37-4b59-98f8-bb052028d8e2 · outbound

This paper cites Meanwhile, our method displays a much lower standard deviation than InterX and InterGen.

Co$^{3}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion Meanwhile, our method displays a much lower standard deviation than InterX and InterGen

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:14:51.037920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:14:50.968563Z digest=sha256:4172fcf5875d9e92f5c4a2f6601a23a8005556bc62c97906f3ed553b01c0de0f

Observation 4760762d-83f2-4e32-bc81-37d19cff6283 · outbound

This paper cites Learning to Generate Diverse Dance Motions with Transformer.

Co$^{3}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion Learning to Generate Diverse Dance Motions with Transformer

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-16T04:14:50.933480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:14:50.933480Z digest=sha256:0b771023d7b95ca422635051afab26ebe5a33a0bea55a796071bb069bc184847

Observation 48b2a685-0913-463c-8c0a-16652ca573f0 · outbound

This paper cites Understanding embodied reference with touch-line transformer.

Co$^{3}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion Understanding embodied reference with touch-line transformer

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:14:51.120353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:14:50.936901Z digest=sha256:9079d09ad84d88ac0d69adca8f05f4af0ab1685115b9d59d71d38caa6e0239ee

Observation 0a684c7c-47a9-45e8-be4b-324be00d72e2 · outbound

This paper cites Pyannote.

Co$^{3}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion Pyannote

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:14:51.164489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:14:50.917816Z digest=sha256:7dfb75f78006b27e81d03a19b8af67884d71ff2680e838d1275948b2d30e7d71

Observation a08e4497-bc55-46a0-b620-87f7c7ca4ca2 · outbound

This paper cites Whisperx: Time-accurate speech transcription of long-form audio.

Co$^{3}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion Whisperx: Time-accurate speech transcription of long-form audio

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-16T04:14:50.913998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:14:50.913998Z digest=sha256:625db4b95485114e16e9409bb5255db9dcf78e0f4f92ab73635dd3797845e229

Observation d585e588-b83a-4bb6-a22b-ad4762936c79 · outbound

This paper cites Generative agents: Interactive simulacra of human behavior.

Co$^{3}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion Generative agents: Interactive simulacra of human behavior

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:14:51.095730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:14:50.942891Z digest=sha256:c1324d7b72efc13584c34ae163d32b021513f88720203eedc1b8683b18192463

Observation 5bc23898-1dab-435d-9385-f0fe161120f9 · outbound

This paper cites Denoising Diffusion Implicit Models.

Co$^{3}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion Denoising Diffusion Implicit Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-16T04:14:50.949196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:14:50.949196Z digest=sha256:e54b9af475aefc3792a6e744cd25e49f0866bc0d1a369063ef9a3b221df3c075

Observation 0bb5b99d-1708-4ee2-99f0-e9937d56f5f4 · outbound

This paper cites Alpacafarm: A simulation framework for methods that learn from human feedback.

Co$^{3}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion Alpacafarm: A simulation framework for methods that learn from human feedback

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:14:51.153929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:14:50.921058Z digest=sha256:c0c3dbf86af558558c367119664b7e1d3134b409fbf94e25b4681479e23d0e07

Observation 4773e41b-6694-41f5-9460-5fd2b8caa360 · outbound

This paper cites Diffusestylegesture: stylized audio-driven co-speech gesture generation with diffusion models.

Co$^{3}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion Diffusestylegesture: stylized audio-driven co-speech gesture generation with diffusion models

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:14:51.083665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:14:50.952774Z digest=sha256:39f011211215e4eb7e5484b50d3360f379c6aa5094c2480bc2dec0506c89dc83

Pith citing papers

No inbound Pith citation observations are available.