Pith. sign in

Paper Citation Record · LEDGER

Human Action CLIPs: Detecting AI-generated Human Motion

As of 13 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 3 inbound Pith citation observations for arXiv:2412.00526.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00526 v2

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T05:23:20.684103Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T21:12:23.019873Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T15:13:25.052578Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 85c9dbce-8aa6-46ec-9002-eca253eaa487 · outbound

This paper cites People are poorly equipped to detect AI-powered voice clones.

Human Action CLIPs: Detecting AI-generated Human Motion People are poorly equipped to detect AI-powered voice clones

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.615674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.615674Z digest=sha256:dce427fd58339022ac0a8cbcf8224fe954aefc224a872b07fbc9bf375ee12c31

Observation 9f07d460-e61c-4e4d-bdc7-9dcf7f770f24 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Human Action CLIPs: Detecting AI-generated Human Motion Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.620799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.620799Z digest=sha256:ce10a308bbb7218dfd9d649b8ec2e291f769340bd0fad7bdff16333a8fedf731

Observation ed0b7a9b-2803-4c74-a8e5-10033bf0e925 · outbound

This paper cites Playing for 3D Human Recovery.

Human Action CLIPs: Detecting AI-generated Human Motion Playing for 3D Human Recovery

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.625984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.625984Z digest=sha256:73f8c13088a11bec4f14d1554ed1be320285c43814e4ac314eab6791a934aebe

Observation fd524c48-284d-4e88-8b0d-6c44be0146b5 · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

Human Action CLIPs: Detecting AI-generated Human Motion PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.631282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.631282Z digest=sha256:206110c079cd3f6af989a5ed0b9e6bd4ab14141a24e6a13058addffb996e145d

Observation 9cdfbb66-0a0b-4d1b-a09b-b090c93b4404 · outbound

This paper cites Exposing Lip-syncing Deepfakes from Mouth Inconsistencies.

Human Action CLIPs: Detecting AI-generated Human Motion Exposing Lip-syncing Deepfakes from Mouth Inconsistencies

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.636834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.636834Z digest=sha256:785a343fc3f620538ab24ac70b091baec87fdaa2a9f6e9e88c9a689b6c7baeb4

Observation b1a2d967-7dcc-487f-a665-7a7c303f991f · outbound

This paper cites Exploring the Adversarial Robustness of CLIP for AI-generated Image Detection.

Human Action CLIPs: Detecting AI-generated Human Motion Exploring the Adversarial Robustness of CLIP for AI-generated Image Detection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.641475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.641475Z digest=sha256:ea7d9c4a74aedd915246bcaa7ada5950aded349eb1982a42c495834f22772cf7

Observation 6442f9c9-2048-40d6-9bda-349bc54598f1 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

Human Action CLIPs: Detecting AI-generated Human Motion VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.646151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.646151Z digest=sha256:6d1d1aa9bd101d0902b29d3329e8dfa15578af93472935efa09c3e93c7289ebb

Observation 71d4e8e1-5848-4744-aeb1-65df39c7d66b · outbound

This paper cites AnimateDiff-Lightning: Cross-Model Diffusion Distillation.

Human Action CLIPs: Detecting AI-generated Human Motion AnimateDiff-Lightning: Cross-Model Diffusion Distillation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.655486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.655486Z digest=sha256:c6c039bd2375152dcadb14cda9eceb85702566a9a7ef2c265e60d62514e22b54

Observation 1f938311-0eef-4ee4-86e1-db65f2fa704b · outbound

This paper cites DeCLIP: Decoding CLIP representations for deepfake localization.

Human Action CLIPs: Detecting AI-generated Human Motion DeCLIP: Decoding CLIP representations for deepfake localization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.670086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.670086Z digest=sha256:777c18667ff2fe46b11bf36defbb5d758978d7ae00c7bfd3c62c884d2599be9e

Observation 15b8295a-9ff7-47af-bd7f-be6e5e19c222 · outbound

This paper cites CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment.

Human Action CLIPs: Detecting AI-generated Human Motion CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.674694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.674694Z digest=sha256:b4bac256f05d70305d2abf415358f87700bc4cce2d35d736a30668220c27be5c

Observation 6e7baf02-b2af-48f4-8b46-e556881897fa · outbound

This paper cites 4Real: Towards Photorealistic 4D Scene Generation via Video Diffusion Models.

Human Action CLIPs: Detecting AI-generated Human Motion 4Real: Towards Photorealistic 4D Scene Generation via Video Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.684103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.684103Z digest=sha256:e1642ca75f64037041fc01cf3bb2f292932d8636c3adee4c02ddc4092999aaf3

Observation 81aa69e6-e2e5-4fa3-baf2-59251b8623b2 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Human Action CLIPs: Detecting AI-generated Human Motion CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.679515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.679515Z digest=sha256:8c4482a5386a503519f8b6198c53ce3a0790fd41a9937c46ae8abefd8ddc100e

Observation 2931710b-9f08-4443-9809-b6230cff4857 · outbound

This paper cites Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models.

Human Action CLIPs: Detecting AI-generated Human Motion Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.604699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.604699Z digest=sha256:175f4c74ae3af0158c3da9d85db36fd7e1e86de3decedb0dbf34ce3a48dafb91

Observation 0de5dad8-3326-4922-b73d-f82e5443a8b0 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Human Action CLIPs: Detecting AI-generated Human Motion LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.660085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.660085Z digest=sha256:5615092406c577d1cee3eea3afa2c9776e78523c22142c5d6613c25c70d121d7

Observation d0fa75b4-7dfa-451b-bcfd-ee04f69247cc · outbound

This paper cites How Much Can CLIP Benefit Vision-and-Language Tasks?.

Human Action CLIPs: Detecting AI-generated Human Motion How Much Can CLIP Benefit Vision-and-Language Tasks?

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.665056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.665056Z digest=sha256:b5cd256519155439a67f3e7a068b42369273b3c36a1579de8a58c29a3abdde14

Observation 3997a9ce-1f6b-45bc-a7c4-e62e457061ee · outbound

This paper cites Jina CLIP: Your CLIP Model Is Also Your Text Retriever.

Human Action CLIPs: Detecting AI-generated Human Motion Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.650879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.650879Z digest=sha256:cbf1407ee501acb6ae34db4623342df915d804f345f78c7d81469e27109efacc

Observation 993eeb03-e35c-4bb5-8cb1-4e4ccf310ffc · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

Human Action CLIPs: Detecting AI-generated Human Motion Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.610451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.610451Z digest=sha256:cc0323d19bb57a665f4bb960628b3dfa65f8d7c3073f64286341bafc07f7c83c

Pith citing papers

Observation d2dd6fc4-7569-4714-bdf3-b51737b727bd · inbound

Survey on AI-Generated Media Detection: From Non-MLLM to MLLM cites this paper.

Survey on AI-Generated Media Detection: From Non-MLLM to MLLM Human Action CLIPs: Detecting AI-generated Human Motion

Reference 141

Resolution
unresolved
no resolver link, observed 2026-08-08T21:12:23.019873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:12:23.019873Z digest=sha256:571a347c4f6d86604f353b7210381d55fb42721dc4a54011fcb059a88fd30905

Observation 1fac0a74-999a-4860-9013-1b5e30fddcdd · inbound

SynthForensics: Benchmarking and Evaluating People-Centric Synthetic Video Deepfakes cites this paper.

SynthForensics: Benchmarking and Evaluating People-Centric Synthetic Video Deepfakes Human Action CLIPs: Detecting AI-generated Human Motion

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:37:33.016331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T07:32:30.580643Z digest=sha256:54777a02bd76d64b4ec7f66c33b6ea3a9440e572d19ebf2d655c04f03123c792

Observation 23b91fb0-a20f-4881-a85c-ab2c1b4876da · inbound

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection cites this paper.

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection Human Action CLIPs: Detecting AI-generated Human Motion

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:13:25.054227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T15:08:25.309094Z digest=sha256:3c2f3445f59f13bbe4bd9947870e00c68a0097ffb9c64e4ed7585e79163d1da0