Pith. sign in

Paper Citation Record · LEDGER

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation

As of 12 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 0 inbound Pith citation observations for arXiv:2412.04189.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04189 v5

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:45:10.247547Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

82 of 82 outbound references displayed

  • verified exact1
  • verified fuzzy42
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f642428d-8390-4af5-9e2e-5c12478cec54 · outbound

This paper cites The mug facial expression database.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation The mug facial expression database

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:09.848461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:09.848461Z digest=sha256:6a03578fa69239be21827191ea050b87f6151aebbeef19dd94c28376e91b0b09

Observation 3729c403-f6e6-4ebd-be85-0bfa87f6a5b4 · outbound

This paper cites Detours for navigating instructional videos.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Detours for navigating instructional videos

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:09.853490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:09.853490Z digest=sha256:178d049e62eedc661e6e486b68c9a0a54ff26c6247a5f81527da24091e93d1ea

Observation 0d056796-4471-4d98-8e3a-c9b3035531ba · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:09.857437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:09.857437Z digest=sha256:d32a3b5e8677f8c9bc7b80e92536887e4d6cbf6e27d390ba13be0b14a3519d1a

Observation ef9e9c3f-5e94-4e05-ae3d-8469909202c7 · outbound

This paper cites Hrs-bench: Holistic, reliable and scalable benchmark for text-to-image models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Hrs-bench: Holistic, reliable and scalable benchmark for text-to-image models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.512917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:09.860935Z digest=sha256:45bca960f6c522e2bc034dd45e382786cd6bb570fed0958a6afb262810fad237

Observation 081a8a1c-63a7-44b8-916d-1cb2173a61ee · outbound

This paper cites EditVal: Benchmarking Diffusion Based Text-Guided Image Editing Methods.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation EditVal: Benchmarking Diffusion Based Text-Guided Image Editing Methods

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:09.864718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:09.864718Z digest=sha256:3b2ef5cc9ed3b618513dbc30489bd47a55799806d78ae1f13b2aed4a67d62c60

Observation ea11565d-4198-43ef-8e98-275e2cda8b11 · outbound

This paper cites Brooks, A.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Brooks, A

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.496580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:09.868784Z digest=sha256:70458c50204ae7f87757b7bccf6c7b14ac4fda8169725d683ac47df79f4a958e

Observation 31a2ed05-e1eb-4f98-a350-6f0915c6de41 · outbound

This paper cites Gener- ating human motion in 3d scenes from text descriptions.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Gener- ating human motion in 3d scenes from text descriptions

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.483499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:09.873338Z digest=sha256:83a6e8b10af869f0eef0d48792b74d60a88d3f3a9cd5a3685784eef0c30b8659

Observation 164c07f7-1778-4547-9ded-584d90d528d5 · outbound

This paper cites Ceylan, C.-H.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Ceylan, C.-H

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.469388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:09.876682Z digest=sha256:85e56d836aaf411584e9082b24a5e925061e3601a2a4402f681dae7f1041c8e6

Observation 653a53e4-930a-455c-9b5c-2321bfceb89d · outbound

This paper cites Cognitive load theory and the format of instruction.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Cognitive load theory and the format of instruction

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.457298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:09.883210Z digest=sha256:ff5b48ee73f7f97f1617a5fb7dd0a27b3b925e9ca80b659208da27ca2083e2e8

Observation 2446aa98-4e4f-453a-84e2-436b1a2a8a33 · outbound

This paper cites Learning video-conditioned policies for unseen manipula- tion tasks.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Learning video-conditioned policies for unseen manipula- tion tasks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.439261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:09.887050Z digest=sha256:952e18517ef312f868569b9d2c5a1554487db9fcd31110df809f87334b8b736f

Observation 3bf0f642-0624-4c18-b885-de977ce05134 · outbound

This paper cites Cheikh Youssef, A.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Cheikh Youssef, A

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.425456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:09.892699Z digest=sha256:7654c5921ed297f683756818efb92fe6421e36945668d3ddf1a606fa7b9f2383

Observation f1dc980f-22a6-483d-b76c-4f6416783538 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffusion models, 2024.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Videocrafter2: Overcoming data limitations for high-quality video diffusion models, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:09.897586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:09.897586Z digest=sha256:570b606c607d2ce09f44d947d18a83380e4cb7fe6fc1965b92ae9351b542e168

Observation d422d952-13fb-484c-8976-fac8f8d821a2 · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:09.902375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:09.902375Z digest=sha256:77e7614397d3d9b0c8e0cb6583084238ab39f90c099739c510c6f7a9e2c93e0d

Observation b93ae10d-05e5-4231-aa50-6a683b3c535d · outbound

This paper cites AnimateAnything: Fine-Grained Open Domain Image Animation with Motion Guidance.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation AnimateAnything: Fine-Grained Open Domain Image Animation with Motion Guidance

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:09.906734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:09.906734Z digest=sha256:bacc0c73ef2f41adc358578a989a1c6c11d1f180c68bf5458244a6b7f832ce65

Observation a078f0a4-2dce-4c56-95c3-d08364ef75ad · outbound

This paper cites Rescaling egocentric vision: Collection, pipeline and chal- lenges for epic-kitchens-100.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Rescaling egocentric vision: Collection, pipeline and chal- lenges for epic-kitchens-100

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.384219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:09.914009Z digest=sha256:c600b889d767177380518d336723a67783e8bfe43561bbc8ba2688815569cfc9

Observation af9eabd8-b54e-4094-8547-fa15793678a9 · outbound

This paper cites Learning universal policies via text-guided video genera- tion.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Learning universal policies via text-guided video genera- tion

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.367378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:09.919008Z digest=sha256:c6379a334fd55325b293e8ddd0c2b0190210c5b5b4e1f7a1c5b0a26e89c251ce

Observation 66797e2b-787c-43e0-ba36-820b7eeb3cd5 · outbound

This paper cites Structure and content-guided video synthesis with diffusion models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Structure and content-guided video synthesis with diffusion models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:09.924083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:09.924083Z digest=sha256:78ffaa70c1eafe0307e73b15be493baf9f912f801c571ce0d3045c02a7b837d3

Observation c68ac1b6-f258-4cdc-a606-cdcb743c14bb · outbound

This paper cites Handrawer: Lever- aging spatial information to render realistic hands using a conditional diffusion model in single stage, 2025.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Handrawer: Lever- aging spatial information to render realistic hands using a conditional diffusion model in single stage, 2025

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.342122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:09.927920Z digest=sha256:f33c31c15ee66de65794bdf0c9298b40447a22b4cef3c3d69583431cb748c05d

Observation 5be0a0e2-f208-4406-9f00-0b9e5df9ebd9 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Ego4d: Around the world in 3,000 hours of egocentric video

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.328972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:09.932703Z digest=sha256:5228a7b54e4ff34088d00c1d3415c8d70e4bf53d83f9c489d603a5a435bea94b

Observation 300d4442-7dcd-420d-9ed7-4b57719550eb · outbound

This paper cites Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.315625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:09.937048Z digest=sha256:f993e41afb4cf6ff348959a76e1ee3e89dcfd6b808e9492dff9031e0a8ad5d42

Observation c172d233-926b-4cbf-be3a-8421f482e350 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:09.941469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:09.941469Z digest=sha256:51982488c9ebe8691aa8112a573c22a72846f8afab019c18ca5b6593bed6a4e4

Observation ec3f5388-06bb-44f9-8e5e-8ffcae4ad229 · outbound

This paper cites Denoising dif- fusion probabilistic models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Denoising dif- fusion probabilistic models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:09.947071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:09.947071Z digest=sha256:0ca8bf41b673cdc220039b6b17a54c0834342b07dd9a59fb44c310bb3121057a

Observation 0c3099ee-0f63-4835-b28c-08df527f70b3 · outbound

This paper cites Video dif- fusion models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Video dif- fusion models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.278384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:09.952593Z digest=sha256:c99bf2228bcaeb46fc52d968bafd830f417c5dcf14f7129dd1995328a0909e60

Observation a0aa761a-7ed0-40bf-9aec-66ccbc44ea45 · outbound

This paper cites T2i-compbench: A comprehensive bench- mark for open-world compositional text-to-image genera- tion.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation T2i-compbench: A comprehensive bench- mark for open-world compositional text-to-image genera- tion

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.261520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:09.957306Z digest=sha256:990cfe082f420b192ba962f8d969608b9a8bfe88b73ea75f5d3645cd7ad56494

Observation 5824e7d4-8696-4c6c-a8a4-8a053a525956 · outbound

This paper cites Vbench: Comprehensive bench- mark suite for video generative models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Vbench: Comprehensive bench- mark suite for video generative models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.246127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:09.961560Z digest=sha256:9977a1410a0ebb27fcae9850a868667e690bc2465fc28d5baafac36392c7166d

Observation 8f37f44a-9c29-430d-85bf-2d5840cbbf56 · outbound

This paper cites VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:09.965337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:09.965337Z digest=sha256:7ad4b767b0438a493aeaa5a4178577b4eb437981c9bb92f60e9c6b92a70cc274

Observation c817a730-3f91-4b55-8beb-f5f1a6508de5 · outbound

This paper cites Vid2robot: End-to- end video-conditioned policy learning with cross-attention transformers, 2024.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Vid2robot: End-to- end video-conditioned policy learning with cross-attention transformers, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.230494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:09.970349Z digest=sha256:2ef281790491a4f11cd5086aefa226b6573ed7376e0e2c83545bd5f308b4e6a2

Observation eb93ff5f-db30-46f1-9b59-6d6aebd881da · outbound

This paper cites LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:09.974729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:09.974729Z digest=sha256:02e8b1833e2567fbdabf336f17443c069e478ed3c3dac5df8a9ed49efad24765

Observation 680d6f3b-8e02-4a57-b672-610afca8998b · outbound

This paper cites Temporal convolutional networks for ac- tion segmentation and detection.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Temporal convolutional networks for ac- tion segmentation and detection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.217323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:09.979486Z digest=sha256:f36f3167a2f68994df9785b4d4b28c5144258389fdb144465b541023fa2b3de9

Observation e39195db-2bee-4e88-8eb2-6b62179a1336 · outbound

This paper cites Gradient-based learning applied to document recog- nition.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Gradient-based learning applied to document recog- nition

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.203636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:09.983851Z digest=sha256:ef3c5380b5f86ae7c8e5af254da4610f2c976e24242b45e64931b39943bf2169

Observation 7b190191-8e12-4d35-ae9e-5f8621dff8e6 · outbound

This paper cites Holis- tic evaluation of text-to-image models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Holis- tic evaluation of text-to-image models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:09.988588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:09.988588Z digest=sha256:87e8305cd4dedcc8c42b22d251eedd28834fda22718119ed0679f9a2371dea19

Observation ca687851-383c-4a8d-984f-15c92b7ccb87 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.175203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:09.993512Z digest=sha256:522fbff4d11da7f4588fdc25b188d64574a5478e7deaf69987e6ed8a74030ff9

Observation f94d3f16-8793-46ab-8654-9082073dac07 · outbound

This paper cites Egocentric video-language pretraining.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Egocentric video-language pretraining

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.160981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:09.997890Z digest=sha256:689740cf2cc2625a85e1803a65c6d5868d133c90572470fa94727f9be5f8c230

Observation 98c6eb90-b37a-4244-8bf4-5efb8029c341 · outbound

This paper cites Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.002519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.002519Z digest=sha256:82afbfb5b18cb1ce8dda8f15c578047651631fda25c188991c33fd1102b0381a

Observation 8d20e3a2-82c3-40c9-824d-5b1d9fec4651 · outbound

This paper cites DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.007032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.007032Z digest=sha256:e0dd9737f92e6523d0ca102c7bb03779cff3e9e47cb10e68aaf2e6c3663108ff

Observation 50bc3a62-629e-481a-8aa1-fbb9cc8b7afb · outbound

This paper cites Handrefiner: Refining malformed hands in generated images by diffusion-based conditional inpainting.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Handrefiner: Refining malformed hands in generated images by diffusion-based conditional inpainting

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.132941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:10.011270Z digest=sha256:2298c76ddabc35c02a8db48c6c9f81c44cd3f80663363bd1b92de604b8f2c3e5

Observation df996d96-a8ef-47e7-a982-d9058f18d572 · outbound

This paper cites MediaPipe: A Framework for Building Perception Pipelines.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation MediaPipe: A Framework for Building Perception Pipelines

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.017361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.017361Z digest=sha256:b3acd34d21e09702712e32d6fb3d8f445c8141d52783f2489083858969a7291c

Observation 49015ec3-e604-4884-8243-cf92068f8b29 · outbound

This paper cites Dexvip: Learning dexterous grasping with human hand pose priors from video.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Dexvip: Learning dexterous grasping with human hand pose priors from video

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.117322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:10.021924Z digest=sha256:e4983b7ee0e7bb12d3fae5a98c77b4f288bf1ab1fef0cf229c41e534cd29c53b

Observation bf88c734-6f0a-4a9f-aa4c-b6d1cf86dbce · outbound

This paper cites Recipe1M+: A Dataset for Learning Cross-Modal Embeddings for Cooking Recipes and Food Images.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Recipe1M+: A Dataset for Learning Cross-Modal Embeddings for Cooking Recipes and Food Images

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.025263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.025263Z digest=sha256:fdd3bb809861986019bd50ba0a21b7c8adc71b0d9638a73c551086ea386853f8

Observation 86664d79-ab7f-40d2-b01b-9fddadd615b9 · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.028678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.028678Z digest=sha256:0ab257fc4a719bd9b1cf8277402105b5e4a4ebec680558234a932b08a2bd4c89

Observation fa91b37d-daa1-4561-938b-4830fd64619d · outbound

This paper cites Han- diffuser: Text-to-image generation with realistic hand ap- pearances.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Han- diffuser: Text-to-image generation with realistic hand ap- pearances

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.095541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:10.032120Z digest=sha256:9f9d417724967c8fe9f4971849396a893151e4ac8e9dc2d158f60fc34bacaee5

Observation 0af6a45f-b2e8-4f19-9e17-be6b9a04bc55 · outbound

This paper cites Conditional image-to-video gener- ation with latent flow diffusion models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Conditional image-to-video gener- ation with latent flow diffusion models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.078521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:10.036469Z digest=sha256:b2f4b815b6edd0be44079d95a31432d504db68c1e4aef7d299f350b98e2b2a4a

Observation f6768da8-c1d6-4744-bbc8-cf575ce84005 · outbound

This paper cites Ti2v-zero: Zero-shot image condition- ing for text-to-video diffusion models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Ti2v-zero: Zero-shot image condition- ing for text-to-video diffusion models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.061746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:10.040254Z digest=sha256:402ea91e02958b8433c3cc9078fd5ec02116740b0f5f85e669886d4c1e0f1b09

Observation 78ee2f4e-1f22-4847-8472-d2e9bcb3e993 · outbound

This paper cites Improved denoising diffusion probabilistic models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Improved denoising diffusion probabilistic models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.043511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.043511Z digest=sha256:a4330ec01687c9f2b78bf676a10f4638dd2b30e7aa1f3ffbeaae10c66a17e7a9

Observation a4102e6b-ab42-4357-a7fb-06d1fcc38034 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Learning transferable visual models from natural language supervi- sion

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.046644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.046644Z digest=sha256:53a7c6fea670709cba5e72055050aa65f96879e000b77a4dfec6397e0a2be7c9

Observation e52dd587-8d4e-46dc-b45b-69cbee120526 · outbound

This paper cites The meccano dataset: Understanding human-object interactions from egocentric videos in an industrial-like domain.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation The meccano dataset: Understanding human-object interactions from egocentric videos in an industrial-like domain

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.050808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.050808Z digest=sha256:e01941e2fd1de8ec53435333bc2f95ba10ff15f6af8a9ef2d7f7ddb59c2b2e83

Observation ef77aa70-7a2d-415a-9c0f-a102233f68e6 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation High-resolution image synthesis with latent diffusion models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.991089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:10.055175Z digest=sha256:5f392e260b18c14dd3cf59912ef180edb4596196bc6df9ec5efc62dc736f88ab

Observation bc4e86e4-e2c6-4ed0-89e9-ecd873b53b03 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Photorealistic text-to-image diffusion models with deep language understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.058737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.058737Z digest=sha256:49b5c04e4aa1f365d28aeac053fb049899f09e36e71c5a269c5d7fd0425962ad

Observation 6c642cb2-8f57-4ffd-9513-3129aa1ebdb9 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.063091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.063091Z digest=sha256:e69108187a3f7db8c22a90de903eed70042b9ad7ad045e2731912789be768270

Observation 53feb2f4-ee4d-4333-943d-3dbf01ed72a6 · outbound

This paper cites As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.067305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.067305Z digest=sha256:fd18bbc99906ce5896a899006b16996b4d37224d84d71611b695be7405a7cce4

Observation db523071-473f-4964-9366-d7e124177089 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.071463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.071463Z digest=sha256:48d792593009ac6d18d16a038b3e31177e88564eb033ae78f1b6146c6327dc1b

Observation c4df0f93-d0e2-4c0e-b988-69f5664eb36d · outbound

This paper cites Many turn to youtube for children’s content, news, how-to lessons.Pew Research Center, 7, 2018.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Many turn to youtube for children’s content, news, how-to lessons.Pew Research Center, 7, 2018

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.959083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:10.079620Z digest=sha256:19c489c351b56afa286e9c6cd86e761de6865662a8f7b00d25d9dd59c01d988c

Observation 66cdd9d8-2910-4151-a238-2f0e5b04c7f3 · outbound

This paper cites Denoising Diffusion Implicit Models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Denoising Diffusion Implicit Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.084373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.084373Z digest=sha256:dc5c49e952379a43f334bf1117eeaa1c80c436aafc3925a8cf42e7983e54fe36

Observation a897df8d-f6e3-4e7a-8dfe-13e88d87d733 · outbound

This paper cites VideoAgent: Self-Improving Video Generation.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation VideoAgent: Self-Improving Video Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.088431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.088431Z digest=sha256:9b771f0e7ed363b97633bd16b902273fa4ed28acc95eee626836f7c2d731d7fc

Observation c95c5cdd-8d6b-448d-83a9-3522f8228f47 · outbound

This paper cites Genhowto: Learning to generate actions and state transformations from instructional videos.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Genhowto: Learning to generate actions and state transformations from instructional videos

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.944483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:10.092424Z digest=sha256:ead6828db48745bd30f1e169749d6e5e845adc8342adfe34693a0b6ffb91db8a

Observation c43752f6-ee06-4057-b6f4-c3d3a596b5f7 · outbound

This paper cites Fvd: A new metric for video generation.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Fvd: A new metric for video generation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.929599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:10.097624Z digest=sha256:822d5443263150cc47351418a5109959a5b5558f6ba322972f72fb8cc76209b2

Observation 1b24bb18-76bf-4fe2-9906-07aeb04a19af · outbound

This paper cites Long-term temporal convolutions for action recognition.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Long-term temporal convolutions for action recognition

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.911202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:10.101769Z digest=sha256:c9ef207094b9c39535070b1126fed9656bde4f9f7c96b4e31a9ada10dbfc80fc

Observation 31ee753b-b15c-464d-88bb-54425c944b60 · outbound

This paper cites Attention is all you need.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Attention is all you need

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.108198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.108198Z digest=sha256:cff128aeaf422bb857587b08393a2b6c7fa5dd190c06f77730c1d8dc2c6dba64

Observation 0b87055f-1e6c-4f93-8fb3-8744c03b4455 · outbound

This paper cites Imagen editor and editbench: Advancing and evaluating text-guided im- age inpainting.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Imagen editor and editbench: Advancing and evaluating text-guided im- age inpainting

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.113909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.113909Z digest=sha256:073e9199cdbf2a4eaef9f1891e24b2f141022d78dd5c6b6a85e522070a5103c0

Observation 57f8ec96-e070-4281-9377-522da41f2d4e · outbound

This paper cites Holoassist: an egocentric human interaction dataset for interactive ai assistants in the real world supple- mentary material.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Holoassist: an egocentric human interaction dataset for interactive ai assistants in the real world supple- mentary material

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.878103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:10.118962Z digest=sha256:b54f78daa71d16e5f418f27dafb69165f4ef2075b3610fc9ead60123e05997e3

Observation 31f1c55b-a807-4f1e-80fe-b475490bfc37 · outbound

This paper cites Multi- modal augmented-reality assembly guidance based on bare- hand interface.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Multi- modal augmented-reality assembly guidance based on bare- hand interface

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.859282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:10.123297Z digest=sha256:6e9ada2a0e79f7ab8008165b8b948938bf486169ecf10226ad24b10ce52dfd70

Observation 01d2e141-519f-4a36-bfba-906867de73c2 · outbound

This paper cites Holoassist: an egocen- tric human interaction dataset for interactive ai assistants in the real world.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Holoassist: an egocen- tric human interaction dataset for interactive ai assistants in the real world

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.127496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.127496Z digest=sha256:bb5df2b4d3b05876dd2b9228a5d2857137714060895ffaa8a5a4932f9b8db784

Observation d2f5d3fa-986b-484a-b3ef-a2dd0c6c8b6b · outbound

This paper cites Videocomposer: Compositional video synthesis with motion controllability.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Videocomposer: Compositional video synthesis with motion controllability

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.132365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.132365Z digest=sha256:98fadf72bba2e7efe1a8a1f06c6ec05d9f6976d189ee4bfe48ff811de96bf67f

Observation 6f5f1db1-2755-4779-a441-58553b189d1e · outbound

This paper cites EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.136640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.136640Z digest=sha256:0ba1f3cfa1c8ef8dde2ae538b39946a6ccead795181c19944462c7a945f7bb06

Observation 95626d99-19c3-4dac-a75d-04c1eacc7e6c · outbound

This paper cites Lavie: High-quality video generation with cascaded latent diffusion models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Lavie: High-quality video generation with cascaded latent diffusion models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.825113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:10.141927Z digest=sha256:954483a68c634fe8fd6b456b27dc5fe0ff38fdd39a8618391834c5835e23918f

Observation c3a86d49-5613-41b7-80cf-553565c35d36 · outbound

This paper cites Towards A Better Metric for Text-to-Video Generation.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Towards A Better Metric for Text-to-Video Generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.147322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.147322Z digest=sha256:d06b63aaae06e5d701bd8de9e4e7f54df712cbb9ecea8cc36992e12fc2b6bec0

Observation c2d1a377-f961-4b2d-9d36-2c0252e97de2 · outbound

This paper cites Freeinit: Bridging initialization gap in video dif- fusion models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Freeinit: Bridging initialization gap in video dif- fusion models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.809310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:10.152539Z digest=sha256:d66ec22075c6347f6a8ef9dab40728b8f19c702607886e68eb7c36695be3ec52

Observation a319beee-3af4-4f41-ba47-b358f3743d64 · outbound

This paper cites Dynamicrafter: Animating open-domain images with video diffusion priors.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Dynamicrafter: Animating open-domain images with video diffusion priors

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.791674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:10.156459Z digest=sha256:c8af526adb33bb2412259b3f624a1a59f8f4ecbb1422aca90e43acf6973a25a5

Observation 69cbdec1-6388-4ac8-b1f9-b99551f26838 · outbound

This paper cites X-gen: Ego-centric video prediction by watching exo-centric videos.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation X-gen: Ego-centric video prediction by watching exo-centric videos

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.775728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:10.161433Z digest=sha256:c43a76aff89cce1f5a5d69f333436c8178c5393f1b0a8f70dc79084af6b92632

Observation 1ed1cdca-47e3-4de5-aa48-68d8ba40aacb · outbound

This paper cites Ad- vancing high-resolution video-language representation with large-scale video transcriptions.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Ad- vancing high-resolution video-language representation with large-scale video transcriptions

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.761706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:10.166491Z digest=sha256:227342ed976bff8e0f0be41e863696b523ef34dde85c551d030607e3dd40f1d3

Observation f616454d-f686-4b37-812b-d70680323439 · outbound

This paper cites Stat: Spatial-temporal attention mechanism for video cap- 11 tioning.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Stat: Spatial-temporal attention mechanism for video cap- 11 tioning

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.739904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:10.172257Z digest=sha256:9652baf5a281b9a9a6dca157fd466e7c8a7357600cc57a2be90bbb54db32feb8

Observation 7fe80857-1b75-4ccd-bb03-8787d860ba96 · outbound

This paper cites Learning Interactive Real-World Simulators.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Learning Interactive Real-World Simulators

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.177014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.177014Z digest=sha256:1783877a447cb113fae386b991556cd230a16ab071287378277a3709294480dd

Observation d4d2ce0d-5e36-412d-9e67-848dc9fd2a39 · outbound

This paper cites Annotated Hands for Generative Models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Annotated Hands for Generative Models

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-08-11T21:45:10.372031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:10.184737Z digest=sha256:2a69bf2e4e5aa22d5ff19dde1b3e828b308bad9e760ec42bca14b862537cf5e3

Observation 3b0d8d95-bca7-4ccd-9b57-b725e623dc4f · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.192103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.192103Z digest=sha256:04c6c9eab5672c23ba3291a83c2d2bb435b49b258c6fcaceeb109490f2c29ddf

Observation eb9af77c-8673-4a21-9319-b37c2a7c5d11 · outbound

This paper cites Learning universal policies via text-guided video generation.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Learning universal policies via text-guided video generation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.718568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:10.200996Z digest=sha256:97275427bf0db7374f04f84e04e3c2815f478dbb2f05fcb7439a3d326f7feb78

Observation 833fe562-249e-434f-8fb5-ff90f3d117ad · outbound

This paper cites DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.209618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.209618Z digest=sha256:c9854f5b3325611afe9e31d35e280e0b6cef18ffe9196ba106fd9520c12258e0

Observation 34151f3a-bd05-4fbf-9a39-1978b2539da7 · outbound

This paper cites an unresolved cited work.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:45:10.703802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:10.215010Z digest=sha256:84129c2c282ceeeffee609cb02c2e3b4c839632b9c6746c932080b99f4a65433

Observation 1800aa40-4ab1-4cab-85ee-30f1711d8c9e · outbound

This paper cites MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.224664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.224664Z digest=sha256:5987bd3e53001a2c6dcc42c27254e477fa73427d2eab9c0d8e9d111640f50dd8

Observation def0720d-1a4a-4ebb-835e-c95ad5a8c746 · outbound

This paper cites Pia: Your personalized image animator via plug-and-play modules in text-to-image models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Pia: Your personalized image animator via plug-and-play modules in text-to-image models

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.683823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:10.231845Z digest=sha256:9e9e9c2bcebc68ba9068f6116d48a38206eb6368ae5e049e54a47dc1daec7ec0

Observation 503325e1-b18b-4b60-b170-f8db274132fb · outbound

This paper cites Open-sora: Democratizing efficient video production for all, 2024.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Open-sora: Democratizing efficient video production for all, 2024

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.668299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:10.237158Z digest=sha256:edb696f2f37072fe42cda37cad7e467cd823f76f6cc24d8802f90f89c64759e1

Observation 180a3c7e-14d0-45b2-bee9-7b78ef2a8c7f · outbound

This paper cites Towards automatic learning of procedures from web instructional videos.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Towards automatic learning of procedures from web instructional videos

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.645763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:45:10.243509Z digest=sha256:7bd5267fcecb9cb2e83dc0ed419d5cc1b37b6df0ff3f8f7b0b16b7cc85c480ed

Observation e119dd37-e116-46a1-90e5-68fbe57e7a6b · outbound

This paper cites Motion Control for Enhanced Complex Action Video Generation.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Motion Control for Enhanced Complex Action Video Generation

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.247547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.247547Z digest=sha256:9a072571d6c12fae40e4be6208811e6eedce70fcac04139d6bc2fd92745c6e5d

Pith citing papers

No inbound Pith citation observations are available.