Pith. sign in

Paper Citation Record · LEDGER

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction

As of 20 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 4 inbound Pith citation observations for arXiv:2412.13187.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.13187 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:25:27.472156Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:04:30.302310Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T22:32:11.045924Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact3
  • verified fuzzy5
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0d3e6144-e524-4436-8027-aac64176c2e6 · outbound

This paper cites GPT-4 Technical Report.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.315027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.315027Z digest=sha256:ba41a05e51632e8d17cd748cac619e45d4906295ad33f3ea46cbef8e977d809a

Observation 6346ffee-166d-4ce6-ab59-9127fedfa295 · outbound

This paper cites Black, Danica Kragic, and Hedvig Kjellstr ¨om.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction Black, Danica Kragic, and Hedvig Kjellstr ¨om

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:25:28.202850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:25:27.347061Z digest=sha256:053a05e6d5c0346850239cf0e3368feecd5a68bd6a918033a892876d78ac4a41

Observation 9cce5b6c-6412-479e-9e93-5e5e52b30d7a · outbound

This paper cites We also conduct an ablation study on the zero-shot chain-of-thought (Wei et al., 2022; Kojima et al.,.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction We also conduct an ablation study on the zero-shot chain-of-thought (Wei et al., 2022; Kojima et al.,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:25:28.096673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:25:27.472156Z digest=sha256:d06ba53d30961094c25fe752a3f06b6679fd2fcce7475181de6cc78b89a6a7c2

Observation e7e6acfc-e4dd-421f-9f8b-6b9e7beb43ac · outbound

This paper cites Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:25:28.180273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:25:27.383823Z digest=sha256:7be944541620672c81a4efeb64858e070a895fa7905437a7c1906fee99bcc0ea

Observation 92c1b10a-8acc-4c17-9bce-f7faafe0f1b7 · outbound

This paper cites Learning a hierarchy of discriminative space-time neigh- borhood features for human action recognition.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction Learning a hierarchy of discriminative space-time neigh- borhood features for human action recognition

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:25:28.154917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:25:27.415090Z digest=sha256:1d0ff3f60aa3846e8568b6636fa2cc2b3d370006528f1c7e09b8d9198bfbc9e6

Observation 076257f0-de31-4348-a465-e026704f16b8 · outbound

This paper cites Forecasting human-object interaction: joint prediction of motor attention and actions in first person video.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction Forecasting human-object interaction: joint prediction of motor attention and actions in first person video

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:25:28.122154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:25:27.425157Z digest=sha256:a6275267d592112939a9e9dcb6104fdcd051f708d0d28e9af7f4af9e0a19ba50

Observation 96832fac-19f8-44d5-bfb5-4f9399d95154 · outbound

This paper cites Madiff: Motion-aware mamba diffusion models for hand trajectory prediction on egocentric videos.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction Madiff: Motion-aware mamba diffusion models for hand trajectory prediction on egocentric videos

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.431585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.431585Z digest=sha256:d53dc5c3ee3b1c16f7825f4b5e55f954d4d4ad82052a8feb33e3eb7238b9b326

Observation ca1d9f18-2315-4fd7-abbe-e261016efe07 · outbound

This paper cites R3M: A Universal Visual Representation for Robot Manipulation.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction R3M: A Universal Visual Representation for Robot Manipulation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.438085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.438085Z digest=sha256:b76e841b62bfce528ead0b394f0d42c1e80394f2ed04a32ff420f930940a4d21

Observation cfcc605b-ea90-4a26-8f5f-014e99cc8098 · outbound

This paper cites FrankMocap: Fast Monocular 3D Hand and Body Motion Capture by Regression and Integration.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction FrankMocap: Fast Monocular 3D Hand and Body Motion Capture by Regression and Integration

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.444539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.444539Z digest=sha256:55d0fbc2bb27c10e5b7bb2b97112388e7b286bc8fdbd55acc6966a03c2c710b4

Observation 0535b06f-adc5-45a9-9480-c72ded5d9914 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.450101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.450101Z digest=sha256:2b687a60bbdb294fb68825acf6ded8ebaefa4e4edd03fbeafde365db50886beb

Observation 1a71a723-d1f5-46f7-8e1c-f8af06a7f56f · outbound

This paper cites LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.455904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.455904Z digest=sha256:9ea3005df86bef0728fac6779d0c6d76631900a4914eab2cbde9584651a93000

Observation f1de86b4-749c-4e62-aea8-a3a61477779d · outbound

This paper cites NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.464809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.464809Z digest=sha256:572ddaea365dc23846d8e4fe2e8512b02749d1fb08e09b6bfed7d43451b65e74

Observation a89cbccd-29c9-442b-bd8e-955d67dff461 · outbound

This paper cites Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.330201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.330201Z digest=sha256:a5f419bb95125ff9f72a8d0fd6d132bff5455ca84b4c6d0229d85ca1f4d5eea0

Observation 61021fcc-290b-4904-9116-01d79faf9edf · outbound

This paper cites Pix2seq: A Language Modeling Framework for Object Detection.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction Pix2seq: A Language Modeling Framework for Object Detection

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.369987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.369987Z digest=sha256:dc7277be32bbb5cdef733efb8f51c8cb345b2f557dfca901cc1067b3eb44fc9c

Observation 6a9f8887-2b0d-4364-9118-47c4967f1d40 · outbound

This paper cites Judith B¨utepage, Hedvig Kjellstr¨om, and Danica Kragic.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction Judith B¨utepage, Hedvig Kjellstr¨om, and Danica Kragic

Reference 2017

Resolution
verified exact
doi, observed 2026-08-11T13:25:27.544923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:25:27.354684Z digest=sha256:592435423e86bcac063eae17222f72c70e72e00bf5fd043c3aaa60055b197140

Observation dfde9f0b-b4d6-455d-9079-bf0a30581932 · outbound

This paper cites an unresolved cited work.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction Unresolved cited work

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.361665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.361665Z digest=sha256:0da0b7c63c41beee83a40a6d0722e7834a9fbcbacfa29033ae99e3020aa17e32

Observation 60e9aecc-25c2-475e-84be-71171ad4ec03 · outbound

This paper cites LITA: Language Instructed Temporal-Localization Assistant.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction LITA: Language Instructed Temporal-Localization Assistant

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.399478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.399478Z digest=sha256:418d13e291e49207116259c5d413a6dbc16103e2802272a96f1d415dcb4d9c4d

Observation 621224e9-6890-40f9-94c7-cf596e8ae6f8 · outbound

This paper cites SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.377004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.377004Z digest=sha256:a78da3150ed662390bc15c7a8c367b530d2fb0135fd6ea5c9b626ddac28cada0

Observation 2e7e6801-9363-4d16-9c09-aa326ae7d8b2 · outbound

This paper cites Expressive Forecasting of 3D Whole-body Human Motions.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction Expressive Forecasting of 3D Whole-body Human Motions

Reference 2022

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:25:27.875796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:25:27.390474Z digest=sha256:9d5393884ecc60384a31c685002f61d64ab5cb6c5cf4ed09774ab831bdabab88

Observation 6413f722-0f89-42c6-98a4-9dc500c3ce26 · outbound

This paper cites Uncertainty-aware State Space Transformer for Egocentric 3D Hand Trajectory Forecasting.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction Uncertainty-aware State Space Transformer for Egocentric 3D Hand Trajectory Forecasting

Reference 2023

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:25:28.040507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:25:27.322124Z digest=sha256:37965c64f397a10b22ab7f9985aa27a04418aad0a2a1853e8e0241e40a2c49de

Observation b78afefe-1b21-4649-9f04-ad02675583b9 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction OpenVLA: An Open-Source Vision-Language-Action Model

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.407545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.407545Z digest=sha256:9d85e1afcecdabe0652c534b81de0cb2e09ea434e558f580fb67ed5d0409e0ff

Observation 635f13ad-a235-49ee-8cee-bfc845328e1a · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.337527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.337527Z digest=sha256:1f08ccf02e40508edfb1c57d305bebb528185ee3bffc8e406dee79a4d3566068

Pith citing papers

Observation d3e68d82-77d7-40aa-8952-8f19c7507f26 · inbound

MEgoHand: Multimodal Egocentric Hand-Object Interaction Motion Generation cites this paper.

MEgoHand: Multimodal Egocentric Hand-Object Interaction Motion Generation HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:04:30.302310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:04:30.302310Z digest=sha256:06c4c6b671a6b9ac12250834876fd19e21ebd61d66d3979d47c7e8df0e6b3565

Observation 8cea29c6-c471-4e35-83b3-4124dd3243cf · inbound

Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views cites this paper.

Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:32:11.048567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T22:31:13.391859Z digest=sha256:c5f5dccac68546812dd8054b1e2cfa447759fb02dfae68a604edfd4c66492edd

Observation 35dbf9cc-5449-4140-a493-34be081ab8b5 · inbound

EggHand: A Multimodal Foundation Model for Egocentric Hand Pose Forecasting cites this paper.

EggHand: A Multimodal Foundation Model for Egocentric Hand Pose Forecasting HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:57.582720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T02:11:17.739464Z digest=sha256:ee19e0c2b011b984755b55653bab6d00048bbad78ed25862fbd3533c5b0230c6

Observation 89482c55-4d31-4fb9-8ee6-08a30eaafcdd · inbound

MotionForesight: Re-purposing Video Models for Future 3D Scene-Flow Prediction cites this paper.

MotionForesight: Re-purposing Video Models for Future 3D Scene-Flow Prediction HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T21:08:06.701306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:08:06.701306Z digest=sha256:7e89c922e69f4100ad1577e7ea1291183c150265403238c191ab8dedaa48e8a9