Pith. sign in

Paper Citation Record · LEDGER

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction

As of 20 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 4 inbound Pith citation observations for arXiv:2412.13187.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.13187 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:25:27.472156Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:04:30.302310Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T22:32:11.045924Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact3
  • verified fuzzy5
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0d3e6144-e524-4436-8027-aac64176c2e6 · outbound

This paper cites GPT-4 Technical Report.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.315027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.315027Z digest=sha256:ba41a05e51632e8d17cd748cac619e45d4906295ad33f3ea46cbef8e977d809a

Observation 6346ffee-166d-4ce6-ab59-9127fedfa295 · outbound

This paper cites Black, Danica Kragic, and Hedvig Kjellstr ¨om.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction Black, Danica Kragic, and Hedvig Kjellstr ¨om

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:25:28.202850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T13:25:27.347061Z digest=sha256:248977f977b6957641437f5195a16d4bf8d6562989d5e6784166c4443de859ef

Observation 9cce5b6c-6412-479e-9e93-5e5e52b30d7a · outbound

This paper cites We also conduct an ablation study on the zero-shot chain-of-thought (Wei et al., 2022; Kojima et al.,.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction We also conduct an ablation study on the zero-shot chain-of-thought (Wei et al., 2022; Kojima et al.,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:25:28.096673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T13:25:27.472156Z digest=sha256:4cadd596af158e7da5ace7af7f33b78ffd486db3b3d85379088c81dc685a1136

Observation e7e6acfc-e4dd-421f-9f8b-6b9e7beb43ac · outbound

This paper cites Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:25:28.180273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T13:25:27.383823Z digest=sha256:1308f810ecad9e0ede029a515e3e59dfd29a2acf02a66853de7b305ad4456f45

Observation 92c1b10a-8acc-4c17-9bce-f7faafe0f1b7 · outbound

This paper cites Learning a hierarchy of discriminative space-time neigh- borhood features for human action recognition.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction Learning a hierarchy of discriminative space-time neigh- borhood features for human action recognition

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:25:28.154917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T13:25:27.415090Z digest=sha256:575eb0f4935146dfee9bcf0f6977da6235fcbd399f1b29cdd632e4219de05cc1

Observation 076257f0-de31-4348-a465-e026704f16b8 · outbound

This paper cites Forecasting human-object interaction: joint prediction of motor attention and actions in first person video.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction Forecasting human-object interaction: joint prediction of motor attention and actions in first person video

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:25:28.122154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T13:25:27.425157Z digest=sha256:f8c368942a2519df672649505d8b04bdf1d181570e7abe2954687e34834eaed3

Observation 96832fac-19f8-44d5-bfb5-4f9399d95154 · outbound

This paper cites Madiff: Motion-aware mamba diffusion models for hand trajectory prediction on egocentric videos.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction Madiff: Motion-aware mamba diffusion models for hand trajectory prediction on egocentric videos

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.431585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.431585Z digest=sha256:d53dc5c3ee3b1c16f7825f4b5e55f954d4d4ad82052a8feb33e3eb7238b9b326

Observation ca1d9f18-2315-4fd7-abbe-e261016efe07 · outbound

This paper cites R3M: A Universal Visual Representation for Robot Manipulation.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction R3M: A Universal Visual Representation for Robot Manipulation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.438085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.438085Z digest=sha256:b76e841b62bfce528ead0b394f0d42c1e80394f2ed04a32ff420f930940a4d21

Observation cfcc605b-ea90-4a26-8f5f-014e99cc8098 · outbound

This paper cites FrankMocap: Fast Monocular 3D Hand and Body Motion Capture by Regression and Integration.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction FrankMocap: Fast Monocular 3D Hand and Body Motion Capture by Regression and Integration

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.444539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.444539Z digest=sha256:55d0fbc2bb27c10e5b7bb2b97112388e7b286bc8fdbd55acc6966a03c2c710b4

Observation 0535b06f-adc5-45a9-9480-c72ded5d9914 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.450101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.450101Z digest=sha256:2b687a60bbdb294fb68825acf6ded8ebaefa4e4edd03fbeafde365db50886beb

Observation 1a71a723-d1f5-46f7-8e1c-f8af06a7f56f · outbound

This paper cites LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.455904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.455904Z digest=sha256:9ea3005df86bef0728fac6779d0c6d76631900a4914eab2cbde9584651a93000

Observation f1de86b4-749c-4e62-aea8-a3a61477779d · outbound

This paper cites NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.464809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.464809Z digest=sha256:572ddaea365dc23846d8e4fe2e8512b02749d1fb08e09b6bfed7d43451b65e74

Observation a89cbccd-29c9-442b-bd8e-955d67dff461 · outbound

This paper cites Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.330201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.330201Z digest=sha256:a5f419bb95125ff9f72a8d0fd6d132bff5455ca84b4c6d0229d85ca1f4d5eea0

Observation 61021fcc-290b-4904-9116-01d79faf9edf · outbound

This paper cites Pix2seq: A Language Modeling Framework for Object Detection.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction Pix2seq: A Language Modeling Framework for Object Detection

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.369987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.369987Z digest=sha256:dc7277be32bbb5cdef733efb8f51c8cb345b2f557dfca901cc1067b3eb44fc9c

Observation 6a9f8887-2b0d-4364-9118-47c4967f1d40 · outbound

This paper cites Judith B¨utepage, Hedvig Kjellstr¨om, and Danica Kragic.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction Judith B¨utepage, Hedvig Kjellstr¨om, and Danica Kragic

Reference 2017

Resolution
verified exact
doi, observed 2026-08-11T13:25:27.544923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T13:25:27.354684Z digest=sha256:3bc8eb0ad5b6bb6fb57181925c6790f9720737fac51a771e62eecd95bbfcc9c4

Observation dfde9f0b-b4d6-455d-9079-bf0a30581932 · outbound

This paper cites an unresolved cited work.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction Unresolved cited work

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.361665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.361665Z digest=sha256:0da0b7c63c41beee83a40a6d0722e7834a9fbcbacfa29033ae99e3020aa17e32

Observation 60e9aecc-25c2-475e-84be-71171ad4ec03 · outbound

This paper cites LITA: Language Instructed Temporal-Localization Assistant.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction LITA: Language Instructed Temporal-Localization Assistant

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.399478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.399478Z digest=sha256:717529fc3756dcea0a56e8d4ccce7303c31b16bd0812f38c59380cad061322b7

Observation 621224e9-6890-40f9-94c7-cf596e8ae6f8 · outbound

This paper cites SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.377004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.377004Z digest=sha256:a78da3150ed662390bc15c7a8c367b530d2fb0135fd6ea5c9b626ddac28cada0

Observation 2e7e6801-9363-4d16-9c09-aa326ae7d8b2 · outbound

This paper cites Expressive Forecasting of 3D Whole-body Human Motions.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction Expressive Forecasting of 3D Whole-body Human Motions

Reference 2022

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:25:27.875796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T13:25:27.390474Z digest=sha256:0b7869c43cd0ff25525611f57b26587afcae15a64d537064088a20bc8cc1a4cb

Observation 6413f722-0f89-42c6-98a4-9dc500c3ce26 · outbound

This paper cites Uncertainty-aware State Space Transformer for Egocentric 3D Hand Trajectory Forecasting.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction Uncertainty-aware State Space Transformer for Egocentric 3D Hand Trajectory Forecasting

Reference 2023

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:25:28.040507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T13:25:27.322124Z digest=sha256:3abc3190fc48c0abd796a8bae650411e6c2f99c0adf97f7dfd38163f018a87ef

Observation b78afefe-1b21-4649-9f04-ad02675583b9 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction OpenVLA: An Open-Source Vision-Language-Action Model

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.407545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.407545Z digest=sha256:9d85e1afcecdabe0652c534b81de0cb2e09ea434e558f580fb67ed5d0409e0ff

Observation 635f13ad-a235-49ee-8cee-bfc845328e1a · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.337527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.337527Z digest=sha256:1f08ccf02e40508edfb1c57d305bebb528185ee3bffc8e406dee79a4d3566068

Pith citing papers

Observation d3e68d82-77d7-40aa-8952-8f19c7507f26 · inbound

MEgoHand: Multimodal Egocentric Hand-Object Interaction Motion Generation cites this paper.

MEgoHand: Multimodal Egocentric Hand-Object Interaction Motion Generation HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:04:30.302310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:04:30.302310Z digest=sha256:06c4c6b671a6b9ac12250834876fd19e21ebd61d66d3979d47c7e8df0e6b3565

Observation 8cea29c6-c471-4e35-83b3-4124dd3243cf · inbound

Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views cites this paper.

Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:32:11.048567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T22:31:13.391859Z digest=sha256:03a0ae2ebba6c0c3143e1c6f60e343326f3e7b232aa08f48dd9e4dd5829abe2a

Observation 35dbf9cc-5449-4140-a493-34be081ab8b5 · inbound

EggHand: A Multimodal Foundation Model for Egocentric Hand Pose Forecasting cites this paper.

EggHand: A Multimodal Foundation Model for Egocentric Hand Pose Forecasting HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:57.582720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T02:11:17.739464Z digest=sha256:0c96535879c7542ca36843bc123317ebd22c3d111be7980cc83fd52011af2a6d

Observation 89482c55-4d31-4fb9-8ee6-08a30eaafcdd · inbound

MotionForesight: Re-purposing Video Models for Future 3D Scene-Flow Prediction cites this paper.

MotionForesight: Re-purposing Video Models for Future 3D Scene-Flow Prediction HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T21:08:06.701306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:08:06.701306Z digest=sha256:7e89c922e69f4100ad1577e7ea1291183c150265403238c191ab8dedaa48e8a9