Pith. sign in

Paper Citation Record · LEDGER

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception

As of 19 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 1 inbound Pith citation observation for arXiv:2501.00510.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00510 v2

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:53:06.792201Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T12:52:15.138790Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T12:53:17.381905Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact2
  • verified fuzzy11
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 666a7316-de87-41b1-bcba-d6cd65a08b5d · outbound

This paper cites https://polyhaven.com.

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception https://polyhaven.com

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:53:07.254902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:53:06.698629Z digest=sha256:4891c2643d3bc332677ae0b291cfdc5054b7d4dbea2332227ee97bd7473109ee

Observation 4acd3479-055d-43d4-85b6-d4afd9523ad3 · outbound

This paper cites Neural feels with neural fields: Visuo-tactile perception for in-hand manipulation.

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception Neural feels with neural fields: Visuo-tactile perception for in-hand manipulation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:53:06.726354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:53:06.726354Z digest=sha256:25de658b28f0571cf1b180dc7fcd99f76fe2ed67244e29e6b93804a10682cccb

Observation 8d10f87b-96ce-4ea5-93a8-3babc0353160 · outbound

This paper cites In the upper row, we have an AllegroHand equipped with whole-hand tactile sensors, while in the lower row, we have a Trx-hand with sensors.

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception In the upper row, we have an AllegroHand equipped with whole-hand tactile sensors, while in the lower row, we have a Trx-hand with sensors

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:53:07.177228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:53:06.754284Z digest=sha256:eb5e2f161c54a68cedc9a0939b13754c1e07d960eea0cbcf748d3eaaa115bc60

Observation 69100ff1-03ec-41e5-8129-205e4215c038 · outbound

This paper cites UniDexGrasp++: Improving Dexterous Grasping Policy Learning via Geometry-aware Curriculum and Iterative Generalist-Specialist Learning.

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception UniDexGrasp++: Improving Dexterous Grasping Policy Learning via Geometry-aware Curriculum and Iterative Generalist-Specialist Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:53:06.740556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:53:06.740556Z digest=sha256:3f90a2af963c15a66012ae7177a8afb6634c491898d3222cebf58a02eb09cbfb

Observation ca5ae062-b5ab-4263-b851-9f58a562b093 · outbound

This paper cites an unresolved cited work.

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:53:07.192113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:53:06.745215Z digest=sha256:04d25a085c39f5de63a836dee1f55b0ee8c4fed28fe4886dba303430c9102897

Observation 928556ce-6dc1-44c0-b2d1-9c3b5e5fc05d · outbound

This paper cites PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes.

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:53:06.749871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:53:06.749871Z digest=sha256:7968d9c3d44a0a9468a799cd47a4d6fda519f70f8849a3a7f9e46007caeef5e3

Observation 91e95be2-3494-44dc-a554-9d2d4b9ceb71 · outbound

This paper cites (2) (Dikhale et al., 2022; Li et al., 2023; Rezazadeh et al.,.

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception (2) (Dikhale et al., 2022; Li et al., 2023; Rezazadeh et al.,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:53:07.162629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:53:06.759114Z digest=sha256:3b21cdf7a14c169071433a5c5454c0f2ab47c129b55b301327a479578f780783

Observation adac323f-a0ae-4eb5-aafc-8ab872d35186 · outbound

This paper cites an unresolved cited work.

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:53:07.148525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:53:06.763702Z digest=sha256:b403bb6eaed72f54ae98795de512e6a47725a1c68062ce45d22739f6a007c512

Observation be2256af-2e51-4c3e-8400-967e3cd1160f · outbound

This paper cites However, these approaches often do not address the critical aspect of stable holding in hand and readiness for manipulative actions, which are key focuses of our research.

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception However, these approaches often do not address the critical aspect of stable holding in hand and readiness for manipulative actions, which are key focuses of our research

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:53:07.133016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:53:06.768472Z digest=sha256:da9f62f671c7bbcc3daa4c0fb632389287f911d67709f317ccd2b56a23c8d51e

Observation d3c566f5-8399-428a-8d18-231e994527b5 · outbound

This paper cites We present a set of photo-realistic rendered images that depict objects held in hand.

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception We present a set of photo-realistic rendered images that depict objects held in hand

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:53:07.118350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:53:06.772581Z digest=sha256:388945b9746f8b9ffd1f2c0ef2b6ee06bb986fab5e280c067d03e47fc2493f64

Observation 31be5084-c1a9-4bdb-bc07-3f301224d834 · outbound

This paper cites The motion capture system accurately captures the poses of the object, hand, and robot base.

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception The motion capture system accurately captures the poses of the object, hand, and robot base

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:53:07.103785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:53:06.777561Z digest=sha256:d489b98a5a772c3cd2b09be77f794d8644748975057d1768134ae0d9d0550c24

Observation d21fa7be-229a-4e38-ba5a-e9b952afb8c6 · outbound

This paper cites an unresolved cited work.

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:53:07.059212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:53:06.792201Z digest=sha256:d6ef2b0835c457feeae5087e3ad503e4927685592dd18b6d214981a677339392

Observation 5ec26a06-1034-4ae4-b4e8-aceab0e66bff · outbound

This paper cites After distortion correction, we reproject the depth image into the right camera’s frame, resulting in two well-aligned vision images with resolutions of 640 x 400,.

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception After distortion correction, we reproject the depth image into the right camera’s frame, resulting in two well-aligned vision images with resolutions of 640 x 400,

Reference 565

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:53:07.089123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:53:06.782168Z digest=sha256:18ffde5d4d7963601abc9720ad2999c166fbab7c19c3c5fa66e34ac666989a97

Observation 810b5c62-ac58-4dda-98d5-58095b25e236 · outbound

This paper cites Tactile Pose Estimation and Policy Learning for Unknown Object Manipulation.

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception Tactile Pose Estimation and Policy Learning for Unknown Object Manipulation

Reference 2011

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:53:07.045312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:53:06.703397Z digest=sha256:39aacf5720d3de4b4ff4ef961160c5f59cee4fb0941bd6af7e6134b789b6adcc

Observation d576de0d-1f54-4f60-856f-5ab6706f43e6 · outbound

This paper cites Learning a state estimator for tactile in-hand manipulation.

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception Learning a state estimator for tactile in-hand manipulation

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:53:07.225242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:53:06.717746Z digest=sha256:49aa1e98c388f48dfdd420e9e087cdcf5b2c435bd48c74ad32780e5ef6394035

Observation 63c9aa5e-7acd-4e44-a150-d0f10b2ef7ac · outbound

This paper cites Smith, E., Calandra, R., Romero, A., Gkioxari, G., Meger, D., Malik, J., and Drozdzal, M.

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception Smith, E., Calandra, R., Romero, A., Gkioxari, G., Meger, D., Malik, J., and Drozdzal, M

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:53:07.208955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:53:06.722087Z digest=sha256:9a6f97f74b9dd9494263ffaa5dafa2c03b6452f6b03f6b41a4a53d7f2a58ad24

Observation 7e6c5da5-4bb0-47f4-a421-f26a0cbea85c · outbound

This paper cites However, we believe that simply sticking markers on objects may not yield high-quality pose data due to the uncertainty in marker positions relative to the object frame.

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception However, we believe that simply sticking markers on objects may not yield high-quality pose data due to the uncertainty in marker positions relative to the object frame

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:53:07.073368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:53:06.787336Z digest=sha256:8d23ed7b6da19ef412ad059caaf0204be6aceca2ce342fd71c8b83dcc67d614e

Observation e2afa834-ffef-4175-9a8e-30ee7089598c · outbound

This paper cites Tracking object’s pose via dynamic tactile interaction.

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception Tracking object’s pose via dynamic tactile interaction

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:53:07.240535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:53:06.713303Z digest=sha256:de71b31ca19d4f2f86b56423f8aeb42c425396c8c726d67dd3712ad2d34ccd44

Observation cf1e99c9-ea1f-4a0e-a11e-7f4f460dad33 · outbound

This paper cites PoseFusion: Robust Object-in-Hand Pose Estimation with SelectLSTM.

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception PoseFusion: Robust Object-in-Hand Pose Estimation with SelectLSTM

Reference 2021

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:53:06.995626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:53:06.730772Z digest=sha256:0c00c312dff607b5263aaf9ee09679271a27a0fdb032cdcdced029e5138b9351

Observation a29da9b6-3d7b-4ecf-b7f1-33f0ffd39a62 · outbound

This paper cites Segment Anything.

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception Segment Anything

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T22:53:06.708310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:53:06.708310Z digest=sha256:0f89a7b6b1c9058db38f445d2ff1f15e9aabbfd1ecf7cb171af9b489f48cacf7

Observation 80f6faa6-7224-4f0f-a8cf-4349805614c6 · outbound

This paper cites Fast-Grasp'D: Dexterous Multi-finger Grasp Generation Through Differentiable Simulation.

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception Fast-Grasp'D: Dexterous Multi-finger Grasp Generation Through Differentiable Simulation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T22:53:06.735353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:53:06.735353Z digest=sha256:fc915bf4ff60fbf3bbef30c8d7b85b87ec289f21d1694c4ec7c9bdb5b5e572aa

Pith citing papers

Observation afef0867-5d2e-4ce2-a4c5-4c8db5e74efc · inbound

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms cites this paper.

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.383723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T12:52:15.138790Z digest=sha256:4518dbdba229149aef61099458ed485b668f19c012416efb0793f08b85e233b1