{"id":"daf19c16-5ea4-4dde-9aed-5521d3f0a9d4","arxiv_id":"2501.00510","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A new large-scale vision, touch, and proprioception dataset for estimating the 6D pose of objects held in multi-fingered robotic hands, with a baseline network showing improved accuracy when touch is added.","lead":"This paper introduces VinT-6D, a dataset of two million simulated and one hundred thousand real measurements of objects held in robotic hands, combining camera images, whole-hand touch sensors, and joint positions. It provides a benchmark for estimating an object's 6D pose while it is gripped, a capability needed for robots to manipulate objects they cannot fully see.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's core real-split value rests on an unmeasured 'sub-millimeter' Vicon object-pose claim (Sec 3.2.1/App A.2.4); without fixture calibration and tracking-loss statistics, benchmark tables cannot be trusted.","rationale":"The reader identifies the same load-bearing premise: VinT-Real's benchmark value depends on unverified sub-millimeter object-pose ground truth. I agree. The paper claims sub-millimeter accuracy (Sec 3.2.1, App A.2.4) but provides no measurement of fixture rigidity, calibration residual, tracking continuity, or independent reprojection. Because all real-split experiments (Tables 3-6) evaluate against this label source, any undisclosed error propagates into every quantitative claim. This makes the paper conditionally acceptable at best: the dataset concept and pipeline are plausible, the simulated split's visual/tactile synthesis is described in detail, and the availability of fixture models is promised, but the decisive validation data are absent. The sample-count arithmetic (2,500 x 25 x 4 x 3 = 750k, not 2M; 125k x 4 x 3 = 1.5M, not 2M) is a separate internal inconsistency that also needs correction, but it is less fundamental than GT accuracy: a corrected count of 1.5M would still be a large dataset, whereas uncorrected GT error would invalidate the empirical conclusions. Therefore the reader's CONDITIONAL verdict should stand, with the conditions being: release the data, fix the sample count, and publish a GT-validation protocol with measured error distributions. No change is needed to the verdict category.","tokens_in":18106,"tokens_out":5397,"duration_ms":58863,"concrete_test":"Run a validation pass on VinT-Real: for a fixed Trx-hand grasp of one object, record 5k stationary and 5k moving Vicon frames; compute object-pose standard deviation and 95th-percentile translation error relative to the hand flange, the fraction of frames with fewer than 3 visible markers or invalid rigid-body fits, and the mean reprojection error of the mocap-derived CAD model onto the aligned RGB-D point cloud. If median translation error exceeds 1.0 mm or more than 1% of frames lose tracking, the sub-millimeter claim and Tables 3-6 are unsupported. If errors are within 1.0 mm with no significant tracking loss, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"VinT-Real's claim to be the largest high-quality real dataset (Sec 3.2, App A.2.4) depends on object-pose ground truth being sub-millimeter while held and manipulated. The paper states that custom marker fixtures and Vicon 'ensure' this accuracy, but reports no data supporting it: no fixture-to-object calibration residual, no static/dynamic repeatability, no marker-visibility/tracking-loss statistics, and no independent reprojection of the mocap-derived CAD model onto the RGB-D point cloud. The hand occludes markers, fixtures can flex under grasp forces, and 'toddler-like' exploratory motions (Sec 3.3.7) can cause intermittent tracking loss; any of these would inject millimeter-to-centimeter label noise. Since every ADD(S) AUC number in Tables 3-6 is computed against this GT, and the touch-point alignment in Fig. 4 is validated only visually, the real split's central quality claim is currently an assertion, not a measured result. This is distinct from the sample-count inconsistency in Sec 3.3.7 (2,500 x 25 x 4 x 3 = 750k, not 2M), which mainly affects the scale claim; GT accuracy affects every quantitative conclusion drawn from VinT-Real.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces VinT-6D, a multi-modal dataset for 6D object-in-hand pose estimation with two splits: VinT-Sim (synthesized) and VinT-Real (collected on a custom robotic platform). The dataset combines RGB-D vision, whole-hand tactile sensing, and proprioception, and is built for three- and four-fingered robotic hands. The authors describe a MuJoCo/Blender simulation pipeline, a real platform with Vicon motion capture, and a baseline network VinT-Net that fuses vision, touch, and proprioception. Experiments on the real split report that adding touch and proprioception improves pose accuracy, especially under hand-induced occlusion. The appendices describe the hardware setup, tactile sensor modeling, camera calibration, and segmentation procedure in detail.","tokens_in":18304,"tokens_out":6739,"duration_ms":61345,"significance":"If the dataset is released as described, it is a potentially valuable community resource: it is larger than existing real object-in-hand datasets, covers whole-hand tactile sensing rather than only fingertips, and provides an independent motion-capture ground truth that is not derived from the proposed baseline. The simulation pipeline, particularly the 3D scanning of taxel distributions, is concrete and useful for sim-to-real transfer. The VinT-Net baseline is simple and gives initial evidence that touch helps under occlusion. The main open threats are the unquantified sub-millimeter claim for the real-split ground truth and an inconsistent sample-count arithmetic; both are fixable in revision but are load-bearing for the paper's central claims.","major_comments":[{"comment":"The dataset-scale arithmetic is internally inconsistent. Section 3.3.7 states \"2,500 interactions for each object\" with \"four different environments, with three camera positions for each,\" which yields 2,500 × 25 × 4 × 3 = 750,000 VinT-Sim samples, not the claimed \"2 million\". If the object count is 26 as suggested by Section 3.3.1 (21 YCB objects plus 5 added objects), the total is 780,000; if one instead uses the 125,000 grasps stated in Section 3.1.2, 125,000 × 4 × 3 = 1,500,000. None of these reproduces the abstract's 2M figure. The same inconsistency propagates to Table 1 and Section 6. Please reconcile the counting definitions (per-object, per-hand, per-view) and correct the stated scale.","section":"§3.3.7, §3.1.2, Abstract, Table 1"},{"comment":"The real-split ground-truth pose accuracy is asserted but not measured. The text claims that the Vicon system plus custom fixtures \"ensure sub-millimeter accuracy\" (Appendix A.2.4), yet no calibration residual for the fixture-to-object transform, no static or dynamic repeatability test, no fixture-deflection test under grasp forces, and no marker-occlusion or tracking-loss statistics are reported; Figures 4 and 15 provide only visual checks. Because every ADD(S) AUC number in Tables 3–6 is evaluated against this ground truth, any unmodeled calibration error would propagate directly into all quantitative conclusions drawn from VinT-Real. Please add a quantitative accuracy characterization, such as reprojection error of the mocap-derived CAD model onto the RGB-D point cloud, worst-case error across the ten objects, and the tracking-loss fraction during the toddler-like motions of Section 3.3.7.","section":"§3.2.1 and Appendix A.2.4"},{"comment":"No error bars or repeated-run statistics are reported for any benchmark result. All tables give single point estimates, so the reported improvements (for example, the large shaker row in Table 4, 87.04 vs. 94.66, or the occlusion-robustness gap in Table 6) are not shown to be significant relative to training noise or random split variability. For a paper whose purpose is to establish a benchmark, please provide mean and standard deviation over at least three seeds or train/test splits, and specify the exact splits used.","section":"Tables 3–6"},{"comment":"The comparison against prior methods is not sufficiently specified. The paper states that the Object-Hand-Pose method (Wen et al., 2020) was reproduced \"by replacing its two-finger gripper with our three-finger hand, which has reduced degrees of freedom. This and other adjustments were made...\" but it does not enumerate the other adjustments, the optimization settings, the training data used for the baseline, or the compute budget. Without these details, or released reproduction code and configurations, the claim that VinT-Net \"significantly outperforms\" prior work in Table 5 is not independently verifiable.","section":"§5.3, Table 5"}],"minor_comments":[{"comment":"The object count is inconsistent: Section 3.3.1 says 21 YCB objects plus 5 added objects (26 in total), while Section 3 and Figure 18 say 25 objects. Please reconcile this number.","section":"§3.3.1, Section 3, Figure 18"},{"comment":"There are several typos and terminology inconsistencies: \"Tencnet\" in the affiliations, \"Internationl Journal of Robotics Research\" in the Calli et al. reference, \"motion caption system\" in the Figure 4 caption, \"Benifit\" in Section 5.3, the duplicated phrase \"robust object segmentation reliable object segmentation\" in Section 3.3.6, and inconsistent capitalization between \"VinT-Net\" and \"Vint-Net\" in Section 3.3.4.","section":"Throughout"},{"comment":"The text refers to \"SAM model\" and \"SAM modal\" in places where the intended meaning is the Segment Anything Model; please use consistent terminology such as \"SAM\" or \"the Segment Anything model.\"","section":"§3.3.6 and §4.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is earnest and technically detailed, and the dataset, if released, will likely be widely used. In my assessment there is no circularity: the ground truth comes from MuJoCo and Vicon, independent of VinT-Net. The two load-bearing issues are the unsupported sub-millimeter ground-truth claim for VinT-Real and the inconsistent sample count in the abstract and Section 3. Both are fixable in revision but require the authors to add measurement evidence and correct the arithmetic. I would also encourage the authors to provide error bars and a fuller description of the reproduced baselines, since the benchmark's credibility depends on them."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The dataset is genuinely new: real three- and four-finger hands with whole-hand tactile arrays (620/679 taxels), 100k real samples, and mocap object poses. No prior dataset combines these. If the data actually ships, it could become a standard benchmark for in-hand pose estimation and sim-to-real transfer. The touch simulation strategy—simulating contact positions rather than raw taxel readings—is sensible and addresses a real gap.\n\nThat said, there are two soft spots that matter. First, the simulated sample count doesn't multiply. Section 3.3.7 says 2,500 interactions per object × 25 objects × 4 environments × 3 cameras = 2M, but that's 750k. Earlier in 3.1.2 they say 125,000 successful grasps, which ×4×3 = 1.5M. Neither gives 2M. This is a claim about the headline number, so it should be fixed or explained. Second, and more important, the paper asserts 'sub-millimeter' object pose accuracy from Vicon markers on custom fixtures, but reports no calibration residual, no tracking-loss statistics, no independent reprojection check. Given that the hand occludes markers and the fixtures can flex under grasp forces, this is a load-bearing assertion. Every result in Tables 3–6 inherits that label quality. The paper needs to measure and report that error.\n\nMinor points: no error bars in any table, and main evaluations cover 7 of the 10 real objects. The baseline VinT-Net is a standard keypoint-voting architecture, which is fine for a dataset benchmark. The limitations section is honest about object/scene diversity.\n\nIf the GT accuracy is confirmed and the data is released, this is a solid contribution worth serious refereeing. As it stands, the paper is conditionally acceptable: the dataset concept is strong, but the two quantitative claims need support.","headline":"Strong, potentially benchmark-setting dataset paper held back by unmeasured ground-truth accuracy and an arithmetic inconsistency in the headline sample count.","tokens_in":18962,"tokens_out":1962,"would_cite":false,"duration_ms":18394,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"VinT-6D is a 2.1-million-sample dataset pairing vision, touch, and proprioception for 6D object-in-hand pose estimation, with a baseline showing touch improves accuracy under occlusion.","keywords":["object-in-hand pose estimation","multi-modal robot perception","tactile sensing","proprioception","sim-to-real transfer","6D pose estimation","robotic manipulation","dataset benchmark"],"falsifier":"Mount a high-resolution external camera or a second motion-capture setup and independently measure the same object pose during grasps; if the difference between the fixture-based pose and the independent measurement exceeds a millimeter at typical grasp forces, or if marker dropout occurs during recorded segments, the sub-millimeter ground-truth claim fails.","tokens_in":17844,"feed_emoji":"🖐️","tokens_out":5436,"duration_ms":50389,"temperature":0.7,"pith_summary":"This paper presents VinT-6D, a dataset built to support 6D object-in-hand pose estimation for multi-fingered robotic hands by combining three modalities that are usually collected separately: vision (RGB-D), whole-hand touch, and proprioception. The authors claim it is the first extensive dataset of its kind, with two million simulated samples generated in MuJoCo and rendered in Blender, plus one hundred thousand real samples collected on a custom robot platform. Real objects and hands carry custom marker fixtures tracked by a motion capture system, which the paper asserts provides sub-millimeter object-pose ground truth while the object is held. On top of the dataset, the paper introduces VinT-Net, a baseline that fuses color, depth, and touch point clouds; its experiments show that adding touch to vision raises pose accuracy and makes it degrade more slowly as hand occlusion increases. A sympathetic reader would care because accurate in-hand pose labels at this scale are the missing ingredient for training and benchmarking perception models for dexterous manipulation.","feed_headline":"2.1M multimodal samples back object-in-hand pose estimation","feed_subtitle":"Simulated and real splits cover whole-hand touch; adding touch keeps pose accuracy high as hand occlusion reaches 50%.","key_machinery":"The load-bearing mechanism is the aligned multi-modal ground-truth pipeline. On the real side, custom-printed fixtures with motion-capture markers are attached to each object, and a flange plate with markers on the hand lets forward kinematics convert activated piezoresistive taxels into a local touch point cloud in the palm frame; the same transformed poses align with RGB-D from a Kinect Azure and stereo cameras. On the simulation side, each taxel is modeled as a force-sensing cuboid on the finger surface, so contact positions rather than raw readings are the shared representation for pressure and vision-based tactile sensors. VinT-Net then treats touch as a complementary point cloud fused with depth features before a keypoint-voting head predicts the object's center and 3D keypoints.","core_discovery":"The paper's central claim is that VinT-6D supplies the scale, modality coverage, and label quality needed to train object-in-hand pose estimators that work in real multi-finger grasps. VinT-Sim contributes two million samples by simulating whole-hand taxel distributions on a three-fingered and a four-fingered hand, generating stable grasps in MuJoCo, and rendering photo-realistic RGB-D in Blender. VinT-Real contributes one hundred thousand samples from a platform that uses a motion capture system with custom object fixtures to record object and hand poses, forward kinematics to build touch point clouds, and SAM guided by touch prompts to segment the object from the hand. The dataset covers 25 simulated and 10 real objects across household materials, with occlusion rates from 10% to 90% in simulation. The accompanying VinT-Net uses a U-Net for color features, PointNet++ for depth and touch, and a keypoint-voting pose head; in their experiments the vision-plus-touch model outperforms vision-only on all seven reported real objects and loses less accuracy as occlusion rises from 20% to 50%.","pith_inferences":["If the ground-truth accuracy claim holds, the same collection platform could be extended to category-level or zero-shot object-in-hand pose estimation, which the paper itself names as open future work; the fixture-based labels would then serve as a supervisory signal independent of object-specific meshes.","The dataset's choice to store tactile data as contact positions rather than raw taxel readings, if adopted more widely, could let models trained on one tactile sensor type transfer to another with less recalibration; the paper argues for this representation but does not demonstrate cross-sensor transfer.","A natural testable extension is to push occlusion beyond 50% to find where touch alone, without vision, becomes the better estimator; the paper's occlusion experiments stop before that crossover.","The transparent and reflective objects added to the simulated split could be used to test whether touch and motion-capture labels enable pose estimation where RGB-D alone fails, a regime the paper does not evaluate."],"forward_implications":["Training on the combined simulation-plus-real split gives higher ADD(S) AUC on the reported objects than training on either split alone, supporting the paper's claim that the pipeline narrows the sim-to-real gap.","Adding touch and proprioception to vision raises ADD(S) AUC on every real object in the ablation, with the largest gains on objects that are hard for vision, such as the large shaker and tomato soup can.","Under increasing hand occlusion, vision-plus-touch accuracy falls from 94.76% to 88.81% at 50% occlusion, while vision-only falls to 80.43%, so the tactile modality is what keeps the estimate usable in heavy occlusion.","The dataset's three-finger and four-finger splits with whole-hand taxel distributions give a benchmark that two-finger gripper datasets cannot provide, for methods that need multi-finger contact reasoning."],"supporting_citations":[{"why":"The prior visuotactile in-hand pose dataset and method that VinT-6D extends and compares against in scale and real-world coverage.","marker":"(Dikhale et al., 2022)"},{"why":"Fast-Grasp'D supplies the multi-finger grasp generation baseline that VinT-Sim's selected stable grasping is contrasted with.","marker":"(Turpin et al., 2023)"},{"why":"The YCB object set provides most of the simulated household objects and the prior ArUco-based pose ground truth approach the paper replaces.","marker":"(Calli et al., 2017)"},{"why":"SAM is the segmentation model prompted by touch and proprioception to produce object-in-hand segmentation labels in VinT-Real.","marker":"(Kirillov et al., 2023)"},{"why":"PVN3D supplies the 3D keypoint voting architecture and ADD/ADD-S evaluation conventions that VinT-Net is built on and compared with.","marker":"(He et al., 2020)"},{"why":"MuJoCo is the physics engine used to simulate stable object-grasp interactions and taxel contacts in VinT-Sim.","marker":"(Todorov et al., 2012)"},{"why":"The two-finger Hand-Object dataset and pose method are the main prior object-in-hand benchmark that VinT-Real is compared against.","marker":"(Wen et al., 2020)"},{"why":"The characterization of Kinect Azure depth noise guides the simulated depth post-processing used to narrow the sim-to-real visual gap.","marker":"(Tolgyessy et al., 2021)"}],"fun_headline_variants":["VinT-6D: 2.1M multimodal samples for in-hand pose estimation","Touch boosts object-in-hand pose accuracy under heavy occlusion","Multimodal robot hand dataset: vision, touch, and proprioception","2.1M samples fuse touch and vision for in-hand pose estimation","VinT-6D: Largest vision-touch dataset for in-hand manipulation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that VinT-Real's object poses are sub-millimeter accurate rests on custom marker fixtures and motion capture staying rigid and fully tracked while the hand holds and moves each object; the paper reports no calibration error, fixture-flex test, reprojection check, or tracking-loss statistics.","fun_headline_variants_meta":{"raw":{"variants":["VinT-6D: 2.1M multimodal samples for in-hand pose estimation","Touch boosts object-in-hand pose accuracy under heavy occlusion","Multimodal robot hand dataset: vision, touch, and proprioception","2.1M samples fuse touch and vision for in-hand pose estimation","VinT-6D: Largest vision-touch dataset for in-hand manipulation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000469,"raw_usage":{"total_tokens":2349,"prompt_tokens":971,"completion_tokens":1378,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":587,"completion_tokens_details":{"reasoning_tokens":1281}},"tokens_in":587,"tokens_out":1378,"duration_ms":12230,"temperature":1.0,"reasoning_tokens":1281,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:49:24.205175+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Mount a high-resolution external camera or a second motion-capture setup and independently measure the same object pose during grasps; if the difference between the fixture-based pose and the independent measurement exceeds a millimeter at typical grasp forces, or if marker dropout occurs during recorded segments, the sub-millimeter ground-truth claim fails.","supporting_citations":[{"cited_title":"However, we believe that simply sticking markers on objects may not yield high-quality pose data due to the uncertainty in marker positions relative to the object frame","cited_arxiv_id":null,"evidence_quote":"The prior visuotactile in-hand pose dataset and method that VinT-6D extends and compares against in scale and real-world coverage."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The two-finger Hand-Object dataset and pose method are the main prior object-in-hand benchmark that VinT-Real is compared against."}],"review_version":1}