{"id":"4c52d160-61b6-41d1-8a9f-c4ee0b7c8160","arxiv_id":"2504.15329","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new interactive tool for manually annotating 6D object poses by aligning 3D models onto 2D images, evaluated with a user study on Linemod and HANDAL.","lead":"Vision6D is an open-source desktop tool that lets users align 3D object models onto 2D images to annotate 6D positions and orientations. A user study with 11 annotators on public datasets reports average angular errors around 5 to 6 degrees and annotation times near 100 seconds per object.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Best-of-three selection and top-90% trimming inflate reported annotation accuracy; with no statistical tests, the 'statistically comparable' claim is not yet supported.","rationale":"The user-study evidence is the only support for the central claim of statistically comparable annotations, so anything that changes the reported error magnitudes is load-bearing. The paper's own protocol introduces two selection mechanisms that each lower the reported error: choosing the best of three repeated annotations by ground-truth ADD, and reporting only the top 90% of results. A concrete symptom is that the intra-personal repeatability mean (7.11°/10.08°) is larger than the inter-personal best-of-three mean (4.77°/5.88°): with trial-to-trial variation of about 7-10°, selecting the closest-to-ground-truth of three trials can easily produce a mean error several degrees below the error of any single annotation. The \"statistically comparable\" wording in H1 is also unsupported by any significance test or equivalence bound; reporting means and standard deviations after error-based trimming is not a statistical comparison. The reader's intrinsics concern is legitimate for the tool's general use, but in the benchmark evaluation the datasets ship calibration matrices and the tool consumes K directly, so K error would only matter if the study used the wrong matrices, and no evidence suggests that. The selection problem is internal to the evaluation and therefore more directly undermines the headline numbers. I do not recommend REJECT: the tool is open source, the interface design is plausible, and a re-analysis of already-collected trial logs could settle the issue. The appropriate verdict remains CONDITIONAL, with the condition being a bias-free re-analysis. This is why I set verdict_should_be to UNCHANGED (the reader's CONDITIONAL is correct) and mark agreement as partial: the reader mentioned selection in the rationale but made intrinsics the headline weakest assumption.","tokens_in":13976,"tokens_out":5366,"duration_ms":51128,"concrete_test":"Re-analyze the raw per-trial annotation logs and recompute the inter-personal metrics without best-of-three selection and without top-90% trimming: use the first trial for every participant-sample pair and keep all 10 samples per dataset. Compare the resulting mean angular and ADD errors against the reported 4.77°/15.27 mm and 5.88°/30.53 mm. If the first-trial mean angular error exceeds 8° or mean ADD exceeds 40 mm on either dataset, or if the median error rises by more than 2°, the headline accuracy is an artifact of post-hoc selection and the \"statistically comparable\" claim needs to be withdrawn or substantially qualified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing problem is the evaluation protocol, not the intrinsics assumption. Section IV.B.4 states that \"we present the top 90% results to eliminate outliers,\" and Section V.A.1 reports inter-personal errors \"calculated using each participant's best pose annotation from the three repeated trials, defined as the one with the lowest ADD score.\" Both choices select on the error being measured. Taking the minimum of three trials and discarding the worst 10% of trials guarantees a lower mean error than a single first attempt, and because the selection uses ground-truth ADD, it cannot be replicated by a user who does not already know the pose. The reported 4.77°/5.88° angular errors are therefore a best-case, post-hoc bound, not expected annotation accuracy. This is visible in the paper's own intra-personal numbers: the mean within-user repeat variability is 7.11° (Linemod) and 10.08° (HANDAL), larger than the best-of-three inter-personal means, so trial-to-trial variation is large enough to move the headline metric. In addition, H1 claims annotations are \"statistically comparable\" to ground truth, but no statistical test, confidence interval, or equivalence bound is reported; the claim rests entirely on descriptive statistics after selection. The intrinsics concern is real but secondary: Linemod and HANDAL distribute known calibration matrices, and the tool's stated use case is annotation \"using only the camera intrinsic matrix,\" so in the reported evaluation K is presumably fixed and correct. The selection bias, by contrast, directly determines the magnitude of the headline accuracy numbers.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"Vision6D is an interactive 3D-to-2D visualization and annotation tool for 6D object pose estimation. The paper describes the tool's interface and mathematical formalism (projection through a pinhole camera model using a known intrinsic matrix K), then presents a user study with 11 participants who annotated 10 samples from the Linemod and HANDAL datasets three times each. The study measures inter-personal and intra-personal annotation errors against ground-truth poses using angular distance, Euclidean distance, and ADD, plus annotation time and NASA-TLX/SUS questionnaires. The central claim is that Vision6D enables users to produce 6D pose annotations 'statistically comparable' to ground truth, with reported best-of-three mean angular errors of 4.77° (Linemod) and 5.88° (HANDAL).","tokens_in":14238,"tokens_out":6528,"duration_ms":55360,"significance":"If the accuracy and efficiency results are confirmed after correcting the evaluation protocol, Vision6D would be a valuable open-source contribution to the 6D pose annotation community. The paper uses external ground truth from standard datasets, applies commonly accepted metrics, and includes usability instruments (NASA-TLX, SUS). The reproducible open-source release is a clear strength. However, the current evaluation contains selection bias and lacks inferential statistics, so the headline accuracy numbers and the 'statistically comparable' claim (H1) are not yet established. With appropriate revisions, the tool's contribution could be solid.","major_comments":[{"comment":"The evaluation protocol selects the best of each participant's three trials by lowest ADD and additionally reports only the 'top 90%' of results to eliminate outliers. Both procedures condition on the error being measured: the best-of-three selection cannot be replicated by an end user who does not know the ground truth, and the top-90% trimming removes the worst trials without a principled criterion. Consequently, the reported mean angular errors (4.77° for Linemod, 5.88° for HANDAL) are best-case post-hoc bounds, not expected annotation accuracy. The paper's own intra-personal data support this concern: the mean intra-personal angular distance is 7.11° (Linemod) and 10.08° (HANDAL), larger than the corresponding inter-personal best-of-three means, showing that trial-to-trial variability is large enough to change the headline metric. Please report results for all trials, for the first trial only, and the number of trials excluded by the trimming rule.","section":"IV.B.4, V.A.1"},{"comment":"Hypothesis H1 claims that Vision6D annotations are 'statistically comparable' to ground truth, but the results section provides only descriptive statistics and no statistical test, confidence interval, or equivalence bound. After the selection procedures described above, the descriptive statistics cannot support an inferential claim. To substantiate H1, the authors should pre-specify an equivalence margin for the angular and ADD metrics and apply a standard equivalence test (e.g., two one-sided tests) using the full per-trial data, reporting confidence intervals and effect sizes.","section":"IV.A (H1), V.A"},{"comment":"The projection pipeline depends entirely on the camera intrinsic matrix K being known and correct. The paper acknowledges that the tool uses 'only the camera intrinsic matrix,' but it performs no sensitivity analysis for errors in K, and it does not discuss how estimated intrinsics (e.g., from a calibration algorithm or a different camera) would affect annotation accuracy. This is a key limitation for real-world use, since users of the tool may have imperfect intrinsics. Please add a sensitivity experiment that perturbs the components of K (fx, fy, cx, cy) over a plausible range and reports the resulting increase in angular and ADD errors, or explicitly state the accuracy requirement on K.","section":"III.A, Eq. (3), IV.B.4"},{"comment":"The study does not address object symmetries in the evaluation. The limitation section correctly notes that symmetric and textureless objects create pose ambiguities, but the ADD metric used in Equations (4)-(5) is not symmetry-aware. If any of the selected Linemod or HANDAL objects have rotational symmetries, a visually correct annotation could be scored as a large ADD error, potentially biasing the reported accuracy. The authors should either exclude symmetric objects from the accuracy analysis, use a symmetry-aware metric (e.g., ADD-S), or justify that the selected objects do not have symmetries.","section":"VI, V.A.1"}],"minor_comments":[{"comment":"The phrase 'we present the top 90% results to eliminate outliers' is ambiguous; specify the unit of trimming (participants, trials, or samples) and the rationale for the 90% threshold.","section":"IV.B.4"},{"comment":"The table rows are labeled L0/L1... and H0/H1..., but the text refers to participants as L0 to L4 and H0 to H5; the connection to samples S0-S9 is unclear. The table should include the overall mean and standard deviation reported in the text.","section":"Table I"},{"comment":"The example uses the Linemod-Occluded dataset [21], while the user study uses Linemod [10]; clarify the relationship between the datasets to avoid confusion.","section":"III.B"},{"comment":"Typographical and notation issues: 'LineMod' and 'Linemod' are used inconsistently; the acronym 'ADD' is written as 'Add' in the captions of Figures 4 and 5; and Equation (3) uses 'Puvw' without explicitly defining the subscript convention. Please ensure consistent terminology.","section":"Throughout"},{"comment":"Figures 4, 5, and 7 are referenced but do not appear in the manuscript text; in the final version, ensure that all figures are embedded with clear axis labels, legends, and statistical annotations.","section":"Figures 4, 5, 7"},{"comment":"The conclusion states that Vision6D 'has supported several deep-learning-based studies' [24]-[28]; please briefly describe the role of the tool in these studies to support this claim.","section":"VII (Conclusion)"},{"comment":"The limitation section discusses symmetric objects, but the user study section does not report which specific objects were selected from Linemod and HANDAL, making it impossible to assess the impact of symmetries and object difficulty on the reported errors. Please include the object list and, if relevant, object-level error analysis.","section":"IV.B (Stimuli)"},{"comment":"The study does not include a baseline annotation method (e.g., manual numerical pose entry or another interactive tool). While not required to validate the tool's internal accuracy, a comparison would strengthen the claim that Vision6D is more efficient and intuitive than alternatives.","section":"IV.B (Study Design)"}],"recommendation":"major_revision","confidential_remarks":"The paper describes a useful tool, but the evaluation protocol is the main obstacle: best-of-three selection and top-90% trimming inflate the reported accuracy, and the 'statistically comparable' claim needs proper inferential testing. The lack of a baseline comparison also weakens the novelty statement, though it is not fatal. The small sample (5-6 participants per dataset) further limits generalizability. I recommend major revision with a focus on transparent reporting of all trials and appropriate statistical analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: Vision6D is a genuine, open-source annotation tool with a plausible interface, but the paper overstates what the user study proves. The headline accuracy numbers come from best-of-three selection and post-hoc trimming, so the 'statistically comparable' claim isn't supported.\n\nWhat's genuinely useful: the tool itself. It does interactive 3D-to-2D alignment with immediate projection feedback, which is handy for generating first-frame annotations when you only have intrinsics and a mesh. The UI is thoughtfully described, the math is standard pinhole, and the software is open source. The user study is a real attempt: 11 users, two benchmark datasets, standard metrics, plus NASA-TLX and SUS. That's more evaluation than most tools papers bother with.\n\nThe soft spots are in the evaluation protocol. Section IV.B.4 says they report the top 90% of results, and Section V.A.1 says inter-personal errors use each participant's best of three trials, defined as the one with the lowest ADD score. Both choices select on the error being measured. The intra-personal repeat variability (7.11° Linemod, 10.08° HANDAL) is larger than the best-of-three inter-personal means (4.77° and 5.88°), which shows trial-to-trial noise is big enough to move the headline number. The authors disclose these choices, which is good, but the reported accuracy is a best-case bound, not expected annotation performance. And H1's 'statistically comparable' has no statistical test or equivalence bound behind it.\n\nThe 'first tool' novelty claim is also overreach; the related work doesn't establish that the interaction pattern is new. The intrinsics assumption is real but secondary here, since Linemod and HANDAL ship known calibration matrices.\n\nBottom line: this is a useful contribution to the annotation-tool niche, and it deserves a serious referee, but the evaluation section needs to be reframed. The tool is fine; the claims need to match the evidence.","headline":"Vision6D is a real, open-source annotation tool with a plausible interface, but the evaluation's best-of-three and top-90% selections inflate the headline accuracy numbers, leaving the statistical-comparability claim unsupported.","tokens_in":14756,"tokens_out":1793,"would_cite":true,"duration_ms":16358,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Vision6D claims that by interactively dragging a 3D model over a 2D photograph, a user can produce 6D pose annotations statistically comparable to dataset ground truth—mean angular errors of 4.77 degrees on Linemod and 5.88 degrees on…","keywords":["6D pose estimation","pose annotation","3D-to-2D registration","interactive visualization","user study","pinhole camera model","Linemod","HANDAL"],"falsifier":"Take a set of images from Linemod or HANDAL, deliberately perturb the supplied $K$ (for example, scale focal lengths by 5 to 10 percent while leaving the principal point fixed), have users annotate with Vision6D, and measure ADD against ground truth. If mean error does not grow substantially, the claim survives; if it grows in proportion to the perturbation, the claim's dependence on exact intrinsics is confirmed.","tokens_in":13785,"feed_emoji":"🖱️","tokens_out":4249,"duration_ms":35718,"temperature":0.7,"pith_summary":"Vision6D is an open-source desktop tool that lets a person drag a 3D object model on top of a 2D photograph until its projected outline matches the object in the image, and then exports the implied 6D camera pose. The paper argues this is the first purpose-built interactive 3D-to-2D annotation system for 6D pose estimation, and backs it with a user study on the Linemod and HANDAL datasets. In the study, each user's best of three annotations landed within mean angular errors of 4.77 degrees (Linemod) and 5.88 degrees (HANDAL) of dataset ground truth, with mean annotation times around 100 seconds. If true, this gives pose-estimation researchers a way to generate training labels for custom or prerecorded scenes using only the camera intrinsic matrix, without fiducial markers or motion capture.","feed_headline":"Drag a 3D model onto a photo and get a 6D pose within ~5 degrees","feed_subtitle":"Users can annotate 6D poses in about 100 seconds with accuracy rivaling dataset ground truth, with no markers required.","key_machinery":"The load-bearing object is the pinhole projection model of Equation (3), $P_{uvw} = K M P_w$, where $K$ is the camera intrinsic matrix and $M=[R|t]$ is the unknown extrinsic pose the user is annotating. Given $K$, aligning the rendered 3D model to the 2D image determines $M$ by construction. The supporting machinery is the multi-view interface: a free-navigation scene camera for spatial understanding, an original-camera view locked to the photograph, and immediate re-projection on drag, so the user closes the loop between 3D manipulation and 2D evidence.","core_discovery":"The central discovery is that a 3D-to-2D interactive alignment task, mediated by the pinhole projection equation $P_{uvw} = K[R|t]P_w$, is accurate and efficient enough for humans to reproduce dataset-grade 6D poses. Users manipulate meshes in a 3D scene while a synchronized original-camera view re-renders the projection in real time; the final matrix $[R|t]$ is read out directly from the aligned model. Across 11 participants, annotations were statistically comparable to ground truth, with inter-personal mean ADD of 15.27 mm on Linemod and 30.53 mm on HANDAL, intra-personal repeatability around 7 to 10 degrees of angular distance, and NASA-TLX and SUS feedback indicating moderate workload and generally good usability. The paper frames this as filling the gap between existing 3D GUIs and the need for fast, marker-free 6D pose labels.","pith_inferences":["If the camera intrinsics are uncertain, the tool's practical ceiling may be set by calibration quality rather than user skill; a calibration-refinement step or in-tool estimation of $K$ would be the natural next test.","The same interaction loop could be extended to semi-automatic pose propagation, where the user annotates only the first frame and tracked 2D features propose subsequent poses for correction.","Combining Vision6D with a learned initial pose estimate from a rough detector or keypoint-based PnP could cut the roughly 100-second annotation time substantially, turning the tool into a refinement interface.","The best-of-three protocol suggests that majority voting or an automated consistency check across repeated annotations could further improve final pose quality."],"forward_implications":["Pose labels can be produced for arbitrary images and custom objects without fiducial markers or known camera extrinsics, requiring only the intrinsic matrix $K$.","First-frame annotations can seed video sequences or downstream tracking and pose-estimation training pipelines.","Texture-less and cluttered objects, such as those in Linemod-Occluded scenes, can be annotated by direct 3D-to-2D overlay rather than by feature matching.","Mean annotation times around 100 seconds per sample make small-scale dataset labeling practical for research groups without specialized capture equipment.","The reported accuracy and repeatability metrics support using Vision6D annotations as pseudo-ground-truth when dataset ground truth is unavailable."],"supporting_citations":[{"why":"Supplies the pinhole image-formation model that grounds the projection equation used for 3D-to-2D alignment.","marker":"[9]"},{"why":"Provides the Linemod scenes and ground-truth poses against which user annotations are compared.","marker":"[10]"},{"why":"Provides the HANDAL real-world category-level dataset and its ground-truth poses for the more challenging evaluation.","marker":"[11]"},{"why":"Supplies the Linemod-Occluded dataset used to demonstrate multi-object annotation in cluttered scenes.","marker":"[21]"},{"why":"Provides the NASA-TLX questionnaire used to measure perceived cognitive workload during annotation.","marker":"[22]"},{"why":"Provides the SUS usability scale used to assess how intuitive and learnable the interface is.","marker":"[23]"}],"fun_headline_variants":["Drag 3D models onto photos for 6D pose annotations","Interactive 3D-to-2D alignment yields accurate 6D poses","6D pose labeling tool rivals ground truth accuracy","Open-source Vision6D makes 6D pose labeling intuitive","Annotation tool creates 6D poses via drag-and-drop"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire pipeline assumes the camera intrinsic matrix $K$ is known and correct for every image; any error in focal length or principal point shifts the projection in Equation (3) and biases every annotation, no matter how perfect the user's visual alignment.","fun_headline_variants_meta":{"raw":{"variants":["Drag 3D models onto photos for 6D pose annotations","Interactive 3D-to-2D alignment yields accurate 6D poses","6D pose labeling tool rivals ground truth accuracy","Open-source Vision6D makes 6D pose labeling intuitive","Annotation tool creates 6D poses via drag-and-drop"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00022,"raw_usage":{"total_tokens":1494,"prompt_tokens":1038,"completion_tokens":456,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":654,"completion_tokens_details":{"reasoning_tokens":370}},"tokens_in":654,"tokens_out":456,"duration_ms":4152,"temperature":1.0,"reasoning_tokens":370,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:29:40.077643+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of images from Linemod or HANDAL, deliberately perturb the supplied $K$ (for example, scale focal lengths by 5 to 10 percent while leaving the principal point fixed), have users annotate with Vision6D, and measure ADD against ground truth. If mean error does not grow substantially, the claim survives; if it grows in proportion to the perturbation, the claim's dependence on exact intrinsics is confirmed.","supporting_citations":[{"cited_title":"Peng, Image Formation","cited_arxiv_id":null,"evidence_quote":"Supplies the pinhole image-formation model that grounds the projection equation used for 3D-to-2D alignment."},{"cited_title":"Model based training, detection and pose estimation of texture-less 3d objects in heavily cluttered scenes,","cited_arxiv_id":null,"evidence_quote":"Provides the Linemod scenes and ground-truth poses against which user annotations are compared."},{"cited_title":"Handal: A dataset of real-world manipulable object categories with pose annotations, affordances, and reconstructions,","cited_arxiv_id":null,"evidence_quote":"Provides the HANDAL real-world category-level dataset and its ground-truth poses for the more challenging evaluation."},{"cited_title":"Learning 6d object pose estimation using 3d object coordi- nates,","cited_arxiv_id":null,"evidence_quote":"Supplies the Linemod-Occluded dataset used to demonstrate multi-object annotation in cluttered scenes."},{"cited_title":"Sus: A quick and dirty usability scale,","cited_arxiv_id":null,"evidence_quote":"Provides the SUS usability scale used to assess how intuitive and learnable the interface is."}],"review_version":1}