{"id":"e5699bc9-1807-4aec-93b9-b1cc873aace7","arxiv_id":"2508.13775","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"MR6D is a new benchmark of 92 real industrial scenes for 6D pose estimation on mobile robots, and current unseen-object pipelines achieve only 0.35 average recall with ground-truth masks.","lead":"MR6D is a new dataset for testing robots that estimate the position and orientation of large industrial objects from mobile platforms. Current 6D pose pipelines score noticeably lower on it than on household-object benchmarks, suggesting it captures real perception challenges.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Benchmark validity rests on unverified ground-truth pose accuracy; the paper's own §3.2 quality statement is based on visual checks, not independent measurement.","rationale":"I agree with the reader's weakest assumption: the benchmark's validity depends on ground-truth pose accuracy, and the paper provides no independent measurement of it. The reader also mentions a missing controlled comparison to existing datasets, but that is secondary to the annotation-accuracy question—even a controlled comparison would be uninterpretable if the reference poses contain systematic errors. The paper itself flags the relevant limitations in Sec. 3.2: visual checks only, deformable objects, pallet dimensional variation, manual scale refinement for O³dyn, and fully manual scaling for MR-like. These are not accusations of carelessness; they are unverified components of an otherwise useful resource. The proposed test is a bounded, feasible validation step: independently re-measure a sample of poses and compare. If errors are within the claimed low-centimeter range, the central claim holds and the conditional acceptance can be lifted; if not, the reported AR numbers and cross-subset comparisons would need re-evaluation. The reader's CONDITIONAL verdict is appropriate, and my analysis does not change it.","tokens_in":9965,"tokens_out":2983,"duration_ms":33337,"concrete_test":"Select a random sample of 3-5 scenes from each subset. Independently recover object poses using a high-accuracy reference (e.g., a laser tracker or a fixed, calibrated LiDAR scan with ICP fitting of the provided meshes; for the dynamic subset use an independent VICON marker set or manual re-annotation). Compare these references to the published annotations, computing per-object translation error (cm) and rotation error (deg). If the median translation error exceeds ~5 cm or median rotation error exceeds ~5°—both plausible given stated deformations—the 'low-centimeter' claim is unsupported and the benchmark validity needs qualification; if errors are below these thresholds, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"MR6D's central claim—that it is a valid benchmark on which current unseen-object pipelines underperform—requires trustworthy ground-truth 6D poses. The annotation section (Sec. 3.2, 'Quality of the annotations') provides only: 'Based on our annotation process and visual checks, we expect static scenes to be accurate within the low-centimeter range.' No independent reference measures this error. The dynamic subset tracks both camera and objects with VICON but can have larger deviations when markers are occluded; the O³dyn subset uses VGGT trajectories whose scale is manually refined when odometry is noisy, and the MR-like subset scales manually. The statement also concedes that object IDs 9, 10, 15 deform and Euro pallets vary by centimeters. If unmeasured pose errors are larger than the low-centimeter estimate—particularly for long-range scenes where small angular errors translate into large translation errors at the object—the reported AR values could be artifacts of the annotations rather than properties of the scenes. Since BOP metrics compare predictions to these annotations, this is the most load-bearing assumption for the dataset's validity.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MR6D, a dataset and benchmark for 6D pose estimation in mobile robotics, comprising 92 real-world scenes with 16 objects across static and dynamic subsets captured from low-mounted, long-range, and occluded perspectives. The authors evaluate two unseen-object pipelines on the dataset: FoundationPose with ground-truth masks and FoundationPose with CTL-generated masks, reporting BOP metrics (AR 0.3462 and 0.1841 on test subsets, respectively). The central claims are that MR6D fills a gap in mobile-robot-relevant pose estimation benchmarks and that current unseen-object pipelines underperform in these settings.","tokens_in":10125,"tokens_out":3586,"duration_ms":37056,"significance":"If the ground-truth annotations are trustworthy, MR6D addresses a genuine gap: existing BOP-style datasets focus on small household objects and arm-mounted cameras, whereas MR6D targets long-range, low-perspective, large-object, and occluded conditions relevant to industrial mobile robots. The dataset is publicly released, uses BOP-compatible format and metrics, and provides 3D meshes for all objects. The authors also identify segmentation as a bottleneck in fully unseen pipelines, which is a useful finding. However, the strength of these contributions depends on the accuracy of the pose annotations and on a controlled comparison to existing datasets, both of which are currently lacking or underdeveloped.","major_comments":[{"comment":"The central validity claim of the benchmark rests on unmeasured ground-truth accuracy. The paper states that 'Based on our annotation process and visual checks, we expect static scenes to be accurate within the low-centimeter range' without any independent reference measurement. The dynamic subset relies on VICON tracking with possible marker occlusions, and the O³dyn and MR-like subsets use VGGT trajectories whose scale is manually refined when odometry is noisy. Given that BOP metrics compare predictions directly against these annotations, unquantified errors—especially at long range where small angular errors translate into large translation errors—could account for the reported AR values. The authors should provide a quantitative annotation-error analysis, e.g., against an independent high-precision reference or via multi-annotator agreement, and report per-subset error bounds.","section":"3.2, 'Quality of the annotations'"},{"comment":"The claim that current 6D pipelines 'underperform' in mobile-robot settings is not supported by a controlled comparison. The paper reports AR only on MR6D; it does not run the same pipelines on established BOP datasets (e.g., YCB-V, T-LESS, ITODD) under identical protocols. Without such baselines, the absolute AR values (0.3462 with GT masks, 0.1841 with CTL masks) cannot be interpreted as evidence that the mobile-robot setting is especially challenging, as opposed to reflecting properties of the chosen methods or of the annotation quality.","section":"4.1, Table 1"},{"comment":"The evaluation that underlies the 'underperform' conclusion uses only one pose estimator (FoundationPose) and one segmentation method (CTL). The paper's abstract and conclusion generalize to 'current 6D pipelines', but a single pipeline is insufficient to establish a systematic deficiency. At minimum, the authors should add at least one more unseen-object pose estimator (e.g., MegaPose or GigaPose) and one more segmentation method (e.g., CNOS), or explicitly state that the results hold only for the FoundationPose/CTL combination.","section":"4.1, Section 2"}],"minor_comments":[{"comment":"The point labeled 'MR6D (All)' is the mean over subsets rather than a global per-scene average, which could mislead readers; the caption discloses this, but the figure itself should use distinct markers or a footnote.","section":"Figure 2"},{"comment":"References [9] and [10] cite the same BOP 2018 paper twice; please consolidate into a single reference or clearly differentiate the 2018 and 2020 BOP publications.","section":"References [9] and [10]"},{"comment":"The caption for Figure 9 says 'Quantitative results' but the figure shows qualitative visualizations of pose projections; the caption should read 'Qualitative results'.","section":"Figure 9"},{"comment":"The statement that the IKEA objects 'will remain available through at least the end of 2026' is a useful reproducibility detail, but it also underlines that the dataset's long-term availability may depend on commercial product lifecycles; consider noting this in the dataset documentation.","section":"Section 3.1"},{"comment":"The sentence attributing poor results to 'poor initialization due to depth values for FoundationPose' is not substantiated by an experiment; either add evidence or soften the claim.","section":"Section 4.2"}],"recommendation":"major_revision","confidential_remarks":"The paper uses CTL, a segmentation method from the same group, and reuses objects from MTevent, also from the same group. This is not circular, because the benchmark's value does not depend on CTL's success, but the self-citation pattern is worth monitoring. The ground-truth accuracy issue is the main risk: if an independent evaluation later shows larger annotation errors than the 'low-centimeter' estimate, the reported AR values could be artifacts. The paper would be strengthened by adding a controlled comparison on existing BOP datasets and at least a second pose estimator. Given that the paper is already accepted at an ICCV workshop, the authors may have limited space, but these points are essential for the journal version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: MR6D is a real, usable dataset that fills a gap in 6D pose estimation benchmarks. If you work on mobile robotics perception, it's worth knowing about. The main caveat is that the ground-truth pose accuracy is never independently measured, which is the one issue I'd want addressed before relying on it.\n\nWhat's new: the dataset targets large objects at long range, low camera perspectives, self-occlusion, and dynamic scenes with a moving camera and moving objects. That's a genuinely different regime than YCB-V, T-LESS, or the household-focused BOP datasets. They built 92 scenes with 16 objects, provide meshes (BundleSDF, manual for the pallet), and release it on HuggingFace. The annotation pipeline uses VICON where available, BOP toolkit for manual refinement, and VGGT plus manual scale for the odometry-only subsets. That is a reasonable recipe and they document it clearly.\n\nThe evaluation is honest but thin. They run FoundationPose with GT masks and with CTL masks, report BOP AR numbers per subset. The numbers are low (0.35 with GT masks, 0.18 with CTL masks), which supports the claim that mobile-robot settings are harder than the household benchmarks. What they don't do is run the same pipeline on an existing dataset like YCB-V or T-LESS under the same protocol, so the \"underperform\" claim is not actually controlled. That is a real gap, but it doesn't weaken the dataset itself.\n\nThe bigger soft spot is the ground-truth quality. Section 3.2 says static scenes are 'expected to be accurate within the low-centimeter range' based on visual checks, but there is no independent measurement against, say, a laser tracker or a CAD-based re-projection error audit. For long-range scenes, sub-degree angular errors turn into multi-centimeter errors at 2 meters, and they admit that some objects deform and pallets vary. The dynamic subset, with both camera and object tracked by VICON, is probably fine, but the O³dyn and MR-like subsets rely on manually scaled trajectories where the error could be larger. For a benchmark, this is the load-bearing assumption. It's not fatal, but it's worth pushing on: measure the error on a few scenes and report it.\n\nThe self-citation concern is minor. CTL is from the same group, but they compare it to DINOv2 and the benchmark value doesn't depend on CTL's success. Reusing MTevent objects is fine because the tasks don't overlap.\n\nBottom line: this is a useful resource for the mobile-robotics and pose-estimation community. It deserves serious peer review, with the request that the authors measure and report annotation error, and ideally add a controlled comparison on at least one existing dataset. I'd accept it with revisions.","headline":"MR6D is a genuinely useful dataset for mobile-robot 6D pose estimation, with the main caveat that ground-truth annotation accuracy is asserted rather than measured.","tokens_in":10719,"tokens_out":2389,"would_cite":true,"duration_ms":23420,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MR6D introduces a benchmark that exposes the gap between current 6D pose estimators and mobile-robot perception needs.","keywords":["6D pose estimation","mobile robotics","benchmark dataset","unseen object pose estimation","industrial objects","long-range perception","occlusion","robotic perception"],"falsifier":"Measure the ground-truth poses against an independent high-accuracy reference, such as a laser tracker or a second motion-capture system, on a sample of MR6D scenes—especially the static subsets; if typical errors exceed the stated low-centimeter bound by a wide margin, the reported underperformance could be an artifact of annotation noise rather than genuine algorithm limits.","tokens_in":9732,"feed_emoji":"🤖","tokens_out":7718,"duration_ms":71949,"temperature":0.7,"pith_summary":"The paper introduces MR6D, a benchmark dataset for 6D object pose estimation aimed at mobile robots rather than fixed robot arms. It contains 92 real-world scenes with 16 industrial objects—pallets, storage bins, and consumer containers—captured from low, distant, and moving viewpoints that match how mobile platforms perceive their surroundings. The central claim is that these conditions are underrepresented in existing pose benchmarks, and the paper backs this by showing that current pose pipelines designed for unseen objects reach only modest accuracy on MR6D: about 0.35 average recall with ground-truth segmentation and about 0.18 when segmentation must be predicted as well. If the benchmark is sound, it gives mobile-robotics research a realistic testbed for a perception problem that household-object datasets do not exercise.","feed_headline":"6D pose pipelines underperform on new mobile-robot benchmark","feed_subtitle":"92 real scenes with large, distant, occluded industrial objects; even perfect segmentation gives only ~0.35 recall.","key_machinery":"The load-bearing artifact is the dataset itself, MR6D, with its annotation methodology. Four capture setups produce four subsets, and each subset's ground truth is built through a different chain: motion-capture tracking with eye-in-hand calibration for the static validation and dynamic sets; a learned multi-view reconstruction model plus odometry-derived scale, manually refined where odometry is noisy, for the low-mounted outdoor set; and fully manual scale alignment for the simulated-wheeled-robot set. Object meshes come either from manual design (the Euro pallet) or from a neural reconstruction method applied to high-accuracy depth images. These choices let the authors argue that the scenes are genuinely harder—larger objects, longer ranges, severe self-occlusion, sunlight-degraded depth—so the low reported accuracy reflects the gap between current algorithms and mobile-robot perception needs.","core_discovery":"MR6D's core discovery is that mobile-robotic conditions form a distinct evaluation regime in which current pose estimators generalize poorly. The dataset spans four capture setups: a static validation set and a dynamic set with both camera and object motion tracked by a multi-camera motion capture system; an indoor-outdoor set filmed from a low-mounted robot camera at roughly 40 cm height, with camera trajectories recovered by a feed-forward multi-view reconstruction model and scale adjusted from odometry or manual alignment; and a set simulating a wheeled robot's approach to objects. On two unseen-object evaluation pipelines—one seeded with ground-truth segmentation masks and one fully automatic—the paper reports average recall scores of 0.3462 and 0.1841, and identifies misidentification under occlusion, closely stacked similar objects, and similar-textured faces as recurring failure modes.","pith_inferences":["A natural test of the distance hypothesis: splitting MR6D frames into distance bins and recomputing average recall would show whether long-range depth degradation, rather than viewpoint or scale per se, drives the low scores.","An independent metrological audit of the ground-truth poses, using a laser tracker or precisely dimensioned fiducials on a sample of scenes, would separate annotation noise from genuine algorithm failure.","The paper notes that Euro pallets deviate from nominal dimensions and that a few objects deform slightly; if mesh-based evaluation ignores these tolerances, reported scores may carry a systematic error that per-instance tolerances could absorb."],"forward_implications":["A mobile-robotics-specific pose benchmark now exists, so new methods can be compared on distance, viewpoint, object scale, and occlusion rather than only on household clutter.","Even with perfect segmentation, current unseen-object pose estimators score near 0.35 average recall, leaving substantial room for improvement on these conditions.","The gap between ground-truth-mask and fully automatic pipelines (0.3462 vs 0.1841) shows that better 2D segmentation directly transfers into pose accuracy.","Because the object meshes are released, seen-object pipelines can be trained on synthetic renderings, enabling a fair comparison between methods specialized to mobile robots and the provided baselines.","The identified failure modes—occlusion-induced misidentification, similar stacked objects, and similar-textured faces—point to concrete refinement targets."],"supporting_citations":[{"why":"Supplies the pose estimation backbone whose accuracy is measured in the benchmark.","marker":"[28]"},{"why":"Provides the segmentation method used for the fully automatic unseen-object evaluation pipeline.","marker":"[5]"},{"why":"Provides the manual annotation tool used to refine object poses during ground-truth generation.","marker":"[10]"},{"why":"Defines the average-recall metrics used to score the pose estimation results.","marker":"[11]"},{"why":"Reconstructs the 3D object meshes that serve as models for pose evaluation.","marker":"[27]"},{"why":"Recovers the camera trajectory for the low-mounted outdoor subset before scale refinement.","marker":"[26]"},{"why":"Provides the eye-in-hand calibration between the motion-captured camera frame and the optical frame.","marker":"[24]"}],"fun_headline_variants":["MR6D: mobile robots trip up 6D pose estimators","New benchmark shows 6D pose fails on mobile robots","6D pose estimation stumbles on large distant objects","MR6D reveals 6D pose pipelines fail in mobile robotics","Mobile robot benchmark: 6D pose estimation gap widens"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's validity rests on ground-truth pose accuracy that is visually checked and expected to be within the low-centimeter range for static scenes, but never independently measured; dynamic scenes and manually scaled trajectories can carry larger unquantified errors.","fun_headline_variants_meta":{"raw":{"variants":["MR6D: mobile robots trip up 6D pose estimators","New benchmark shows 6D pose fails on mobile robots","6D pose estimation stumbles on large distant objects","MR6D reveals 6D pose pipelines fail in mobile robotics","Mobile robot benchmark: 6D pose estimation gap widens"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000861,"raw_usage":{"total_tokens":3719,"prompt_tokens":911,"completion_tokens":2808,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":527,"completion_tokens_details":{"reasoning_tokens":2723}},"tokens_in":527,"tokens_out":2808,"duration_ms":17185,"temperature":1.0,"reasoning_tokens":2723,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:10:30.246977+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the ground-truth poses against an independent high-accuracy reference, such as a laser tracker or a second motion-capture system, on a sample of MR6D scenes—especially the static subsets; if typical errors exceed the stated low-centimeter bound by a wide margin, the reported underperformance could be an artifact of annotation noise rather than genuine algorithm limits.","supporting_citations":[{"cited_title":"FoundationPose: Unified 6d pose estimation and tracking of novel objects","cited_arxiv_id":null,"evidence_quote":"Supplies the pose estimation backbone whose accuracy is measured in the benchmark."},{"cited_title":"Learning embeddings with centroid triplet loss for object identification in robotic grasp- ing","cited_arxiv_id":null,"evidence_quote":"Provides the segmentation method used for the fully automatic unseen-object evaluation pipeline."},{"cited_title":"BOP: Benchmark for 6D object pose esti- mation","cited_arxiv_id":null,"evidence_quote":"Provides the manual annotation tool used to refine object poses during ground-truth generation."},{"cited_title":"BOP challenge 2020 on 6D object localization","cited_arxiv_id":null,"evidence_quote":"Defines the average-recall metrics used to score the pose estimation results."},{"cited_title":"BundleSDF: Neural 6-DoF tracking and 3D reconstruction of unknown objects","cited_arxiv_id":null,"evidence_quote":"Reconstructs the 3D object meshes that serve as models for pose evaluation."},{"cited_title":"Vggt: Visual geometry grounded transformer","cited_arxiv_id":null,"evidence_quote":"Recovers the camera trajectory for the low-mounted outdoor subset before scale refinement."},{"cited_title":"Tsai and R.K","cited_arxiv_id":null,"evidence_quote":"Provides the eye-in-hand calibration between the motion-captured camera frame and the optical frame."}],"review_version":2}