{"id":"d29cfbb1-a5b2-4d29-84c1-799c05620953","arxiv_id":"2509.11740","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"An integrated mobile manipulator system achieved 98.3% pick-and-place success in 724 mock supermarket stocking events, at 68.2 seconds per item, still slower than human workers.","lead":"Robotics engineers built a supermarket stocking robot from off-the-shelf parts, integrating vision, planning, and control software, and tested it on a mock store shelf. It completed about 98% of individual pick-and-place moves in controlled overnight tests, but still needed roughly ten times longer per item than a human stocker, showing the remaining gap to commercial use.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 98.3% object-level success rate is not well-defined: no explicit success rubric, no raw failure counts, and unclear treatment of navigation failures make the headline metric unverifiable.","rationale":"The reader's weakest_assumption is transferability of the 98% figure to real shelves with tight packing and diverse products. That is a valid external-validity concern. However, the most load-bearing issue for the central claim is more basic: the success metric itself is undefined, so even the lab performance cannot be verified. The paper's own text exposes this: Section V.C reports 'minor placement inaccuracies' as task-level partial failures without clarifying object-level classification, and Section V.D claims 'including ... shelf navigation' without an object-level navigation metric. The reader's rationale does mention 'a defined success metric' as a needed addition, but the formal weakest_assumption does not identify it as the primary concern. I agree with the reader's overall conditional posture—the system exists, the experiments are disclosed, and the limitations are listed—but I would strengthen the condition: the 98.3% success rate must be accompanied by an explicit success rubric and raw counts before it can be taken at face value. The proposed concrete test (re-analysis with a defined metric and raw data) would settle whether the number survives minimal scrutiny. The verdict remains CONDITIONAL because the paper can be revised to provide these clarifications; the underlying contribution is not rejected.","tokens_in":11362,"tokens_out":4988,"duration_ms":54162,"concrete_test":"Obtain from the authors the per-object outcome log for all 724 operations (success/failure, retry count, per-item time, task ID), and a written success rubric specifying, e.g., product orientation, allowable displacement from target position, fronting criterion, and whether navigation failures or partial placements count as failures. Then recompute the success rate and its 95% confidence interval, and separately recompute it excluding tasks with navigation failures. If the 98.3% figure changes materially, or if the rubric counts minor misplacements as successes, the headline claim overstates the system's reliability.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central quantitative claim—98.3% object-level success over 724 stocking operations (Sec. V.D)—cannot be evaluated because the paper never defines what counts as a successful 'stocking operation'. Section V.C reports 'object-level success rate exceeding 98%' but gives no success/failure counts, no confidence intervals, and no tolerance for placement accuracy. The text says 'minor placement inaccuracies linked to small ArUco detection errors' cause task-level partial failures, yet it is unclear whether these same inaccuracies are counted as object-level failures or as successes. Likewise, one task completely failed due to 'total loss of the ArUco marker', but the number of objects in that task and whether they appear in the 724 denominator are not stated. The phrase 'including pick and place operations and shelf navigation' in Sec. V.D is not supported by the experimental description, which evaluates individual pick-and-place operations, not navigation failures. Without a precise success definition and raw data, the 98.3% figure is untestable, even within the controlled lab setting. The paper's own admission that 'all failures involved products lacking a dedicated YOLO model' further conditions the result on a curated product set, but this is secondary to the missing metric definition.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an integrated autonomous supermarket stocking system built on the Hello Robot Stretch 3 mobile manipulator. The architecture combines ROS2, Behavior Trees, Nav2 with a spatio-temporal voxel layer, ArUco-based localization with a Kalman filter, a two-step MPC for base/head control, and a perception pipeline using YOLO and SAM2. The authors report laboratory experiments in a mock supermarket: 70 stocking tasks totaling 724 individual stocking operations, an object-level success rate exceeding 98% (and 98.3% in the discussion), a task-level success rate above 87%, and comparisons against teleoperation and human workers, including a joint time-cost performance index and a hardware/software gap decomposition. Several components are released as open-source ROS2 packages.","tokens_in":11668,"tokens_out":3946,"duration_ms":48764,"significance":"If the reported success rates are reproducible and well-defined, the paper would provide a useful baseline for an affordable, integrated retail stocking platform and a quantitative comparison against humans and teleoperation. The open-source release of the docking, navigation, and simulator packages is a concrete strength, as is the focus on a commercially available robot rather than custom high-end hardware. The paper is primarily a systems integration contribution, and the experimental evaluation is central to its claims.","major_comments":[{"comment":"The central claim '98.3% object-level success, including pick and place operations and shelf navigation' is not verifiable as stated. No success rubric is defined: what counts as a successful stocking operation at the object level (placement tolerance, orientation, contact, gripper release, planogram compliance)? No raw success/failure counts by stage are given, and no confidence intervals are reported. The text says 8 of 9 unsuccessful tasks were partial failures due to 'minor placement inaccuracies' but does not say whether those inaccuracies are scored as object-level successes or failures. One task completely failed due to total loss of the ArUco marker; the number of objects in that task and whether they are in the 724 denominator are not stated. The phrase 'including pick and place operations and shelf navigation' is also unsupported: the experimental section reports individual sto","section":"Sec. V.C–V.D"},{"comment":"The reported 98% success rate is conditioned on a narrow product set and an artificial spacing requirement: 'Due to gripper size ... we focused on cans and small boxes' and 'placement required a spacing of 3cm between items, exceeding the common practice of approximately 3mm.' The paper also states that 'all failures involved products lacking a dedicated YOLO model,' meaning failures are concentrated in the GPT-4o-based perception branch. The claim that the system demonstrates 'reliable performance' for supermarket stocking is therefore broader than the evidence supports. The claims should be restricted to the tested product/spacing regime, or additional experiments under tighter spacing and with a more diverse product set should be provided.","section":"Sec. V.B"},{"comment":"The hardware/software gap decomposition—'approximately 78% of the lost performance should be attributed to hardware constraints and 22% to software limitations'—is asserted without a method. No equation, data table, sensitivity analysis, or error propagation is provided to show how these percentages are derived from the timing data. This decomposition is presented as one of the paper's contributions ('a hardware/software gap decomposition outlining concrete paths toward commercially viable retail robots'). It needs a transparent calculation or should be removed.","section":"Sec. V.D"},{"comment":"The statistical evaluation is incomplete. The paper reports point estimates (98% object-level, 87% task-level) over 70 tasks and 724 operations but gives no per-task variation, no confidence intervals, and no error bars in Fig. 7. The comparisons with Spahn et al. [6] and Wu et al. [11] cite their success rates without adjusting for different protocols or task definitions. Please add confidence intervals for the proportions (e.g., Clopper-Pearson) and clearly state the number of objects per task and the definition of task success, so the headline numbers are interpretable and comparable.","section":"Sec. V.C"}],"minor_comments":[{"comment":"The text says 'Finally, Section IV presents extensive laboratory tests' but the experiments are presented in Section V. The section numbering is inconsistent.","section":"Introduction"},{"comment":"The heading 'Performace' should be 'Performance'.","section":"Sec. V.D"},{"comment":"Typos in function names: 'Get_Bouning_Boxes' appears twice; should be 'Get_Bounding_Boxes'.","section":"Algorithm 1"},{"comment":"The performance index 'pi = 1000 / (Annual Cost · Time Per Item)' would benefit from explicit units and a definition of Annual Cost (USD/year) and Time Per Item (seconds). The text uses p_i,AR, p_i,H, p_i,T without defining the subscript convention.","section":"Eq. (13)"},{"comment":"Minor typo: 'unicyle' should be 'unicycle'.","section":"Sec. III.A"}],"recommendation":"major_revision","confidential_remarks":"The paper is essentially a systems integration contribution with an empirical evaluation. The missing success definition and the unsupported 78/22 decomposition are load-bearing and can be fixed with additional analysis and clarity; I do not see a fundamental flaw that would require rejection. The title/abstract inconsistency between the arXiv metadata and the manuscript's actual title should also be resolved during revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"If you work on mobile manipulation in retail, this is worth a look. The authors put a complete ROS2 pipeline on a Stretch 3, let it run 724 stocking operations overnight, and report 98.3% object-level success plus a side-by-side cost/time comparison with human and teleoperated stocking. That comparison is the most honest part of the paper: they flat-out say robots are ten times slower, and they give a joint time-cost index. The integration itself is the real contribution—BT planning, Nav2, ArUco-KF localization, two-step MPC, YOLO/SAM2 perception—and they open-source the key modules. They also clearly state their own limitations: only cans and small boxes, 3cm spacing instead of the typical 3mm, and all failures involved products lacking a dedicated YOLO model. That is good practice.\n\nNow the soft spots. The headline metric is under-defined. The paper reports 'object-level success rate exceeding 98%' and later '98.3%', but there is no explicit rubric for what counts as a successful stocking operation, no raw success/failure counts, no confidence intervals, and no tolerance for placement accuracy. The text says minor placement inaccuracies caused partial task failures—are those counted as object-level successes? The one fully failed task from total ArUco loss presumably takes some objects with it, but we don't know how many. In a lab setup with no humans, the environment is controlled, so the numbers are probably directionally right, but they are not independently checkable as reported. This is a fixable reporting problem, not a fatal flaw.\n\nSecond, the 78/22 hardware/software gap decomposition in Section V.D is asserted without any derivation. It looks like a rough accounting trick rather than a measured result. I'd treat it as an opinion, not a finding.\n\nThird, the comparison with Spahn et al. and Wu et al. is fair, but note that Spahn's 85% excludes placing, so the '2-3x faster pick execution' claim needs care—they are comparing different scopes. Minor.\n\nOverall, this is a solid systems paper that would benefit from a proper evaluation section: define success, give per-module failure counts, add error bars, and either derive the hardware/software decomposition or drop it. I'd send it to peer review—the empirical data and open-source release deserve referee time, and the authors seem willing to engage with criticism. If you're building a baseline for retail stocking, cite it, but don't cite the 98.3% without checking the appendix.","headline":"A genuinely useful, honest systems paper on retail restocking; the 98.3% headline is plausible but not yet verifiable because the paper never pins down what counts as a success.","tokens_in":12178,"tokens_out":2273,"would_cite":true,"duration_ms":23787,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A robot built from commodity hardware can autonomously stock supermarket shelves with 98% object-level reliability, though it remains ten times slower than human workers.","keywords":["autonomous shelf stocking","mobile manipulation","behavior trees","model predictive control","object detection","visual localization","performance benchmarking","retail automation"],"falsifier":"Run the same pipeline in a store aisle where items are spaced ~3 mm apart and include a wider range of package geometries; if object-level success drops materially below 98% for a comparable number of events, the central claim fails. A cheaper, paper-consistent check: count failures among products without a dedicated YOLO model—the paper itself predicts this is where all failures occur.","tokens_in":11234,"feed_emoji":"🤖","tokens_out":4690,"duration_ms":51341,"temperature":0.7,"pith_summary":"This paper sets out to show that autonomous supermarket shelf stocking can be done by a mobile manipulator built from affordable, off-the-shelf parts, without requiring a bespoke research platform. It reports 98.3% object-level success across 724 pick-and-place operations in a mock store run as an overnight closed-door scenario, with a mean full-pipeline time of 68.2 seconds per item. The same experiments benchmark the robot against teleoperation and human workers, showing that autonomy beats teleoperation on a combined time-cost index but still trails human stocking by roughly a factor of ten. The value of the paper is less the absolute numbers than the reproducible baseline: it identifies where the gap lives and openly measures it.","feed_headline":"Robot stocks shelves with 98% accuracy, ten times slower than humans","feed_subtitle":"Affordable mobile manipulator rests on marker-based navigation and fine-tuned vision; 78% of the speed gap is hardware.","key_machinery":"The mechanism that carries the argument is the 'last-meter' navigation stack: a 14-state Kalman filter fusing ArUco marker poses with a constant-velocity model feeds a two-step Model Predictive Controller that drives the robot base toward the shelf while steering the pan-tilt head to keep the marker in view. This solves the problem that the base lidar cannot see objects closer than 20 cm. Around that core, a Behavior Tree orchestrates docking, navigation, picking, and planogram-aware placement, while perception uses SAM2 segmentation, fine-tuned YOLO per-SKU classifiers, and an optional GPT-4o fallback for products without a trained model.","core_discovery":"The central claim is that a modular pipeline—behavior-tree task planning, ArUco-marker localization fused through a Kalman filter, a two-step model-predictive controller for base and head camera, and a perception stack pairing fine-tuned YOLO classifiers with SAM2 segmentation and tracking—can carry a Stretch 3 mobile manipulator through the full pick-navigate-place cycle in a supermarket-like environment with 98.3% object-level success over 724 events and a mean of 68.2 s per item. The paper further argues that most remaining failures come from products without a dedicated YOLO model, that task-level success reaches 87%, and that a joint cost-time index places autonomous stocking between te","pith_inferences":["The reported experiments use 3 cm spacing between placed items, whereas real store shelves commonly pack items at about 3 mm; until the same pipeline is tested at real packing density, 98% should be read as an upper bound.","The GPT-4o fallback's low accuracy and 6.9 s inference time, versus 0.2 s for fine-tuned YOLO, suggests an inexpensive testable improvement: an online mechanism that flags repeatedly-misclassified SKUs for model training, rather than relying on a generic vision-language model.","If two or more of these robots were coordinated overnight in a single store, the architecture's task queue could scale throughput almost linearly; the paper does not address multi-robot collision or task allocation, but nothing in the reported design prevents that experiment."],"forward_implications":["If the 98.3% reliability transfers beyond the mock store, overnight restocking of canned and boxed goods is economically plausible: the robot's cost-time index is about a third of human performance but an order of magnitude better than teleoperation.","The failure concentration among products lacking a dedicated YOLO model implies that expanding the per-SKU model library is the most direct route to higher task-level success.","The 78%/22% hardware/software decomposition suggests that gripper and depth-sensing upgrades will shrink the speed gap more than any algorithmic change.","The open-sourced modules and simulator give other groups a concrete baseline to test new perception or control components against the same 68.2 s/item and 98.3% figures."],"fun_headline_variants":["Robot stocks shelves 98% accurately, but humans still win on speed","Shelf-stocking robot: 98% success, but slower than human labor","98% pick accuracy for robot stocker, but humans stay faster","Autonomous stocker: 98% accuracy, but 10x slower than humans"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The 98% success rate depends on simplified products (cans and small boxes) and simplified packing (3 cm spacing instead of the ~3 mm common in stores); if those constraints are removed, the figure is not guaranteed by the paper's data.","fun_headline_variants_meta":{"raw":{"variants":["Robot stocks shelves 98% accurately, but humans still win on speed","Shelf-stocking robot: 98% success, but slower than human labor","98% pick accuracy for robot stocker, but humans stay faster","Autonomous stocker: 98% accuracy, but 10x slower than humans"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000846,"raw_usage":{"total_tokens":3511,"prompt_tokens":729,"completion_tokens":2782,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":473,"completion_tokens_details":{"reasoning_tokens":2699}},"tokens_in":473,"tokens_out":2782,"duration_ms":21821,"temperature":1.0,"reasoning_tokens":2699,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T16:43:21.784295+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same pipeline in a store aisle where items are spaced ~3 mm apart and include a wider range of package geometries; if object-level success drops materially below 98% for a comparable number of events, the central claim fails. A cheaper, paper-consistent check: count failures among products without a dedicated YOLO model—the paper itself predicts this is where all failures occur.","supporting_citations":[],"review_version":1}