{"id":"ad0ba079-e1b8-45a3-a0b3-a8db58cee344","arxiv_id":"2509.00339","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A vision-guided robotic arm sorted four rock types with 97.5% reported success in 40 trials, using an attention-augmented YOLOv8 detector, stereo depth, and D-H kinematics.","lead":"This paper builds a robotic arm with a stereo camera and an improved YOLOv8 vision model to sort rocks by type. In 40 test grasps across four rock types it reports 97.5% average success, but the evidence is thin and no code or data are provided.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 3.5's experimental table is internally inconsistent: sandstone shows 9/10 successful grasps but 10/10 correctly classified/placed, so the 97.5% average lacks a valid per-trial basis.","rationale":"In good faith, the paper is an integration report, and the 97.5% success rate is the only quantitative evidence for the central contribution. The reader's weakest-assumption flags the small sample and missing confidence intervals. My read finds a stronger, internal-consistency problem: the sandstone row of the Section 3.5 table cannot be true under the paper's own definition of classification success. This does not imply misconduct; it means the evidence as reported cannot support the headline number. Because the issue is in the central result and cannot be resolved from the text, the original REJECT verdict stands. I would not move to UNVERDICTED, since the inconsistency is specific and checkable rather than merely missing information.","tokens_in":13537,"tokens_out":4323,"duration_ms":51840,"concrete_test":"Ask the authors for the per-trial raw log or video for all 40 trials, especially the 10 sandstone trials. For each trial, record three separate outcomes: vision classification, grasp success, and placement bin. Recompute the Table row counts. If the failed sandstone grasp was still counted in the 'Correctly classified quantity' column, the table's 10/10 entry is invalid; recompute the headline average with corrected counts and report a confidence interval. If the counts are correct, specify how a non-grasped aggregate was placed in the specified bin.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim (97.5% average grasping and sorting success, with 39/40 correct placements) rests entirely on Section 3.5's four-group experiment. In that table, sandstone has 10 trials, 9 successfully captured, yet 'Correctly classified quantity' is 10/10. The text defines classification success as 'the correct rate of each type being placed in the specified position,' which cannot be true for a trial where the arm failed to grasp the aggregate. One of the two counts must be wrong, or the classification metric was measured at vision level before grasping rather than after physical placement; the paper does not say which. If the sandstone classification count is corrected to at most 9, the reported average changes materially: with a 40-trial denominator, classification accuracy falls from 97.5% to 95%, and the combined average drops to 96.25%. In addition, no detection/classification accuracy metric for the vision model is reported despite the 'close to 100%' recognition claim in the text, so 'comparable classification accuracy' cannot be separated from manipulation success. With only 40 trials and no confidence interval, this internal inconsistency directly undermines the headline percentage.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript describes an integrated autonomous aggregate sorting system built from a six-degree-of-freedom Hiwonder JetArm, an Orbbec Gemini binocular stereo camera, and a ROS-based control stack. The claimed contributions are an attention-augmented YOLOv8 detector for lithology classification, stereo matching for 3D localization, minimum-enclosing-rectangle particle-size estimation, Denavit–Hartenberg kinematic modeling, and hand–eye calibration. Four aggregate types (limestone, granite, sandstone, marble) were physically tested with 10 grasping trials per type. The abstract and conclusion report an average grasping and sorting success rate of 97.5% and comparable classification accuracy. The paper also asserts in §3.5 that recognition accuracy is 'close to 100%.' The central empirical claim is a high lab-demonstrated success rate for a low-cost robotic sorting pipeline.","tokens_in":13840,"tokens_out":4516,"duration_ms":56916,"significance":"If the 97.5% result were rigorously supported, the paper would offer a useful integration example for construction and mining automation using commercially available hardware. The authors do perform physical trials rather than only simulation and create a labeled image dataset of 1219 aggregate images, which are positive aspects. However, the methodological novelty is incremental: attention-augmented YOLOv8, stereo matching, DH kinematics, and hand–eye calibration are standard techniques, and no ablation or baseline comparison is provided. More importantly, the experimental evidence for the headline number is internally inconsistent and statistically very thin. The promise of the system is therefore not established in the present manuscript, and the significance as written is limited.","major_comments":[{"comment":"The table is internally inconsistent. For sandstone it reports 10 trials, 9 successfully captured, but 10/10 correctly classified/placed. The text defines classification success as 'the correct rate of each type being placed in the specified position.' A trial in which the arm fails to grasp cannot result in a correct placement. If the sandstone classification count is corrected to at most 9, the classification average falls from 97.5% to at best 95% (and the combined average to 96.25%). The authors must provide per-trial records and clarify whether classification was scored at the vision level or at final physical placement; as printed, the headline 97.5% lacks a valid per-trial basis.","section":"§3.5, experimental results table"},{"comment":"The claim that recognition accuracy is 'close to 100%' is unsupported by any detection metric. No precision, recall, mAP, confusion matrix, or training/validation split is reported for the attention-augmented YOLOv8 model. Without this, vision classification performance cannot be separated from manipulation success. The 97.5% mean is based on 40 total trials, with one failure in a category moving the mean by 2.5 percentage points; no confidence interval or significance test is given. This is insufficient statistical support for the paper's central quantitative claim.","section":"§3.5 and §5"},{"comment":"The kinematic and calibration derivations are not verifiable as printed. Equations (4.4)–(4.9) are garbled and partially illegible in the manuscript, the D-H parameter table (Table 2) lists only five rows for a claimed 6-DOF arm, and the forward/inverse kinematics verification is described only as 'basically consistent' with no numerical error. Hand–eye calibration reports no transformation error or residual. Since these components are load-bearing for the claimed grasping precision, the technical support for the 97.5% success rate is incomplete.","section":"§4.2 and §5.2"},{"comment":"Experimental conditions are under-specified: no lighting protocol, camera height, gripper geometry, or aggregate pose distribution are described; there is no baseline comparison with manual or mechanical sorting and no throughput or cycle-time data. The conclusion claims productivity, cost, and safety benefits, but the experiments do not measure these. The limited lab setup and the absence of robustness metrics make the generalization claim disproportionate to the evidence.","section":"§3.4–§3.5 and §5.1"}],"minor_comments":[{"comment":"There are numerous typos and grammatical errors, e.g., 'limeston', 'capture process', 'Jeson' instead of Jetson, and inconsistent use of 'grabbing' vs 'grasping'. The manuscript needs careful language editing.","section":"Throughout"},{"comment":"The software section states Ubuntu 20.04, while the hardware section states Ubuntu 18.04 on the Jetson Nano. Please clarify the actual OS versions used.","section":"§3.2–§3.3"},{"comment":"The D-H table has only five rows for a six-degree-of-freedom arm. Either a joint is missing or the kinematic model is for a reduced set; this should be explained.","section":"Table 2"},{"comment":"The text says the 'eyes on the hand' configuration is 'as shown in Fig. 17', but Fig. 17 is the calibration checkerboard; the correct reference appears to be Fig. 22.","section":"§5.2"},{"comment":"Equation (1) is referenced but the displayed equation is missing from the manuscript. Also, reference [16] on irrational numbers is not an appropriate citation for the Pythagorean theorem; a standard geometry reference would be more suitable.","section":"§3.1.1 and references"},{"comment":"The phrase 'comparable classification accuracy' is ambiguous. The abstract should state classification accuracy explicitly and consistently with the experimental table and text.","section":"Abstract and §3.5"}],"recommendation":"reject","confidential_remarks":"The manuscript appears to be an early draft with substantial presentation and reporting gaps. The central empirical claim is contradicted by the experimental table as printed, and the supporting statistical and detection metrics are absent. Even if the table inconsistency were a typo, the lack of per-trial data, detection metrics, confidence intervals, and baselines means the claim cannot be evaluated. I would need to see a completely rewritten paper with a properly reported experiment and supplementary raw data before reconsidering."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Honestly, this is a small experimental paper that claims 97.5% grasping/sorting success on four aggregate types. The system is a straightforward integration of an off-the-shelf robotic arm (Hiwonder JetArm), an Orbbec stereo camera, and YOLOv8 with an attention module. The attention-augmented YOLOv8 for aggregate detection is a legitimate extension, and the physical experiments on limestone, granite, sandstone, and marble are new in the cited literature. So there is a real feasibility demonstration here, just not a new principle.\n\nThe soft spots are load-bearing. The headline number comes from 10 trials per type, no confidence intervals, no detection mAP, and no comparison to a baseline. Worse, the results table in Section 3.5 is internally inconsistent: sandstone is listed as 9/10 successfully captured but 10/10 correctly classified/placed, which directly contradicts the text's definition of classification success as placing in the specified position. Either the table is wrong or the metric was measured on the vision output rather than after placement; the paper doesn't say. That undermines the central claim.\n\nThe kinematics and stereo equations are essentially illegible in the provided text, and no code or trained weights are released, so the 'improved' detection and matching algorithms can't be inspected. These gaps are exactly the kind that a serious referee would need to see closed.\n\nThat said, the engineering approach is sane, the experiment is a real physical test, and the authors acknowledge limitations (small aggregates, texture-confusion). This is a typical conference-level systems paper that needs a major revision: fix the table, add error bars, report detection metrics, and make the methodology legible. I'd send it to peer review rather than desk reject, because the underlying system is plausible and the flaws are correctable. But as it stands, the evidence does not support the headline claim.","headline":"A plausible but thin feasibility study; the 97.5% headline is undercut by an internally inconsistent 40-trial table and missing detection metrics.","tokens_in":14275,"tokens_out":2756,"would_cite":false,"duration_ms":30263,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A vision-guided robotic arm can sort four types of construction aggregate with 97.5% average success in a lab setting.","keywords":["aggregate sorting","robotic arm","computer vision","YOLOv8 object detection","stereo vision","grasping success rate","construction automation","lithology classification"],"falsifier":"Run the same system on 100 or more randomly selected aggregates per type, in varied lighting and with aggregates overlapping or dusty, and record per-type grasp and classification counts. If the average success rate falls clearly below 97.5%, or if granite-versus-limestone confusion reappears at high rate, the paper's claim is not robust. Also report per-class precision and recall from the detector alone.","tokens_in":13473,"feed_emoji":"🤖","tokens_out":5317,"duration_ms":63841,"temperature":0.7,"pith_summary":"This paper builds and tests an autonomous sorting system for construction aggregates—the crushed stone and gravel used in roads and concrete. The proposed pipeline combines an upgraded object-detection model with stereo-camera depth estimation, particle-size measurement, and arm kinematics, so the robot can recognize limestone, granite, sandstone, and marble, compute where each piece is, pick it up, and place it in the correct bin. In a laboratory experiment with 40 grasps, the system succeeded in 39 grasps and 39 classifications, for a 97.5% average success rate. The authors argue this is a step toward replacing manual or fixed mechanical sorting with a flexible, reprogrammable system. The paper is an extension and validation study rather than a new theoretical result.","feed_headline":"Sorts four rock types at 97.5% success","feed_subtitle":"A six-axis robot uses stereo depth and deep learning to separate limestone, granite, sandstone, and marble.","key_machinery":"The system is carried by a perception-to-action chain: an attention-augmented object-detection network (a modified YOLOv8) identifies lithology and 2D position; stereo matching over left and right infrared images produces a depth map and hence 3D coordinates; a minimum-enclosing-rectangle calculation estimates particle size; hand-eye calibration maps camera coordinates into the arm's coordinate frame; and a standard four-parameter kinematic link model converts the target pose into servo joint angles. The design choice that makes the claim credible is the tight coupling of these modules: each step feeds directly into the next, so the measured end-to-end success rate reflects the whole chain r","core_discovery":"The central claim is that a complete, modular aggregate-sorting robot—perception, localization, size measurement, motion control, and classification—can be assembled from off-the-shelf components and a custom-trained deep detector, and that this integrated system reaches high accuracy in a controlled lab setting: 10/10 grasps for limestone, granite, and marble; 9/10 for sandstone; and classification accuracy of 100% for all types except granite at 90%. The one misclassification was granite read as limestone on a small, texture-poor piece, and the one grasp failure was a 1 cm aggregate too small for the gripper. The authors take this as evidence that the approach can sort typical aggregates a","pith_inferences":["The 97.5% figure rests on only 10 trials per category with no confidence intervals; under variable lighting, dust, overlapping rocks, or unseen aggregate types, real-world performance is likely lower.","A natural testable extension is to run the same system on crowded, mixed piles and measure both grasp cycles per hour and failure modes; if small-particle failures dominate, a two-stage gripper or suction end-effector would be the first fix.","The texture-sensitive deep features suggest transfer to recycling sorting, such as separating glass, metal, and plastic, but that would require a new dataset and a different gripper.","Adding per-class precision and recall for the detector would make the claimed near-perfect recognition accuracy reproducible."],"forward_implications":["If the 97.5% figure holds beyond the lab, sorting can be automated for tasks where manual sorting is slow, costly, or hazardous.","Because classification is learned rather than rule-based, the same pipeline can be retrained for other material types.","The failure pattern points to concrete improvements: gripper redesign for small aggregates and stronger texture features for visually similar rocks.","Since end-to-end accuracy depends on detection, localization, and motion, improving any one module should raise the total success rate."],"supporting_citations":[{"why":"Supplies the base object-detection network that the authors modify with attention and improved detection heads.","marker":"[29]"},{"why":"Provides the planar checkerboard calibration method used to obtain the binocular camera's internal and external parameters.","marker":"[36, 37]"},{"why":"Supplies the hand-eye calibration equations that relate camera coordinates to robot coordinates.","marker":"[42, 43]"},{"why":"Provides the minimum-area encasing rectangle algorithm used to estimate aggregate particle size.","marker":"[14]"},{"why":"Supplies the analytical inverse-kinematics solution used to convert target end-effector poses into joint angles.","marker":"[35]"},{"why":"Provides the lithology-identification approach referenced for classifying rock type.","marker":"[30]"},{"why":"Provides the annotation tool used to create the training dataset for detection.","marker":"[13]"}],"fun_headline_variants":["Vision-guided robot sorts aggregates at 97.5% success","Robot sorts four rock types with 97.5% accuracy","Autonomous rock sorter: 97.5% success on four types","AI robot sorts rocks: perfect on 3, 90% on granite"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The headline success rate rests on only 10 trials per aggregate type in a controlled lab setting, so a single misclassification shifts the average by 2.5 percentage points; if those trials are not representative of real sorting conditions, the claimed reliability does not carry over.","fun_headline_variants_meta":{"raw":{"variants":["Vision-guided robot sorts aggregates at 97.5% success","Robot sorts four rock types with 97.5% accuracy","Autonomous rock sorter: 97.5% success on four types","AI robot sorts rocks: perfect on 3, 90% on granite"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0003,"raw_usage":{"total_tokens":1564,"prompt_tokens":734,"completion_tokens":830,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":478,"completion_tokens_details":{"reasoning_tokens":753}},"tokens_in":478,"tokens_out":830,"duration_ms":9479,"temperature":1.0,"reasoning_tokens":753,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T13:40:56.196761+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same system on 100 or more randomly selected aggregates per type, in varied lighting and with aggregates overlapping or dusty, and record per-type grasp and classification counts. If the average success rate falls clearly below 97.5%, or if granite-versus-limestone confusion reappears at high rate, the paper's claim is not robust. Also report per-class precision and recall from the detector alone.","supporting_citations":[],"review_version":1}