{"id":"67fafc57-591e-4339-b1ac-3bbcae5138c0","arxiv_id":"2501.19395","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A compact cherry tomato harvesting robot with a gripper-integrated camera reached 85.0% of fruit in high tunnel field trials in 10.98 seconds on average.","lead":"Researchers built a compact robot with a camera placed between the gripper fingers and used it with an arm-mounted camera to find and reach cherry tomatoes in crowded high tunnels. In field tests it reached 85% of fruit in about 11 seconds, and the design is meant for tight spaces where larger harvesters cannot fit.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The obstacle-free approach plane assumption in Section III-B is unjustified by the depth detection and is the paper's own most frequent failure source; a controlled occlusion test is needed before the 85% cluttered-environment claim is treated as robust.","rationale":"The reader's weakest-assumption identification is precisely the load-bearing risk I see: Section III-B assumes a plane is obstacle-free solely because a depth camera detected a berry along one ray within it, and Section V confirms that collisions from re-planning without occupancy knowledge were the most frequent failure mode. This is not an external disagreement with consensus; it is an internal gap in the planning argument. The paper does have independent support: the dual-camera codesign is concrete, the field experiments are real, and the failure analysis is candid. The 85% reach rate is the central quantitative claim, but it is measured as 'reached' rather than harvested, which the paper states clearly, so I do not treat that as a hidden flaw. The table inconsistencies are real and should be corrected, but they do not by themselves overturn the central design conclusion; they mainly weaken the specific robustness numbers. My proposed test directly targets the weakest assumption: if the success rate collapses when an occluder is placed inside the assumed-free plane, then the approach-pose computation needs an occupancy map or a different geometric guarantee, and the current claim for cluttered high tunnels is conditional on that fix. Since the paper's own discussion already flags this and the reader's verdict is CONDITIONAL, my stress-test does not move the verdict; it sharpens the condition that must be met.","tokens_in":8467,"tokens_out":2608,"duration_ms":28157,"concrete_test":"Place an artificial stem or leaf inside the approach plane defined in Section III-B, about 5 cm laterally from the camera-to-berry ray so that the berry remains visible to the base RGB-D camera, and run the same 20-trial outdoor or 40-trial lab protocol. If success drops materially below 85% due to link collisions, the obstacle-free-plane assumption is the binding failure and the central cluttered-environment claim is not established without occupancy-aware replanning. A control condition with the leaf placed just outside the plane would isolate the plane assumption from general clutter effects.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Section III-B, the initial approach pose is computed in the plane containing the camera-to-berry vector and the vertical axis, and 'a portion of this plane is assumed to be obstacle-free, as the depth camera has already detected the berry within it.' This inference does not follow: depth detection only certifies line-of-sight along one ray, not that the surrounding plane is clear of stems or leaves. The manipulator is then moved through this supposedly free plane without an occupancy map. Section V reports that the most frequent failure was collision of the manipulator link with the environment, occurring 'when there was a need to re-plan initial approach poses because the target was not in view mostly due to corrupted depth measurements,' and that re-planning is performed 'without the knowledge of plant occupancy in space.' Since the central evidence for the collocated-camera design is the 85.0% (17/20) outdoor and 87.5% (35/40) lab success rates, even a few collisions from this assumption can materially change the headline. The comparison against the distal depth camera (60% vs 85–87.5%) is also vulnerable if that comparison conflates this planning assumption with camera placement. The table inconsistencies (corrupted depth 46.7% vs 73.3%; 13x lighting 60% vs 80%) further reduce confidence in the exact reported numbers, though they are secondary to the missing occupancy reasoning.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a robotic harvesting system for cherry tomatoes in high-tunnel environments, combining a mobile base with a 6-DOF arm, a global RGB-D camera, and a small RGB camera collocated between the fingers of a custom pneumatic gripper. The proposed Detect2Grasp pipeline detects fruit with YOLOv7, computes an initial approach pose from the base depth camera, and then uses closed-loop visual servoing with the gripper camera to center and approach the fruit. Experiments in a lab and in an outdoor high tunnel report an average of 85.0% success in reaching fruit (17/20 trials) in 10.98s on average, along with ablations on corrupted depth, lighting intensity, a hanging-vine environment, and a distal depth camera baseline. The paper claims that the collocated RGB camera is more effective for reaching under-canopy fruit than a distal depth camera.","tokens_in":8889,"tokens_out":2550,"duration_ms":26826,"significance":"If the reported results hold, the work is a useful step toward compact, low-cost harvesting robots for cluttered high-tunnel environments. The paper's strengths include direct field experiments, a realistic comparison setup, and ablations that test robustness to sensor noise and lighting changes. The central claim of 85% reach success is a direct measurement and the paper provides a reasonable failure-mode analysis. However, the significance is tempered by the fact that success is defined as reaching rather than plucking the fruit, by small sample sizes without confidence intervals, by unresolved inconsistencies in the reported numbers, and by an unverified planning assumption that the paper itself identifies as a main failure source. The comparison against the distal depth camera also conflates camera placement with differences in planning heuristics, so the claim that the collocated design is more effective needs sharper experimental control.","major_comments":[{"comment":"The obstacle-free assumption for the approach plane is load-bearing and unjustified. Section III-B states that 'a portion of this plane is assumed to be obstacle-free, as the depth camera has already detected the berry within it,' but depth detection along a single ray does not certify that the surrounding plane is free of stems, leaves, or other plant material. This matters because Section V reports that the most frequent failure was collision of the manipulator link with the environment, occurring when re-planning was needed after corrupted depth measurements, and that re-planning is performed 'without the knowledge of plant occupancy in space.' Since the headline 85.0% and 87.5% success rates could shift with even a few additional collisions, the paper should either provide a controlled occlusion test that specifically varies whether the approach plane contains obstacles, or implement and evaluate an occupancy-aware planning step, before the 'cluttered environment' claim is treated as robust.","section":"Section III-B and Section V"},{"comment":"The reported success rates are internally inconsistent. Table I lists 46.7% for 'Base VS on Artificial Plant + Corrupted Depth' while Section IV-B states 'Our system achieved a 73.3% success rate over 15 trials.' Similarly, Table I lists 60.0% for '+ 13x Light Intensity' while Section IV-C states 'The average reaching time with 13x light intensity was 8.84s with 80% success.' These discrepancies are not cosmetic: they affect the interpretation of the ablation results and the robustness claims. The authors must correct the numbers and explain the source of the mismatch (e.g., which trials were included, whether the table or text is the final result) before the quantitative claims can be trusted.","section":"Table I and Sections IV-B, IV-C"},{"comment":"The success metric is 'reaching' the fruit, not harvesting or plucking it. The abstract says the system 'can reach an average of 85.0% of cherry tomato fruit,' and Section IV-E calls this 'successful reaching of real fruit.' However, the paper title and introduction frame the contribution as 'harvesting.' Since the gripper stops when the berry 'significantly fills the image' and the paper explicitly notes that foliage between the gripper fingers is counted as a failure only because it 'would inhibited the downstream harvesting process,' the current experiments do not demonstrate that the system can harvest. The claims should be consistently phrased as reach success, or additional end-to-end harvesting trials should be reported.","section":"Sections I and IV-E"},{"comment":"The comparison between the collocated RGB camera and the distal depth camera is not sufficiently controlled. In the distal depth camera baseline, the visual servoing stage is replaced by a purely geometric trajectory that passes through a point 4 cm offset from the berry, and the stopping criterion differs from the collocated-camera version. Thus the observed difference (60% vs. 85-87.5%) could be due to the different planning and control heuristics, not solely to camera placement. Please either implement the same visual servoing logic with the distal depth camera or explicitly identify and isolate the effect of camera placement from the effect of the control strategy.","section":"Section IV-F and Table I"}],"minor_comments":[{"comment":"There is a typo in 'limited manipulator workpsace' that should read 'workspace.'","section":"Section V"},{"comment":"Success rates are reported without confidence intervals or statistical significance tests. With N=15-20, a 60% vs. 85% difference is not self-evidently significant; please include exact binomial confidence intervals or a suitable test.","section":"Table I"},{"comment":"The detector was fine-tuned on only 126 images all collected from the tip camera in a lab setting, but the same detector is used on the base camera and in outdoor conditions. The paper states that the model transfers well, but a brief quantitative statement about detection recall in the field would strengthen the claim.","section":"Section III-A"},{"comment":"The failure-mode breakdown in Figure 8 is important, but the figure is not described in enough detail in the text. Please clarify whether the percentages are computed per environment or globally, and define the color coding explicitly in the caption.","section":"Figure 8"},{"comment":"The phrase 'This lower reaching time and variance is due to the highly noisy poses resulting in failure' is confusing: it seems to say that failures reduce the measured time, but a stopped or aborted trial may not be a valid time measurement. Please clarify which trials are included in the average time.","section":"Section IV-B"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a relevant problem and has genuine field data, but the unresolved numeric inconsistencies and the untested obstacle-free plane assumption are serious enough that the current version should not be accepted. The authors should be asked to clarify the table/text mismatches, add confidence intervals, and either weaken the harvest claims or add end-to-end trials. The comparison with the distal depth camera also needs tighter experimental control. I would be willing to review a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing to know about this paper: it's a solid systems integration paper, not a breakthrough. The design novelty is real—a small RGB camera mounted between the gripper fingers, used for closed-loop visual servoing, coupled with a base RGB-D camera for global localization. In high tunnels, the system reached 85% of cherry tomatoes in about 11 seconds. That's a genuine field measurement, not a simulation.\n\nWhat's good: the hardware/software codesign is coherent, and the authors did useful ablations—corrupted depth, lighting, hanging vine, and a distal depth camera comparison. The failure analysis in Section V is honest; they identify collision from re-planning without occupancy knowledge as their most frequent failure and list occupancy grid as future work. That's the mark of a fair empirical paper.\n\nThe soft spots are real but not fatal. First, the obstacle-free plane assumption in Section III-B is conspicuous: the depth camera having detected a berry along one ray does not certify that the whole approach plane is clear. The stress-test note is right about that. However, the reported 85% already includes failures caused by that assumption, so it doesn't invalidate the headline. A controlled occlusion test would make the claim more robust, but it isn't a precondition for the claim as stated.\n\nSecond, the table has internal inconsistencies: corrupted depth is 46.7% in Table I but 73.3% in the text; 13x lighting is 60% in the table but 80% in text. That's a red flag for data curation. It should be fixed before anyone quotes the numbers as benchmarks.\n\nThird, \"reaching\" is not \"harvesting.\" The metric is whether the berry ends up between the grippers, not whether it gets plucked. The paper says a grasp without obstruction would lead to a successful pluck, but that's an inference, not a measurement. Fine for a reach study, but the title says \"precision harvesting.\"\n\nFourth, no code or data released, and the detector is trained on 126 images. Minor for a systems paper, but limits reproducibility.\n\nOverall: this is a useful incremental contribution for compact agricultural robotics. It's not going to change the field, but it's a credible, well-documented integration with real field data. I'd send it to peer review, with the expectation that the authors reconcile the table numbers and be precise about reach vs. harvest. I wouldn't cite the headline numbers until that's done.","headline":"A credible compact harvester with a camera-in-gripper design and real field data, but table inconsistencies and a reach-not-pick metric keep it from being benchmark-grade yet.","tokens_in":9354,"tokens_out":2911,"would_cite":false,"duration_ms":28235,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A two-camera harvesting robot reaches 85.0% of cherry tomatoes in high tunnels in about 11 seconds.","keywords":["cherry tomato harvesting","high tunnel","visual servoing","eye-in-hand camera","RGB-D perception","gripper design","YOLOv7","mobile manipulation"],"falsifier":"Clamp a leafy stem across the approach plane used for a berry that remains visible to the depth camera, then run the pipeline: if the arm collides with the stem, the obstacle-free plane assumption is the cause; if it adapts, re-planning is more robust than the paper's failure analysis suggests.","tokens_in":8275,"feed_emoji":"🍅","tokens_out":7061,"duration_ms":63631,"temperature":0.7,"pith_summary":"High-tunnel tomato harvesting requires robots that fit in narrow rows and work in dense foliage. This paper proposes a codesigned compact system: a global RGB-D camera on the mobile base locates fruit and computes an approach pose, and a tiny RGB camera mounted between the gripper fingers provides closed-loop visual feedback for the final reach. The central claim is that this two-camera Detect2Grasp pipeline reaches 85.0% of cherry tomatoes in an outdoor high tunnel in 10.98 seconds on average, and that the collocated gripper camera succeeds where a distal depth camera fails in cluttered under-canopy settings.","feed_headline":"Gripper camera helps robot reach 85% of cherry tomatoes","feed_subtitle":"Two-camera pipeline thrives in tight, cluttered high-tunnel rows where distal depth cameras collide.","key_machinery":"The load-bearing mechanism is the Detect2Grasp state machine: YOLOv7 detects berries in RGB images; depth from the base camera is overlaid to estimate the 3D position; an initial pose is computed from a plane through the camera-to-berry vector and the vertical axis, with approach angles interpolated from a workspace boundary calibration; local rotations and offsets search for the berry if it is not in the tip camera's view; and a PID controller with a dead band centers the berry and drives the end effector forward until the berry's image size indicates it is between the grippers. The hardware counterpart is the gripper itself: a compact four-bar pneumatic gripper with a camera collocated on its central axis, a slender distal link, and a 90-degree bend to avoid singularities.","core_discovery":"Using only global localization and open-loop reaching is infeasible: base-camera-only trials had an average gripper-to-berry error of 6.8 cm. The paper's discovery is that a low-cost RGB camera placed exactly between the gripper fingers, combined with a visual servoing loop that centers the fruit and stops when the fruit fills the image, can close that gap without a high-fidelity depth sensor at the end effector. In the authors' comparison, this collocated design reached 87.5% of artificial fruit in the lab and 85.0% of real fruit outdoors, while a distal depth camera baseline reached only 60% overall and 77.8% on peripheral berries, failing under the canopy due to its larger profile and collisions.","pith_inferences":["The paper's failure analysis implies that building an occupancy grid of the canopy, listed as future work, would remove the most frequent collision cause; this is an editorial extension because the paper does not test it.","The 6.8 cm open-loop error suggests the method is not tied to a specific depth sensor's accuracy, so cheaper or lower-power global sensors could be substituted as long as the berry stays within the tip camera's field of view.","The same architecture is likely transferable to other small, roughly uniform fruits by recalibrating the image-size threshold and changing the gripper fingers, though the paper only demonstrates cherry tomatoes."],"forward_implications":["A compact harvester can rely on a global RGB-D camera for coarse localization and a tiny gripper camera for the fine reach, eliminating the need for a distal depth sensor.","The visual servoing loop absorbs large depth errors, including corrupted measurements of nearly 20 cm, so the base camera only needs to put the target in the tip camera's view.","The collocated camera's advantage is strongest under the canopy: the distal depth camera baseline falls to 60% overall and 77.8% on peripheral berries, while the collocated design keeps 85 to 87.5% success.","The pipeline transfers across lab, hanging-vine, and outdoor high-tunnel settings, with success degrading only mildly under 13x and 20x light intensity (80% in both cases).","The image-size stopping heuristic provides a direct way to know the fruit is between the grippers, which supports downstream plucking without additional force sensing."],"supporting_citations":[{"why":"YOLOv7 real-time detector supplies the bounding boxes used on both base and tip cameras.","marker":"[24]"},{"why":"A tomato harvesting robot with binocular stereo end-effector cameras, used as a contrasting large-form-factor system.","marker":"[11]"},{"why":"A cherry tomato harvesting robot whose distal depth-camera design is compared against.","marker":"[12]"},{"why":"Reports localizing fruit between grippers as a main failure mode, motivating the collocated camera.","marker":"[15]"},{"why":"An autonomous tomato harvesting robot with rotational plucking gripper, a comparison point for gripper design.","marker":"[22]"},{"why":"High-tunnel row-spacing guidelines that define the compactness constraint.","marker":"[6]"}],"fun_headline_variants":["Gripper camera boosts robotic cherry-picking to 85% in high tunnels","Dual-camera robot reaches 85% cherry tomato harvest in cluttered rows","Eye-in-hand camera helps robot pick 85% of cherry tomatoes","Finger-mounted camera outperforms depth sensor for tomato picking","Closed-loop visual servoing with gripper camera hits 85% harvest rate"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The approach assumes the plane from the base camera to the berry is free of stems and leaves; if foliage occupies that plane, the arm collides and the reach fails.","fun_headline_variants_meta":{"raw":{"variants":["Gripper camera boosts robotic cherry-picking to 85% in high tunnels","Dual-camera robot reaches 85% cherry tomato harvest in cluttered rows","Eye-in-hand camera helps robot pick 85% of cherry tomatoes","Finger-mounted camera outperforms depth sensor for tomato picking","Closed-loop visual servoing with gripper camera hits 85% harvest rate"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000659,"raw_usage":{"total_tokens":2943,"prompt_tokens":802,"completion_tokens":2141,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":418,"completion_tokens_details":{"reasoning_tokens":2044}},"tokens_in":418,"tokens_out":2141,"duration_ms":16212,"temperature":1.0,"reasoning_tokens":2044,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T20:14:39.168208+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Clamp a leafy stem across the approach plane used for a berry that remains visible to the depth camera, then run the pipeline: if the arm collides with the stem, the obstacle-free plane assumption is the cause; if it adapts, re-planning is more robust than the paper's failure analysis suggests.","supporting_citations":[{"cited_title":"Development of a tomato harvesting robot used in greenhouse,","cited_arxiv_id":null,"evidence_quote":"A tomato harvesting robot with binocular stereo end-effector cameras, used as a contrasting large-form-factor system."},{"cited_title":"Design and test of robotic harvesting system for cherry tomato,","cited_arxiv_id":null,"evidence_quote":"A cherry tomato harvesting robot whose distal depth-camera design is compared against."},{"cited_title":"Development and evaluation of a pneumatic finger-like end-effector for cherry tomato harvesting robot in greenhouse,","cited_arxiv_id":null,"evidence_quote":"Reports localizing fruit between grippers as a main failure mode, motivating the collocated camera."},{"cited_title":"Development of an autonomous tomato harvesting robot with rotational plucking gripper,","cited_arxiv_id":null,"evidence_quote":"An autonomous tomato harvesting robot with rotational plucking gripper, a comparison point for gripper design."},{"cited_title":"Planting in a high tunnel,","cited_arxiv_id":null,"evidence_quote":"High-tunnel row-spacing guidelines that define the compactness constraint."}],"review_version":1}