{"id":"ff0fefb6-3447-4d1d-9c80-3e9983ade0b8","arxiv_id":"2505.14029","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"AppleGrowthVision provides 9,317 calibrated stereo images and 31,084 apple labels across six BBCH growth stages, plus benchmarks showing detection and stage-classification improvements.","lead":"A research team has built a large image dataset of apple orchards across a full growing season, using paired stereo cameras and expert-labeled growth stages. It is offered as a resource for training computer vision models to spot fruit, track tree development, and build 3D orchard models.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"BBCH stage labels may not be representative of all images in each date, undercutting the >95% classification claim.","rationale":"The reader's weakest assumption and my load-bearing concern coincide on the representativeness of the BBCH label assignment. That is the right concern: the dataset paper's main scientific novelty is the expert-validated phenological staging across a full growth cycle, and the headline quantitative claim supporting that novelty is the >95% classification accuracy. The concern lands because the labeling protocol in Section 3.3 explicitly assigns one principal stage per date based on a random subset, and Table 2 shows near one-to-one date-to-stage mapping, which makes temporal leakage and feature memorization a concrete mechanism rather than a hypothetical. The detection experiments are also weak (no error bars, no cross-site generalization, and the 31.06% number is a relative comparison against a degenerate MinneApple+MAD baseline), but they are secondary to the dataset's stated contribution and the authors themselves describe them as an initial evaluation. The 3D reconstruction claim is unsupported quantitatively, but the paper explicitly frames it as a baseline for future work, so it does not carry the central weight. Verdict stays CONDITIONAL because the resource itself is plausibly valuable and the requested check is runnable; if the held-out-date split shows robust accuracy, the concern would be resolved and the dataset claim would substantially strengthen.","tokens_in":10219,"tokens_out":1449,"duration_ms":12656,"concrete_test":"Re-run the BBCH classification experiments with a strict temporal or tree-grouped split: train on images from a subset of capture dates (or from the Brandenburg subset only) and test on images from held-out dates (or from the Pillnitz subset only), with no same-tree/same-date images shared between train and test. If accuracy drops substantially below 95%, the stage labels are not generalizing phenologically. Also report per-class accuracy for the original random-split experiment: if the classifier simply memorizes date-specific background features, confusion will concentrate at boundaries such as BBCH 57 vs 60 or 67 vs 71/72.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The >95% BBCH classification claim (Table 4) depends on reliable per-image stage labels for all 9,317 images. Section 3.3 states that only a randomly selected subset of images per date was manually classified by an expert, with trees outside the random sample varying in secondary stage. The paper then treats every image captured on that date as belonging to the principal stage assigned to the subset, so the network may be learning date-specific visual features (background, lighting, camera position) rather than phenological stage. This is supported by the configuration details: each date corresponds to essentially one or two BBCH stages (Table 2), so a random 80/10/10 split at image level leaks multiple views of the same tree, taken seconds apart, into both training and test partitions. As a result, a model could achieve >95% accuracy by temporal and viewpoint memorization, and reported accuracy would not transfer to images from unseen dates, trees, or orchards. The reader flagged this as the weakest assumption; it is indeed the most load-bearing risk because the dataset's claimed novel scientific value is the expert-validated BBCH annotation enabling phenological analysis, and the headline classification accuracy is the quantitative evidence for that value.","agreement_with_reader":"agree"},"referee_report":null,"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: AppleGrowthVision is a genuinely useful resource—the first public apple orchard dataset I know of that pairs calibrated stereo images across a full growth cycle with expert BBCH stage labels. The resource is the contribution; the algorithmic pieces are all standard, and the paper is honest that the 3D reconstruction section is a baseline demo rather than a validated pipeline. If you work in agricultural computer vision, this is worth a look.\n\nThat said, the headline numbers need caution. The classification experiment leaks labels: Section 3.3 says only a random subset of images per date was expert-classified, yet Table 2 maps each date to one or two BBCH stages, and the paper treats all images from that date as belonging to that stage. The random 80/10/10 image split then puts near-duplicate views of the same trees, taken seconds apart, into both training and test. The >95% accuracy mostly measures date and viewpoint memorization, not true phenological generalization. This is the paper's most load-bearing weakness, because the expert-validated BBCH label is the claimed novel value.\n\nThe detection experiments are better but still soft: single runs, no error bars, and the training data mixes auto-generated labels (YOLOv8 with partial human correction), so part of the gain from adding AppleGrowthVision could just be more data or label noise. The 31% F1 improvement for Faster R-CNN is relative to a particularly weak baseline (MinneApple+MAD alone), so it sounds larger than it is.\n\nWhat the paper does well: the dataset itself, the documentation of the stereo setup, the honest limitations section, and the clear use cases. The 3D reconstruction is a proof-of-concept with no quantitative metric—minor for a dataset paper, but the claim should be downgraded.\n\nRecommendation: the dataset deserves peer review. The evaluation flaws are fixable—split by tree or date, report error bars, and clarify exactly which images carry expert labels versus propagated labels. The resource itself is valuable and worth engaging with seriously.","headline":"A genuinely useful new apple orchard dataset, but the headline BBCH classification accuracy is inflated by label leakage and the detection benchmarks lack error bars.","tokens_in":10952,"tokens_out":2091,"would_cite":true,"duration_ms":23242,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":null,"created_at":"2026-08-07T15:41:39.215988+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":null,"supporting_citations":[],"review_version":1}