Pith. sign in

REVIEW 40 references

AppleGrowthVision: A large-scale stereo dataset for phenological analysis, fruit detection, and 3D reconstruction in apple orchards

T0 review · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read AppleGrowthVision provides 9,317 calibrated stereo images and 31,084 apple labels across six BBCH growth stages, plus benchmarks showing detection and stage-classification improvements.

desk verdict A genuinely useful new apple orchard dataset, but the headline BBCH classification accuracy is inflated by label leakage and the detection benchmarks lack error bars. read the letter →

arxiv 2505.14029 v1 pith:RB3LSFWU submitted 2025-05-20 cs.CV cs.AI

classification cs.CVcs.AI
keywords growthanalysisstagesapplegrowthvisionappledatasetfruitstereo
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

AppleGrowthVision is a new image collection from two apple orchards in Germany. The main part comes from a farm in Brandenburg, where the same 33 trees were photographed with two cameras at the same time, a stereo setup that gives depth information, on 18 dates across the 2022 growing season. A second part comes from a research orchard in Pillnitz with additional labeled images. An expert assigned each date a stage from the standard BBCH scale, which describes how buds, flowers, fruit, and leaves develop.

The labels for apples are bounding boxes. Some were drawn by hand, but most were produced by a YOLOv8 object detector and then reviewed by people for a subset of images. The paper reports that adding this dataset to existing apple datasets improves fruit detection scores, and that image classifiers can predict the coarse BBCH stage with very high accuracy. It also shows one 3D reconstruction of a tree row built from the stereo images.

The strongest claims are plausible but not fully proven. The classification test splits images randomly, so pictures of the same tree on the same day are likely in both training and testing, making the accuracy numbers look better than they would on a new orchard. Detection results come from single runs without error bars, and the 3D reconstruction has no quantitative evaluation. The dataset, if released with clean labels and proper splits, would still be a useful resource for precision agriculture.

Extended reading notes

Core claim

AppleGrowthVision is the first publicly available dataset capturing apple orchards throughout a complete phenological cycle with calibrated stereo imagery and expert-validated BBCH growth stages. Adding it to MinneApple improves YOLOv8 F1-score by 7.69% relative, and adding it to MinneApple and MAD improves Faster R-CNN F1-score by 31.06% relative; six principal BBCH stages are classified with over 95% accuracy by standard ImageNet-pretrained networks.

Load-bearing premise

The dataset utility and the classification claim rest on the assumption that the BBCH growth stage assigned to a random subset of images from a given date is representative of all images taken on that date and, more broadly, that the stage is uniform across the orchard sections photographed. Section 3.3 states that the expert classified only a randomly selected subset of images per date, and that trees outside the random sample can vary in secondary stage. If the principal stage also varies spatially on a date, then many images carry incorrect stage labels and the reported >95% classification accuracy, which uses a random image split, would not transfer to new orchards.

Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claims rest on the validity of extrapolated BBCH labels, the accuracy of partially human-reviewed auto-labels, and the reliability of stereo calibration for COLMAP. No free parameters or invented entities are introduced beyond the dataset itself.

assumptions (3)
  • domain assumption BBCH stage labels assigned to a random subset of images on a date are representative of all images taken that date.
    Section 3.3 says only a random subset of images was categorized per date; the labels are then used for the whole dataset. Local variation would make many labels wrong.
  • domain assumption Auto-generated YOLOv8 annotations that were not human-corrected are accurate enough to be used as training ground truth.
    Section 3.2 describes training YOLOv8 on a small manual subset, generating labels for remaining images, and human review for only 70 stereo and 777 non-stereo images. No label-quality metrics are reported.
  • domain assumption Calibrated stereo extrinsics provide reliable initial poses for COLMAP reconstruction.
    Section 4.3 uses an initial pair from extrinsic calibration priors; no calibration error or reconstruction accuracy is reported.

how reviews work

0 comments
Cite this review

Pith. "Pith review of AppleGrowthVision: A large-scale stereo dataset for phenological analysis, fruit detection, and 3D reconstruction in apple orchards." pith.science (2026). https://pith.science/paper/RB3LSFWU

@misc{pith2026250514029,
  author       = {Pith},
  title        = {Pith review of: AppleGrowthVision: A large-scale stereo dataset for phenological analysis, fruit detection, and 3D reconstruction in apple orchards},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RB3LSFWU}},
  note         = {Machine review of arXiv:2505.14029}
}
read the original abstract

Deep learning has transformed computer vision for precision agriculture, yet apple orchard monitoring remains limited by dataset constraints. The lack of diverse, realistic datasets and the difficulty of annotating dense, heterogeneous scenes. Existing datasets overlook different growth stages and stereo imagery, both essential for realistic 3D modeling of orchards and tasks like fruit localization, yield estimation, and structural analysis. To address these gaps, we present AppleGrowthVision, a large-scale dataset comprising two subsets. The first includes 9,317 high resolution stereo images collected from a farm in Brandenburg (Germany), covering six agriculturally validated growth stages over a full growth cycle. The second subset consists of 1,125 densely annotated images from the same farm in Brandenburg and one in Pillnitz (Germany), containing a total of 31,084 apple labels. AppleGrowthVision provides stereo-image data with agriculturally validated growth stages, enabling precise phenological analysis and 3D reconstructions. Extending MinneApple with our data improves YOLOv8 performance by 7.69 % in terms of F1-score, while adding it to MinneApple and MAD boosts Faster R-CNN F1-score by 31.06 %. Additionally, six BBCH stages were predicted with over 95 % accuracy using VGG16, ResNet152, DenseNet201, and MobileNetv2. AppleGrowthVision bridges the gap between agricultural science and computer vision, by enabling the development of robust models for fruit detection, growth modeling, and 3D analysis in precision agriculture. Future work includes improving annotation, enhancing 3D reconstruction, and extending multimodal analysis across all growth stages.

Figures

Figures reproduced from arXiv: 2505.14029 by the authors.

Figure 1
Figure 1. Stereo camera setup with two calibrated Canon EOS [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Positions per tree in a row on an orchard using a stereo [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Data acquisition in Pillnitz these stages are most relevant for fruit visibility and detec￾tion. To improve the quality of the annotations, human anno￾tators reviewed and corrected the AI-generated labels for 70 images of the stereo-format, and for 777 of the non￾stereo images of the dataset. The final annotations consist of bounding boxes around apples in the images, stored in Darknet/YOLO format used by the Ultral… view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Example images for each existing growth stage in BBCH [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Example of a reconstructed orchard scene of apple trees [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

40 extracted references · 37 canonical work pages

  1. [1]

    Apple dataset benchmark from orchard environment in modern fruiting wall, 2019

    Santosh Bhusal, Manoj Karkee, and Qin Zhang. Apple dataset benchmark from orchard environment in modern fruiting wall, 2019. Open-access dataset. 2

  2. [2]

    End-to- end object detection with transformers, 2020

    Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to- end object detection with transformers, 2020. 4, 7

  3. [3]

    O2RNet: Occluder-Occludee Relational Network for Robust Apple Detection in Clustered Orchard Environments

    Pengyu Chu, Zhaojian Li, Kaixiang Zhang, Dong Chen, Kyle Lammers, and Renfu Lu. O2rnet: Occluder-occludee relational network for robust apple detection in clustered orchard environments.arXiv preprint arXiv:2303.04884,

  4. [4]

    Martin Churuvija, Ranjan Sapkota, Dawood Ahmed, and Manoj Karkee. A pose-versatile imaging system for com- prehensive 3D modeling of planar-canopy fruit trees for au- tomated orchard operations.Computers and Electronics in Agriculture, 230:109899, 2025. 3

  5. [5]

    An image is worth 16x16 words: Transformers for image recognition at scale, 2021

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale, 2021. 4, 7

  6. [6]

    Model-Based Camera Calibration Using Analy- sis by Synthesis Techniques

    Peter Eisert. Model-Based Camera Calibration Using Analy- sis by Synthesis Techniques. InVMV, pages 307–314, 2002. 3

  7. [7]

    M. A. Ellis, D. C. Ferree, R. C. Funt, and L. V . Madden. Effects of an apple scab-resistant cultivar on use patterns of inorganic and organic fungicides and economics of disease control.Plant Disease, 82(4):428–433, 1998. 6, 7

  8. [8]

    Novel 3D Imaging Sys- tems for High-Throughput Phenotyping of Plants.Remote Sensing, 13(11):2113, 2021

    Tian Gao, Feiyu Zhu, Puneet Paul, Jaspreet Sandhu, Henry Akrofi Doku, Jianxin Sun, Yu Pan, Paul Staswick, Harkamal Walia, and Hongfeng Yu. Novel 3D Imaging Sys- tems for High-Throughput Phenotyping of Plants.Remote Sensing, 13(11):2113, 2021. 3

Show all 40 references
  1. [9]

    Au- tomatisierte frucht-und pflanzenerkennung in apfelplantagen durch k¨unstliche intelligenz

    Michael Gerstenberger, Mykyta Kovalenko, David Prze- wozny, Jannes Magnusson, Eike Gassen, Jakub Pawlak, Jochen Hirth, Laura von Hirschhausen, Detlef Runde, Anna Hilsmann, Peter Eisert, and Sebastian Bosse. Au- tomatisierte frucht-und pflanzenerkennung in apfelplantagen durch ...

  2. [10]

    H., Bleiholder L., Meier U., Schnockfricke U., Weber E., and Witzenberger A

    Hack. H., Bleiholder L., Meier U., Schnockfricke U., Weber E., and Witzenberger A. Einheitliche codierung der ph ¨anologischen entwicklungsstadien von kultur- und schadpflanzen- erweiterte bbch-skala.Nachrichtenblatt des Deutschen Pflanzenschutzdienstes, 44:265–270, 1992. 2

  3. [11]

    Minneapple: A benchmark dataset for apple detection and segmentation

    Nicolai H ¨ani, Pravakar Roy, and V olkan Isler. Minneapple: A benchmark dataset for apple detection and segmentation. arXiv preprint arXiv:1909.06441, 2020. 2

  4. [12]

    Deep residual learning for image recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, 2016. 5

  5. [13]

    Le, and Hartwig Adam

    Andrew Howard, Mark Sandler, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Ruoming Pang, Vijay Vasudevan, Quoc V . Le, and Hartwig Adam. Searching for mobilenetv3, 2019. 7 7

  6. [14]

    Weinberger

    Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kil- ian Q. Weinberger. Densely connected convolutional net- works. In2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2261–2269, 2017. 5

  7. [15]

    S3ad: Semi-supervised small apple de- tection in orchard environments

    Robert Johanson, Christian Wilms, Ole Johannsen, and Si- mone Frintrop. S3ad: Semi-supervised small apple de- tection in orchard environments. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), United States, 2024. IEEE. Equal contribu- ...

  8. [16]

    3D Gaussian Splatting for Real-Time Radiance Field Rendering.ACM Trans

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimkuehler, and George Drettakis. 3D Gaussian Splatting for Real-Time Radiance Field Rendering.ACM Trans. Graph., 42(4):1–14,

  9. [17]

    Ground- ing image matching in 3d with mast3r

    Vincent Leroy, Yohann Cabon, and J´erˆome Revaud. Ground- ing image matching in 3d with mast3r. InEuropean Confer- ence on Computer Vision, pages 71–91. Springer, 2024. 6

  10. [18]

    Toward sustainability: Trade-off between data quality and quantity in crop pest recognition

    Yang Li and Xuewei Chao. Toward sustainability: Trade-off between data quality and quantity in crop pest recognition. Frontiers in Plant Science, V olume 12 - 2021, 2021. 2

  11. [19]

    LightGlue: Local Feature Matching at Light Speed

    Philipp Lindenberger, Paul-Edouard Sarlin, and Marc Polle- feys. LightGlue: Local Feature Matching at Light Speed. In 2023 IEEE/CVF International Conference on Computer Vi- sion (ICCV), pages 17581–17592, 2023. ISSN: 2380-7504. 6

  12. [20]

    D.G. Lowe. Object recognition from local scale-invariant features. InProceedings of the Seventh IEEE International Conference on Computer Vision, pages 1150–1157 vol.2, Kerkyra, Greece, 1999. IEEE. 6

  13. [21]

    A survey of public datasets for computer vision tasks in precision agriculture.Computers and Electronics in Agriculture, 178:105760, 2020

    Yuzhen Lu and Sierra Young. A survey of public datasets for computer vision tasks in precision agriculture.Computers and Electronics in Agriculture, 178:105760, 2020. 2

  14. [22]

    Lancashire, Uta Schnock, Reinhold Stauß, Theo van den Boom, Elfriede We- ber, and Peter Zwerger

    Uwe Meier, Hermann Bleiholder, Liselotte Buhr, Carmen Feller, Helmut Hack, Martin Heß, Peter D. Lancashire, Uta Schnock, Reinhold Stauß, Theo van den Boom, Elfriede We- ber, and Peter Zwerger. The bbch system to coding the phe- nological growth stages of plants – history and p...

  15. [23]

    CherryPicker: Semantic Skeletonization and Topological Reconstruction of Cherry Trees

    Lukas Meyer, Andreas Gilson, Oliver Scholz, and Marc Stamminger. CherryPicker: Semantic Skeletonization and Topological Reconstruction of Cherry Trees. InIEEE/CVF Conference on Computer Vision and Pattern Recogni- tion Workshops (CVPRW), pages 6244–6253, Vancouver, Canada, 2023. 3, 6

  16. [24]

    Fruitnerf: A unified neural radiance field based fruit counting framework

    Lukas Meyer, Andreas Gilson, Ute Schmidt, and Marc Stam- minger. Fruitnerf: A unified neural radiance field based fruit counting framework. InIROS, 2024. 6

  17. [25]

    Recognition of phenological development stages of apple blossoms using computer vision

    Xuan Khanh Nguyen, Bastian Braun, Nico Heider, and Mar- tin Schieck. Recognition of phenological development stages of apple blossoms using computer vision. In45. GIL- Jahrestagung, Digitale Infrastrukturen f ¨ur eine nachhaltige Land-, Forst- und Ern ¨ahrungswirtschaft, pages...

  18. [26]

    Lorenz Gunreben Nico Heider. A survey of datasets for computer vision in agriculture.Proceedings of the 45th GIL Annual Conference (GIL-Jahrestagung): Digi- tale Infrastrukturen f ¨ur eine nachhaltige Land-, Forst- und Ern¨ahrungswirtschaft, 2025. 2

  19. [27]

    Girshick, and Jian Sun

    Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. InAdvances in Neural Information Pro- cessing Systems (NeurIPS), pages 91–99, 2015. 4

  20. [28]

    Mobilenetv2: Inverted residuals and linear bottlenecks

    Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zh- moginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In2018 IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 4510–4520, 2018. 5

  21. [29]

    From coarse to fine: Robust hierarchical localization at large scale

    Paul-Edouard Sarlin, Cesar Cadena, Roland Siegwart, and Marcin Dymczyk. From coarse to fine: Robust hierarchical localization at large scale. InCVPR, 2019. 6

  22. [30]

    SuperGlue: Learning feature matching with graph neural networks

    Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. SuperGlue: Learning feature matching with graph neural networks. InCVPR, 2020. 6

  23. [31]

    Structure-from-Motion Revisited

    Johannes Lutz Sch ¨onberger and Jan-Michael Frahm. Structure-from-Motion Revisited. InConference on Com- puter Vision and Pattern Recognition (CVPR), 2016. 6

  24. [32]

    Very Deep Con- volutional Networks for Large-Scale Image Recognition

    Karen Simonyan and Andrew Zisserman. Very Deep Con- volutional Networks for Large-Scale Image Recognition. arXiv:1409.1556 [cs], 2015. arXiv: 1409.1556. 5

  25. [33]

    Mingxing Tan, Ruoming Pang, and Quoc V . Le. Efficientdet: Scalable and efficient object detection, 2020. 7

  26. [34]

    Terven and Diana M

    Juan R. Terven and Diana M. Cordova-Esparza. A com- prehensive review of yolo architectures in computer vision: From yolov1 to yolov8 and yolo-nas.Machine Learning and Knowledge Extraction, 5(4):83, 2024. 7

  27. [35]

    Disk: Learning local features with policy gradient

    MichałTyszkiewicz, Pascal Fua, and Eduard Trulls. Disk: Learning local features with policy gradient. InAdvances in Neural Information Processing Systems, pages 14254– 14265. Curran Associates, Inc., 2020. 6

  28. [36]

    Streif, and J

    Meier U., Graf H., Hack H., Hess M., Kennel W., Klose R., Mappes D., Seipp D., Strauss R. Streif, and J. Van den Boom. Ph ¨anologische entwicklungsstadien des kernob- stes (malus domestica borkh. und pyrus communis l.), des steinobstes (prunus-arten), der johannisbeere (ribes-...

  29. [37]

    Localizing small apples in complex apple orchard environ- ments.arXiv preprint arXiv:2202.11372, 2022

    Christian Wilms, Robert Johanson, and Simone Frintrop. Localizing small apples in complex apple orchard environ- ments.arXiv preprint arXiv:2202.11372, 2022. 2

  30. [38]

    ImageNet Training in Minutes

    Yang You, Zhao Zhang, Cho-Jui Hsieh, James Demmel, and Kurt Keutzer. ImageNet Training in Minutes. InProceedings of the 47th International Conference on Parallel Processing, New York, NY , USA, 2018. Association for Computing Ma- chinery. event-place: Eugene, OR, USA. 5

  31. [39]

    Xilei Zeng, Hao Wan, Zeming Fan, Xiaojun Yu, and Hen- grong Guo. MT-MVSNet: A lightweight and highly accurate convolutional neural network based on mobile transformer for 3D reconstruction of orchard fruit tree branches.Expert Systems with Applications, 268:126220, 2025. 3 8

  32. [70]

    Gesellschaft f¨ur Informatik eV , 2024. 2

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.