REVIEW 40 references
AppleGrowthVision: A large-scale stereo dataset for phenological analysis, fruit detection, and 3D reconstruction in apple orchards
T0 review · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read AppleGrowthVision provides 9,317 calibrated stereo images and 31,084 apple labels across six BBCH growth stages, plus benchmarks showing detection and stage-classification improvements.
desk verdict A genuinely useful new apple orchard dataset, but the headline BBCH classification accuracy is inflated by label leakage and the detection benchmarks lack error bars. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
The labels for apples are bounding boxes. Some were drawn by hand, but most were produced by a YOLOv8 object detector and then reviewed by people for a subset of images. The paper reports that adding this dataset to existing apple datasets improves fruit detection scores, and that image classifiers can predict the coarse BBCH stage with very high accuracy. It also shows one 3D reconstruction of a tree row built from the stereo images.
The strongest claims are plausible but not fully proven. The classification test splits images randomly, so pictures of the same tree on the same day are likely in both training and testing, making the accuracy numbers look better than they would on a new orchard. Detection results come from single runs without error bars, and the 3D reconstruction has no quantitative evaluation. The dataset, if released with clean labels and proper splits, would still be a useful resource for precision agriculture.
Extended reading notes
Core claim
AppleGrowthVision is the first publicly available dataset capturing apple orchards throughout a complete phenological cycle with calibrated stereo imagery and expert-validated BBCH growth stages. Adding it to MinneApple improves YOLOv8 F1-score by 7.69% relative, and adding it to MinneApple and MAD improves Faster R-CNN F1-score by 31.06% relative; six principal BBCH stages are classified with over 95% accuracy by standard ImageNet-pretrained networks.
Load-bearing premise
The dataset utility and the classification claim rest on the assumption that the BBCH growth stage assigned to a random subset of images from a given date is representative of all images taken on that date and, more broadly, that the stage is uniform across the orchard sections photographed. Section 3.3 states that the expert classified only a randomly selected subset of images per date, and that trees outside the random sample can vary in secondary stage. If the principal stage also varies spatially on a date, then many images carry incorrect stage labels and the reported >95% classification accuracy, which uses a random image split, would not transfer to new orchards.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Assumptions & free parameters
assumptions (3)
- domain assumption BBCH stage labels assigned to a random subset of images on a date are representative of all images taken that date.
- domain assumption Auto-generated YOLOv8 annotations that were not human-corrected are accurate enough to be used as training ground truth.
- domain assumption Calibrated stereo extrinsics provide reliable initial poses for COLMAP reconstruction.
Cite this review
Pith. "Pith review of AppleGrowthVision: A large-scale stereo dataset for phenological analysis, fruit detection, and 3D reconstruction in apple orchards." pith.science (2026). https://pith.science/paper/RB3LSFWU
@misc{pith2026250514029,
author = {Pith},
title = {Pith review of: AppleGrowthVision: A large-scale stereo dataset for phenological analysis, fruit detection, and 3D reconstruction in apple orchards},
year = {2026},
howpublished = {\url{https://pith.science/paper/RB3LSFWU}},
note = {Machine review of arXiv:2505.14029}
}
read the original abstract
Deep learning has transformed computer vision for precision agriculture, yet apple orchard monitoring remains limited by dataset constraints. The lack of diverse, realistic datasets and the difficulty of annotating dense, heterogeneous scenes. Existing datasets overlook different growth stages and stereo imagery, both essential for realistic 3D modeling of orchards and tasks like fruit localization, yield estimation, and structural analysis. To address these gaps, we present AppleGrowthVision, a large-scale dataset comprising two subsets. The first includes 9,317 high resolution stereo images collected from a farm in Brandenburg (Germany), covering six agriculturally validated growth stages over a full growth cycle. The second subset consists of 1,125 densely annotated images from the same farm in Brandenburg and one in Pillnitz (Germany), containing a total of 31,084 apple labels. AppleGrowthVision provides stereo-image data with agriculturally validated growth stages, enabling precise phenological analysis and 3D reconstructions. Extending MinneApple with our data improves YOLOv8 performance by 7.69 % in terms of F1-score, while adding it to MinneApple and MAD boosts Faster R-CNN F1-score by 31.06 %. Additionally, six BBCH stages were predicted with over 95 % accuracy using VGG16, ResNet152, DenseNet201, and MobileNetv2. AppleGrowthVision bridges the gap between agricultural science and computer vision, by enabling the development of robust models for fruit detection, growth modeling, and 3D analysis in precision agriculture. Future work includes improving annotation, enhancing 3D reconstruction, and extending multimodal analysis across all growth stages.
Figures
Reference graph
Works this paper leans on
-
[1]
Apple dataset benchmark from orchard environment in modern fruiting wall, 2019
Santosh Bhusal, Manoj Karkee, and Qin Zhang. Apple dataset benchmark from orchard environment in modern fruiting wall, 2019. Open-access dataset. 2
work page 2019
-
[2]
End-to- end object detection with transformers, 2020
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to- end object detection with transformers, 2020. 4, 7
work page 2020
-
[3]
Pengyu Chu, Zhaojian Li, Kaixiang Zhang, Dong Chen, Kyle Lammers, and Renfu Lu. O2rnet: Occluder-occludee relational network for robust apple detection in clustered orchard environments.arXiv preprint arXiv:2303.04884,
-
[4]
Martin Churuvija, Ranjan Sapkota, Dawood Ahmed, and Manoj Karkee. A pose-versatile imaging system for com- prehensive 3D modeling of planar-canopy fruit trees for au- tomated orchard operations.Computers and Electronics in Agriculture, 230:109899, 2025. 3
work page 2025
-
[5]
An image is worth 16x16 words: Transformers for image recognition at scale, 2021
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale, 2021. 4, 7
work page 2021
-
[6]
Model-Based Camera Calibration Using Analy- sis by Synthesis Techniques
Peter Eisert. Model-Based Camera Calibration Using Analy- sis by Synthesis Techniques. InVMV, pages 307–314, 2002. 3
work page 2002
-
[7]
M. A. Ellis, D. C. Ferree, R. C. Funt, and L. V . Madden. Effects of an apple scab-resistant cultivar on use patterns of inorganic and organic fungicides and economics of disease control.Plant Disease, 82(4):428–433, 1998. 6, 7
work page 1998
-
[8]
Tian Gao, Feiyu Zhu, Puneet Paul, Jaspreet Sandhu, Henry Akrofi Doku, Jianxin Sun, Yu Pan, Paul Staswick, Harkamal Walia, and Hongfeng Yu. Novel 3D Imaging Sys- tems for High-Throughput Phenotyping of Plants.Remote Sensing, 13(11):2113, 2021. 3
work page 2021
Show all 40 references
-
[9]
Au- tomatisierte frucht-und pflanzenerkennung in apfelplantagen durch k¨unstliche intelligenz
Michael Gerstenberger, Mykyta Kovalenko, David Prze- wozny, Jannes Magnusson, Eike Gassen, Jakub Pawlak, Jochen Hirth, Laura von Hirschhausen, Detlef Runde, Anna Hilsmann, Peter Eisert, and Sebastian Bosse. Au- tomatisierte frucht-und pflanzenerkennung in apfelplantagen durch ...
-
[10]
H., Bleiholder L., Meier U., Schnockfricke U., Weber E., and Witzenberger A
Hack. H., Bleiholder L., Meier U., Schnockfricke U., Weber E., and Witzenberger A. Einheitliche codierung der ph ¨anologischen entwicklungsstadien von kultur- und schadpflanzen- erweiterte bbch-skala.Nachrichtenblatt des Deutschen Pflanzenschutzdienstes, 44:265–270, 1992. 2
1992
-
[11]
Minneapple: A benchmark dataset for apple detection and segmentation
Nicolai H ¨ani, Pravakar Roy, and V olkan Isler. Minneapple: A benchmark dataset for apple detection and segmentation. arXiv preprint arXiv:1909.06441, 2020. 2
1909 arXiv
-
[12]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, 2016. 5
2016
-
[13]
Le, and Hartwig Adam
Andrew Howard, Mark Sandler, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Ruoming Pang, Vijay Vasudevan, Quoc V . Le, and Hartwig Adam. Searching for mobilenetv3, 2019. 7 7
2019
-
[14]
Weinberger
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kil- ian Q. Weinberger. Densely connected convolutional net- works. In2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2261–2269, 2017. 5
2017
-
[15]
S3ad: Semi-supervised small apple de- tection in orchard environments
Robert Johanson, Christian Wilms, Ole Johannsen, and Si- mone Frintrop. S3ad: Semi-supervised small apple de- tection in orchard environments. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), United States, 2024. IEEE. Equal contribu- ...
2024
-
[16]
3D Gaussian Splatting for Real-Time Radiance Field Rendering.ACM Trans
Bernhard Kerbl, Georgios Kopanas, Thomas Leimkuehler, and George Drettakis. 3D Gaussian Splatting for Real-Time Radiance Field Rendering.ACM Trans. Graph., 42(4):1–14,
-
[17]
Ground- ing image matching in 3d with mast3r
Vincent Leroy, Yohann Cabon, and J´erˆome Revaud. Ground- ing image matching in 3d with mast3r. InEuropean Confer- ence on Computer Vision, pages 71–91. Springer, 2024. 6
2024
-
[18]
Toward sustainability: Trade-off between data quality and quantity in crop pest recognition
Yang Li and Xuewei Chao. Toward sustainability: Trade-off between data quality and quantity in crop pest recognition. Frontiers in Plant Science, V olume 12 - 2021, 2021. 2
2021
-
[19]
LightGlue: Local Feature Matching at Light Speed
Philipp Lindenberger, Paul-Edouard Sarlin, and Marc Polle- feys. LightGlue: Local Feature Matching at Light Speed. In 2023 IEEE/CVF International Conference on Computer Vi- sion (ICCV), pages 17581–17592, 2023. ISSN: 2380-7504. 6
2023
-
[20]
D.G. Lowe. Object recognition from local scale-invariant features. InProceedings of the Seventh IEEE International Conference on Computer Vision, pages 1150–1157 vol.2, Kerkyra, Greece, 1999. IEEE. 6
1999
-
[21]
A survey of public datasets for computer vision tasks in precision agriculture.Computers and Electronics in Agriculture, 178:105760, 2020
Yuzhen Lu and Sierra Young. A survey of public datasets for computer vision tasks in precision agriculture.Computers and Electronics in Agriculture, 178:105760, 2020. 2
2020
-
[22]
Lancashire, Uta Schnock, Reinhold Stauß, Theo van den Boom, Elfriede We- ber, and Peter Zwerger
Uwe Meier, Hermann Bleiholder, Liselotte Buhr, Carmen Feller, Helmut Hack, Martin Heß, Peter D. Lancashire, Uta Schnock, Reinhold Stauß, Theo van den Boom, Elfriede We- ber, and Peter Zwerger. The bbch system to coding the phe- nological growth stages of plants – history and p...
2009
-
[23]
CherryPicker: Semantic Skeletonization and Topological Reconstruction of Cherry Trees
Lukas Meyer, Andreas Gilson, Oliver Scholz, and Marc Stamminger. CherryPicker: Semantic Skeletonization and Topological Reconstruction of Cherry Trees. InIEEE/CVF Conference on Computer Vision and Pattern Recogni- tion Workshops (CVPRW), pages 6244–6253, Vancouver, Canada, 2023. 3, 6
2023
-
[24]
Fruitnerf: A unified neural radiance field based fruit counting framework
Lukas Meyer, Andreas Gilson, Ute Schmidt, and Marc Stam- minger. Fruitnerf: A unified neural radiance field based fruit counting framework. InIROS, 2024. 6
2024
-
[25]
Recognition of phenological development stages of apple blossoms using computer vision
Xuan Khanh Nguyen, Bastian Braun, Nico Heider, and Mar- tin Schieck. Recognition of phenological development stages of apple blossoms using computer vision. In45. GIL- Jahrestagung, Digitale Infrastrukturen f ¨ur eine nachhaltige Land-, Forst- und Ern ¨ahrungswirtschaft, pages...
2025
-
[26]
Lorenz Gunreben Nico Heider. A survey of datasets for computer vision in agriculture.Proceedings of the 45th GIL Annual Conference (GIL-Jahrestagung): Digi- tale Infrastrukturen f ¨ur eine nachhaltige Land-, Forst- und Ern¨ahrungswirtschaft, 2025. 2
2025
-
[27]
Girshick, and Jian Sun
Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. InAdvances in Neural Information Pro- cessing Systems (NeurIPS), pages 91–99, 2015. 4
2015
-
[28]
Mobilenetv2: Inverted residuals and linear bottlenecks
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zh- moginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In2018 IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 4510–4520, 2018. 5
2018
-
[29]
From coarse to fine: Robust hierarchical localization at large scale
Paul-Edouard Sarlin, Cesar Cadena, Roland Siegwart, and Marcin Dymczyk. From coarse to fine: Robust hierarchical localization at large scale. InCVPR, 2019. 6
2019
-
[30]
SuperGlue: Learning feature matching with graph neural networks
Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. SuperGlue: Learning feature matching with graph neural networks. InCVPR, 2020. 6
2020
-
[31]
Structure-from-Motion Revisited
Johannes Lutz Sch ¨onberger and Jan-Michael Frahm. Structure-from-Motion Revisited. InConference on Com- puter Vision and Pattern Recognition (CVPR), 2016. 6
2016
-
[32]
Very Deep Con- volutional Networks for Large-Scale Image Recognition
Karen Simonyan and Andrew Zisserman. Very Deep Con- volutional Networks for Large-Scale Image Recognition. arXiv:1409.1556 [cs], 2015. arXiv: 1409.1556. 5
2015 arXiv
-
[33]
Mingxing Tan, Ruoming Pang, and Quoc V . Le. Efficientdet: Scalable and efficient object detection, 2020. 7
2020
-
[34]
Terven and Diana M
Juan R. Terven and Diana M. Cordova-Esparza. A com- prehensive review of yolo architectures in computer vision: From yolov1 to yolov8 and yolo-nas.Machine Learning and Knowledge Extraction, 5(4):83, 2024. 7
2024
-
[35]
Disk: Learning local features with policy gradient
MichałTyszkiewicz, Pascal Fua, and Eduard Trulls. Disk: Learning local features with policy gradient. InAdvances in Neural Information Processing Systems, pages 14254– 14265. Curran Associates, Inc., 2020. 6
2020
-
[36]
Streif, and J
Meier U., Graf H., Hack H., Hess M., Kennel W., Klose R., Mappes D., Seipp D., Strauss R. Streif, and J. Van den Boom. Ph ¨anologische entwicklungsstadien des kernob- stes (malus domestica borkh. und pyrus communis l.), des steinobstes (prunus-arten), der johannisbeere (ribes-...
1994
-
[37]
Localizing small apples in complex apple orchard environ- ments.arXiv preprint arXiv:2202.11372, 2022
Christian Wilms, Robert Johanson, and Simone Frintrop. Localizing small apples in complex apple orchard environ- ments.arXiv preprint arXiv:2202.11372, 2022. 2
2022 arXiv
-
[38]
ImageNet Training in Minutes
Yang You, Zhao Zhang, Cho-Jui Hsieh, James Demmel, and Kurt Keutzer. ImageNet Training in Minutes. InProceedings of the 47th International Conference on Parallel Processing, New York, NY , USA, 2018. Association for Computing Ma- chinery. event-place: Eugene, OR, USA. 5
2018
-
[39]
Xilei Zeng, Hao Wan, Zeming Fan, Xiaojun Yu, and Hen- grong Guo. MT-MVSNet: A lightweight and highly accurate convolutional neural network based on mobile transformer for 3D reconstruction of orchard fruit tree branches.Expert Systems with Applications, 268:126220, 2025. 3 8
2025
-
[70]
Gesellschaft f¨ur Informatik eV , 2024. 2
2024
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.