REVIEW 3 major objections 5 minor 4 references
Review of Fruit Tree Image Segmentation
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This review of 158 papers finds that fruit tree image segmentation is dominated by task- and environment-specific solutions, and that the field's most noticeable deficiency is the lack of a versatile dataset and segmentation model.
desk verdict Useful field map of front-view fruit-tree segmentation, but the headline deficiency claim leans on a corpus that may not be complete; keep the verdict conditional. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The paper's organizing device is a four-level taxonomy that classifies every collected paper in the fixed order method (rule-based versus deep learning), image type (RGB, RGB-D, point cloud, others), agricultural task (phenotyping, harvesting, spraying, pruning, yield estimation, navigation, thinning, training), and fruit species. This taxonomy is what turns the 158-paper corpus into a map of the field, exposing the dominance of deep learning, the concentration on apples and grapes, and the task-specific fragmentation that underlies the deficiency claim. The companion mechanism is the crawling review itself, a citation-following search procedure analogous to web crawling: it starts with a seed paper, pushes cited papers into a queue, processes them until the queue is empty, and optionally adds recent papers from selected journals. Together these two mechanisms define both the evidence base and the viewpoint from which the review reads the literature.
What would settle it
A comprehensive keyword search of the published literature on front-view fruit tree segmentation, followed by checking whether every result falls inside the 158-paper corpus or is legitimately excluded as fruit-only, forest, or top-view work, would test the review's completeness. Finding even one pre-2023 study that supplies a public dataset or a model used across several agricultural tasks and environments would directly weaken the claim that no versatile dataset or segmentation model exists.
Extended reading notes
Core claim
The paper presents itself as the first review of front-view fruit tree segmentation in the agricultural domain, departing from earlier tree-segmentation surveys oriented to top-view UAV images and digital forestry. Using a crawling review that begins with a seed paper, follows citations until the queue is empty, and supplements with recent issues of three journals from 2020 to 2023, it collected 76 rule-based and 82 deep-learning papers from 1990 to 2023. Classified by method, image type, agricultural task, and fruit, the corpus shows a paradigm shift: rule-based papers peaked around 2015-2018 and then declined, while deep-learning papers appeared around 2018 and kept increasing, with RGB images becoming dominant and harvesting overtaking phenotyping as the most frequent task. The review's central conclusion is that no versatile dataset and no versatile segmentation model exist: the 11 public datasets it lists are each highly specific to one task or environment, so performance results from different papers cannot be objectively compared. It therefore recommends building versatile datasets and models, using few-shot and self-supervised learning, fusing CNNs with transformers, monocular depth estimation, and DL-based 3D reconstruction as routes to a general tree segmentation module.
Load-bearing premise
The review's load-bearing premise is that its crawling search, starting from a single seed paper and supplementing with three journals from 2020 to 2023, captured essentially all relevant front-view fruit tree segmentation work; if a sizable body of such work lies outside that citation graph, the statistics and the deficiency conclusion could misrepresent the field.
Editorial extensions
If this is right
- If the field lacks a versatile dataset and model, then each new agricultural task or environment requires designing, training, and testing a new method, which is a major barrier to applying computer vision broadly in orchards.
- Because there is no standard dataset, objective performance comparison between segmentation methods is currently of little value; a shared benchmark would change that.
- The deep-learning era favors cheap RGB and RGB-D sensors; point clouds and multi-spectral images are increasingly rare, so future systems will likely build on smartphone cameras and low-cost depth sensors.
- Few-shot and self-supervised learning are identified as ways to overcome the scarcity of labeled agricultural data and to move toward a versatile model.
- A versatile dataset, built with horticultural expertise, could act as a de facto standard and motivate challenges that drive the field forward.
Reading between the lines
- If the fragmentation thesis is correct, then building a multi-task, multi-environment benchmark should be the field's first priority; such a benchmark would likely be harder than any existing dataset and would expose current model limitations.
- The review's public-dataset inventory suggests that private task-specific datasets are the norm; a coordinated effort to publish and standardize them could be as valuable as any single algorithmic advance.
- The crawling review's completeness could be tested by repeating it from several different seed papers; if the corpus grows substantially, the statistics and the deficiency claim might need revision.
- The proposed CNN-transformer fusion and monocular depth estimation are concrete, testable next steps: one could directly compare fused versus pure-CNN models on thin branches and occluded fruit, which the review identifies as hard cases.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript surveys front-view fruit tree image segmentation research published between 1990 and 2023. The author introduces a 'crawling review' method: starting from the single seed paper [Chehreh 2023], the bibliography is expanded through citation chains and supplemented by scanning three named journals for 2020-2023, yielding 158 papers. The papers are organized through a taxonomy of method (rule-based vs. deep learning), image type (RGB, RGB-D, point cloud, other), agricultural task (phenotyping, harvesting, spraying, pruning, etc.), and fruit species, and each paper is summarized in Sections 3-4. Statistics on publication date, image type, task, and fruit are presented in Figure 5. The paper's central claim is that the most significant shortcoming in prior work is the absence of versatile datasets and segmentation models that transfer across tasks and environments, and it lists six future research directions. The appendix inventories 11 public datasets and provides background on segmentation methods, performance metrics, and agricultural tasks.
Significance. If the corpus were complete, the paper would provide a useful first synthesis of front-view fruit tree segmentation, an important niche between agricultural robotics and computer vision. The author deserves credit for enumerating and narratively summarizing 158 papers (Tables 1-2), making dataset URLs available in Table A.1, and giving a clear taxonomy that practitioners can follow. The discussion of task-specific sensor choices and the six future directions are sensible. However, the paper's central contribution is an absence claim about the field ('lack of a versatile dataset and segmentation model'), and that claim currently rests on a non-validated manual crawl. The value of the review therefore depends on either making the corpus demonstrably complete or explicitly reframing the conclusion as a property of the collected corpus.
major comments (3)
- [§2.2.1, Abstract, §5.3] The central absence claim in the Abstract and Section 5.3 ('the most noticeable deficiency ... lack of a versatile dataset and segmentation model') is not supported by the evidence presented for the corpus. The crawling phase starts from a single seed paper [Chehreh 2023], which is a UAV/top-view digital-forestry-oriented survey, and the supplementary phase scans only Computers and Electronics in Agriculture, Biosystems Engineering, and Journal of Field Robotics for 2020-2023. Front-view agricultural papers in venues such as Precision Agriculture, Sensors, Remote Sensing, Agronomy, IEEE Access/RAL, and Frontiers in Plant Science can enter the corpus only if they are cited by the seed or its citation descendants, so recent or less-cited papers are systematically excluded. The manuscript states that crawling is 'closer to an exhaustive search' but provides no saturation curve, no recall comparison against a database query, and no reproducibility package. I recommend either adding a validation study that demonstrates recall (e.g., a second crawl from multiple seeds or a WoS/Scopus query with overlap analysis) or softening the conclusion to describe what is observed in the 158-paper corpus, not in the field as a whole.
- [§2.2.2, §2.3, Tables 1-2] The taxonomy and the numerical statistics in Figure 5 depend on manual inclusion decisions (front-view vs top-view, fruit tree vs forest tree, rule-based vs deep learning, task label, fruit label), but the manuscript does not report operational definitions for these decisions, a second annotator, or an inter-rater reliability check. Because Tables 1-2 and Figure 5 are the basis for the qualitative trends in Sections 3-5, misclassifications at this stage propagate into the conclusions. Please provide explicit inclusion/exclusion criteria and at least a small reliability study, or present the tables and statistics as illustrative rather than exhaustive.
- [§5.2-§5.3, Table A.1] The versatility-gap conclusion is not operationalized. Section 5.2 supports the claim that multi-species or multi-task models are essentially absent by citing [Siddique 2022] as the single example 'to the best of our knowledge,' and Section 5.3 supports the dataset-deficiency claim with the 11 datasets in Table A.1. The review does not define what counts as 'versatile' (e.g., number of tasks, number of species, varied illumination/season/architecture), nor does it systematically evaluate each of the 158 papers against that criterion. Without such a criterion, the 'most noticeable deficiency' finding cannot be distinguished from a general impression. I suggest adding a small systematic table or analysis that checks each corpus paper for multi-task/multi-environment evaluation, which would turn this claim into a verifiable statement.
minor comments (5)
- [§3.1.2 / References] The text refers to 'Silwal et al.' for the apple-picking robot, but the reference list entry is [Siwal2017]; please unify the spelling (Abhisesh Silwal).
- [Appendix A.1.3] In the paragraph on AlexNet, 'won first place Krizhevsky 2012]' is missing the opening bracket; it should read '[Krizhevsky 2012].'
- [§4.3.1 and A.3.3] 'tomato and maze' and 'tomatoes and maze' should read 'maize' in both occurrences.
- [§4.4.2] In the description of [Hung 2013], 'conditional random file' should be 'conditional random field (CRF).'
- [§2.2.1] For reproducibility, specify whether the citation-crawling phase had any date restriction, how citation chains were followed (forward citations, backward citations, or both), and how duplicate or inaccessible papers were resolved.
Circularity Check
No significant circularity: the review's conclusions are qualitative syntheses of an explicitly scoped corpus, not quantities derived from inputs.
full rationale
This is a literature review with no mathematical derivation chain, fitted parameters, or predictive model. The central finding—that previous studies lacked a versatile dataset and segmentation model—is stated as a qualitative synthesis of the 158 papers collected by the crawling method (Section 2.2.1) and is explicitly hedged where needed (e.g., 'To the best of our knowledge, fruit or flower segmentation for several different species using a single model is the only example [Siddique 2022]', Section 5.2). A review's conclusions are, by nature, statements about the corpus it examines; that does not make them circular in the sense of the target axis unless the corpus itself is constructed from the conclusion. The crawling protocol is described transparently (single seed [Chehreh 2023], supplementary scan of three journals, 2020–2023), and the paper does not claim that its search is formally exhaustive, only that it is 'closer to an exhaustive search.' The first-review claim is qualified with 'to the best of our knowledge.' No equation is defined in terms of another, no fitted parameter is renamed as a prediction, and no load-bearing result is imported from a self-citation chain. The NIHHS-JBNU dataset entry in Table A.1 is descriptive, not used as evidence for a derived quantity, and the conclusion does not reduce to it. Accordingly, the correct finding is no significant circularity.
Assumptions & free parameters
assumptions (3)
- ad hoc to paper The crawling review method yields a complete and unbiased corpus of 158 relevant papers.
- domain assumption Papers can be consistently classified by the taxonomy (method, image type, task, fruit).
- domain assumption Front-view images define the agricultural domain; top-view UAV imagery is excluded as forestry.
Cite this review
Pith. "Pith review of Review of Fruit Tree Image Segmentation." pith.science (2026). https://pith.science/paper/57QNQMNC
@misc{pith2026241214631,
author = {Pith},
title = {Pith review of: Review of Fruit Tree Image Segmentation},
year = {2026},
howpublished = {\url{https://pith.science/paper/57QNQMNC}},
note = {Machine review of arXiv:2412.14631}
}
read the original abstract
Fruit tree image segmentation is an essential problem in automating a variety of agricultural tasks such as phenotyping, harvesting, spraying, and pruning. Many research papers have proposed a diverse spectrum of solutions suitable to specific tasks and environments. The review scope of this paper is confined to the front views of fruit trees and based on 158 relevant papers collected using a newly designed crawling review method. These papers are systematically reviewed based on a taxonomy that sequentially considers the method, image, task, and fruit. This taxonomy will assist readers to intuitively grasp the big picture of these research activities. Our review reveals that the most noticeable deficiency of the previous studies was the lack of a versatile dataset and segmentation model that could be applied to a variety of tasks and environments. Six important future research tasks are suggested, with the expectation that these will pave the way to building a versatile tree segmentation module.
Figures
Reference graph
Works this paper leans on
-
[2]
class” (object class and confidence information) and “box
Figure A.1. Outline of DL segmentation (a guava tree image is segmented into three classes: fruit, branches, and background [Lin 2022]). In 2012, AlexNet achieved a top -5 classification error rate (15.3%) in the ImageNet large scale visual r ecognition challenge(ILSVRC) and won first place Krizhevsky 2012]. This event motivated most computer vision resea...
work page 2022
-
[4]
Fou r and three of the 11 datasets are for apple trees and grape vines, respectively
Table A.1 summarizes them. Fou r and three of the 11 datasets are for apple trees and grape vines, respectively. One dataset each was found for tomato plants, avocado trees, and capsicum annum plants. All of these except the capsicum annum dataset contain real images. The last dataset, Urban Street Tree, was not constructed in an orchard, but on the stree...
-
[5]
Because the private datasets are small, data augmentation is very important. Usually, a combination of various geometric transformations such as rotation or flipping and various photometric transformations such as adding noise or an intensity change is applied to augment the data. Performance metrics : Once the model has been trained, a performance evalua...
work page 2023
-
[306]
Semantic Segmentation Refinement by Monte Carlo Region Growing of High Confidence Detections
[Chene2012] Yann Chene et al., “On the use of depth camera for 3D phenotyping of entire plants,” Computers and Electronics in Agriculture. [Cheng2020] Zhenzhen Cheng et al., “Interlacing orchard canopy separation and assessment using UAV images,” Remote Sensing. [Coll-Ribes2023] Gabriel Coll -Ribes, Ivan J. Torres -Rodriguez, Antoni Grau, Edmundo Guerra, ...
work page Pith review arXiv 2019
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.