{"id":"482b7c08-f340-47e8-a32b-89f6ee780500","arxiv_id":"2411.12897","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"Machine learning on P-band tomographic SAR profiles classifies eight tree species with 72% overall accuracy, but macro F1 is just 0.39 and gains depend on XY coordinates, not radar intensity alone.","lead":"This study tests whether 3D radar images from the TomoSense dataset can identify eight tree species in a German national park. Using automated machine learning on radar height profiles, the model reached 72% overall accuracy, but only the dominant species was reliably distinguished.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"XY-coordinate confound is the load-bearing weak point: Table 1 shows accuracy jumps from ~0.56–0.60 to ~0.72–0.77 when X,Y are added, so the 72% headline cannot be attributed to TomoSAR intensities without an XY-only control.","rationale":"The reader's weakest assumption pinpoints the same load-bearing risk: the XY coordinates, not the tomographic profiles, may drive the 72% accuracy. I read the paper in good faith: it is transparent about class imbalance, reports macro F1 (0.39) and balanced accuracy alongside accuracy, and explicitly notes the accuracy gain when XY is added. Those are positive signs. However, transparency about the gain does not remove the confound. The paper's own Table 1 is the key evidence: without XY, accuracy is essentially at the majority-class baseline; with XY, it rises sharply. Since species are spatially clustered and the test split is geographic, X,Y can act as a leakage channel even when train and test regions are disjoint, because AutoGluon can interpolate or memorize local spatial patterns. The central claim asks what TomoSAR intensities contribute, so an XY-only control is not a nice-to-have but a necessary condition for attribution. I agree with the reader's CONDITIONAL verdict: the case study is useful and honestly reported, but the headline claim needs a leakage test before it can be accepted as evidence for TomoSAR-based species discrimination. No change to the verdict is needed because the reader already identified this condition and set the verdict accordingly.","tokens_in":5946,"tokens_out":3149,"duration_ms":37886,"concrete_test":"Retrain the exact AutoGluon pipeline under the P-square and P-swathe splits with three feature sets: (a) tomographic intensity features only, (b) X,Y coordinates only, and (c) intensities plus X,Y. Report balanced accuracy, macro F1, and per-class F1 for each. If the XY-only model (b) achieves balanced accuracy close to the 0.26/0.36 values in Table 1, then the 72% accuracy is largely a spatial-mapping artifact and the TomoSAR contribution is unsupported. If XY-only is near chance while the full model is high, the confound is resolved and intensities carry real signal.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that P-band tomographic SAR intensity profiles enable tree-species discrimination, with a 'reasonable overall accuracy of 72%' (Section 3). Table 1 directly undercuts the attribution of this number to radar data: with XY coordinates the accuracy is 0.72–0.77, while without them it drops to 0.56–0.60, near the 60% majority-class baseline. The XY features are raw pixel coordinates (Section 2.2), and tree species are strongly spatially clustered in the study area (Aspen alone is 60%, Appendix A.1). Because the model is an AutoGluon ensemble with access to these coordinates, it can learn spatial class maps or interpolate between nearby training and test pixels, especially under the swath split where test pixels may be adjacent to training pixels. The geographic splits reduce spatial autocorrelation but do not eliminate coordinate-based leakage: a flexible tabular model can use X,Y as a lookup key to the local species distribution. Therefore Table 2's 72% accuracy, and the statement 'some discrimination of species is possible with TomoSAR data', are only supported for the version of the model that has spatial coordinates. The no-XY results (balanced accuracy 0.20–0.25 versus 0.125 chance) hint at weak radar-only signal, but without an XY-only baseline the strength and direction of the TomoSAR contribution remain unquantified. This is a correctness-risk issue, not a stylistic one: the central claim is about the information content of tomographic intensities, yet the headline metric is confounded with geographic location.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies tabular machine learning (AutoGluon) to P-band tomographic SAR intensity profiles from the TomoSense dataset to classify eight tree species in a German forest. The authors voxelize the tomographic images, build per-pixel feature vectors from height-resolved intensities optionally augmented with X,Y coordinates, and evaluate two spatial train/test splits. They report an overall accuracy of 72% with a macro F1 of 0.39, note the strong effect of including spatial coordinates, and compare LiDAR-derived height statistics with model outputs. The stated central claim is that tomographic SAR data provide at least some species-discriminative information, with 'reasonable overall accuracy of 72%' in Section 3.","tokens_in":6395,"tokens_out":5559,"duration_ms":56891,"significance":"If the central claim is established, the paper would provide useful early evidence that tomographic P-band SAR height-intensity profiles can contribute to tree species mapping, which is relevant to upcoming missions such as ESA's Biomass. The authors deserve credit for using spatially held-out test sets, reporting class imbalance and balanced accuracy, and using AutoGluon in a reproducible AutoML pipeline. However, the load-bearing evidence for the TomoSAR-intensity claim is currently weak: the headline accuracy is driven by spatial coordinates rather than radar intensities, majority-class performance dominates the metric, and no baseline or error bars are provided. The study is a plausible pilot but the current support does not yet justify the central claim as stated.","major_comments":[{"comment":"The 72% headline in Section 3 is attributed to a model that includes raw pixel coordinates (X,Y) as features, and Table 1 shows that removing X,Y drops accuracy from 0.72-0.77 to 0.56-0.60. Because tree species are strongly spatially clustered (Aspen is 60% of the study area, Appendix A.1) and the swath split leaves test pixels adjacent to training pixels, the with-XY result may reflect spatial label leakage rather than the information content of the tomographic intensities. To make the central claim 'some discrimination of species is possible with TomoSAR data' load-bearing, the authors need an XY-only control (same model, intensities removed) and a distance-buffered spatial split, plus a full classification report for the no-XY configuration.","section":"Section 2.2, Table 1"},{"comment":"The no-XY balanced accuracies (0.20-0.25) are the only numbers that directly support the TomoSAR-intensity claim, and they are reported without confidence intervals or per-class metrics. With a 60% majority class, overall accuracy of 0.56-0.60 is close to a trivial baseline; the macro F1 (0.39) and balanced accuracy of the headline model show poor minority-class performance. The authors should report confidence intervals over repeated geographic splits and give the full classification report for the no-XY configuration, so the reader can assess whether the radar-only signal is statistically distinguishable from chance.","section":"Section 3, Table 2"},{"comment":"The LiDAR comparison is under-specified: it is not stated what RMSE measures (e.g., LiDAR height vs predicted height vs tomographic height) or how the violin plots in Fig. 5 are computed. The train/test RMSE values differ by a factor of about four (e.g., Aspen: 1.22 m vs 5.28 m), which is not interpretable without a definition. This undermines the conclusion in Section 4 that 'the model tends to overestimate the height of certain types of trees'.","section":"Appendix A.3, Table 4"},{"comment":"The square split is described as random square regions set to 5% of the data width, but the number of squares, the random seed, and the number of repeated splits are not specified. Without this information, the reported 0.77 accuracy for the square split cannot be assessed for stability, and the comparison between the swath and square splits remains anecdotal.","section":"Section 2.3"}],"minor_comments":[{"comment":"The text contains a typo: 'ar stakeholders' should be 'are stakeholders'.","section":"Section 1.1"},{"comment":"The AutoGluon configuration, hyperparameter search space, and the list of models that were actually selected are not reported; adding this information would improve reproducibility.","section":"Section 2.4"},{"comment":"The class indices in Table 2 are not mapped to species names in the table caption; please refer to Table 3 or add a class-name column.","section":"Table 2"},{"comment":"The abstract and the methodology promise comparisons across polarimetric configurations and heading directions, but the Results section reports only P-band results; either add the polarization comparison or remove it from the claims.","section":"Abstract and Section 2.2"},{"comment":"The true and predicted maps would be easier to evaluate if the color scale and class legend were included in the figure.","section":"Figure 3"}],"recommendation":"major_revision","confidential_remarks":"The paper's topic is suitable for the journal and the spatially split protocol is a strength, but the current version does not yet support the headline attribution of classification accuracy to tomographic intensities. The XY-coordinate confound is testable with straightforward control experiments, so I see this as a revision issue rather than a rejection issue. I would encourage the editor to request the additional experiments and a full no-XY classification report before further consideration."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a useful new case study: it is the first application of AutoGluon tabular models to the TomoSense P-band tomographic SAR dataset for tree species classification, with systematic geosplit and polarimetry comparisons. The authors are honest about the class imbalance, report class-wise metrics, and include a LiDAR height comparison. That is real work and worth having on record.\n\nThe soft spot is the XY coordinate confound, and it is the load-bearing one. Table 1 shows accuracy jumps from 0.56–0.60 without XY to 0.72–0.77 with XY. Since species are strongly spatially clustered (Aspen is 60% of the area), the model can use X and Y as a lookup key to local species distribution. The paper never runs an XY-only baseline, so we do not know how much TomoSAR intensity actually contributes. The no-XY balanced accuracy of 0.20–0.25 is above chance but weak, hinting at a modest radar signal. The central claim that TomoSAR enables discrimination is only supported for the model with coordinates. That should be fixed before publication. The paper also lacks error bars and simple baselines (majority class, LiDAR-only, or XY-only), and the LiDAR RMSE comparison in Table 4 is under-specified—no methodology for how those RMSEs were computed. Minor: code is not provided.\n\nWho this is for: remote sensing researchers working on TomoSense or SAR-based forest mapping, and people planning experiments with the ESA Biomass mission. It is a reasonable benchmark and the reporting is honest, but the conclusions need reframing around what the model with XY does versus what the radar data alone delivers.\n\nRecommendation: send to peer review—it deserves serious referee time—but ask the authors to add an XY-only control, simple baselines, and error bars, and to soften the claim about TomoSAR discrimination accordingly.","headline":"A useful new benchmark for TomoSense species classification, but the headline 72% accuracy is confounded by spatial coordinates and overstates what the radar data alone delivers.","tokens_in":6841,"tokens_out":1500,"would_cite":false,"duration_ms":16683,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that P-band tomographic SAR intensity profiles can classify eight tree species with 72% overall accuracy, but the result depends on including pixel coordinates and is dominated by the common aspen class.","keywords":["tree species classification","tomographic SAR","P-band radar","machine learning","tabular AutoML","forest inventory","spatial autocorrelation","imbalanced classification"],"falsifier":"Train the same pipeline on the same pixels with the tomographic intensity profile replaced by random noise or permuted across height layers; if the accuracy stays near 72%, the radar signal is not what drives the result. Equivalently, a coordinates-only model that matches the full-model accuracy would falsify the claim that TomoSAR data enables the classification.","tokens_in":5716,"feed_emoji":"🌲","tokens_out":6129,"duration_ms":60356,"temperature":0.7,"pith_summary":"This paper asks whether 3D tomographic radar data—a vertical stack of backscatter intensities reconstructed from many satellite images—can tell eight tree species apart in a German national park. The authors convert each ground pixel's tomographic intensity profile into tabular features and train several machine-learning models with automatic hyperparameter tuning. They report a \"reasonable\" 72% overall accuracy when the pixel's X/Y coordinates are also included, but only 56–60% without them, and a macro F1 of 0.39. In the authors' words, some discrimination of species is possible with TomoSAR data, but the strongest classes dominate and location contributes much of the signal.","feed_headline":"Tree species from 3D radar: 72% accuracy, but location matters","feed_subtitle":"Accuracy drops to 56–60% when pixel coordinates are removed, so radar profiles alone carry less signal.","key_machinery":"The central object is the tomographic height-intensity profile: each ground cell carries a 36-layer vertical profile of radar backscatter intensity reconstructed from a stack of single-look-complex SAR images. The paper flattens these profiles into tabular rows, optionally combines polarimetric channels (HH, HV, VV), and adds the pixel's X/Y coordinates. An automated machine-learning model selection tool then trains tabular classifiers on these rows, with the geographical split (swath or square) controlling how training and test pixels are separated.","core_discovery":"On its own terms, the paper establishes that P-band tomographic SAR height-intensity profiles are usable for tree species classification: an automated machine-learning pipeline classifies eight species with 72% overall accuracy in the best configuration, and 77% on a square geographical split. Balanced accuracy stays low (0.20–0.36), and the gap between macro F1 (0.39) and weighted F1 (0.70) shows that the majority Aspen class drives the score. The authors conclude that some species discrimination is possible with tomographic SAR data, while acknowledging the strong influence of class imbalance and the large accuracy gain from adding spatial coordinates.","pith_inferences":["The accuracy jump from 56–60% to 72–77% when X/Y coordinates are included suggests much of the reported performance may come from the model learning the spatial distribution of species rather than from radar backscatter; a coordinates-only baseline would settle how much TomoSAR truly contributes.","If a model trained on the same coordinates but with tomographic intensities shuffled or replaced by noise kept the same accuracy, the claim that TomoSAR data enables species discrimination would need to be substantially weakened.","The paper uses LiDAR heights only as a post-hoc validation; feeding height statistics derived from LiDAR or from the tomographic profile itself as features might show whether height, rather than intensity shape, is the informative part of the signal."],"forward_implications":["If tomographic P-band SAR can separate tree species, forest inventories could be updated from space, especially in cloudy regions where optical imaging fails.","The 72% accuracy is driven by the dominant Aspen class; rare classes such as Oak and Beech forests would need more training samples before operational use.","The large accuracy jump when pixel coordinates are added means future models should explicitly account for spatial context or use stricter spatial generalization tests.","The same feature pipeline could be applied to future spaceborne tomographic SAR missions to see whether the 2-meter-resolution result scales to lower resolution.","The LiDAR-derived height proxy used for validation suggests that comparing predicted and true canopy heights could become a useful quality check for tomographic classifications."],"supporting_citations":[{"why":"Provides the TomoSense dataset, including the 3D tomographic SAR imagery and the LiDAR point cloud used in all experiments.","marker":"Tebaldini et al. [2023]"},{"why":"Supplies AutoGluon, the automated machine-learning tool that selects, tunes, and ensembles the tabular classifiers.","marker":"Erickson et al. [2020]"},{"why":"Motivates the need for spatially separated train/test splits, leading to the swath and square geosplit configurations.","marker":"Salazar et al. [2022]"}],"fun_headline_variants":["Radar profiles classify 8 tree species, but location drives the score","Tomographic SAR achieves 72% tree species accuracy, but only with location hints","Tree species from 3D radar: accuracy hinges on where you sample","Class imbalance, not radar, limits tree species accuracy to 72%","Radar tree species: 72% accuracy, but balanced accuracy under 0.36"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the model learns from the tomographic radar intensities rather than from the pixel coordinates that are also given as inputs.","fun_headline_variants_meta":{"raw":{"variants":["Radar profiles classify 8 tree species, but location drives the score","Tomographic SAR achieves 72% tree species accuracy, but only with location hints","Tree species from 3D radar: accuracy hinges on where you sample","Class imbalance, not radar, limits tree species accuracy to 72%","Radar tree species: 72% accuracy, but balanced accuracy under 0.36"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000855,"raw_usage":{"total_tokens":3692,"prompt_tokens":898,"completion_tokens":2794,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":514,"completion_tokens_details":{"reasoning_tokens":2692}},"tokens_in":514,"tokens_out":2794,"duration_ms":19543,"temperature":1.0,"reasoning_tokens":2692,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:03:45.793944+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same pipeline on the same pixels with the tomographic intensity profile replaced by random noise or permuted across height layers; if the accuracy stays near 72%, the radar signal is not what drives the result. Equivalently, a coordinates-only model that matches the full-model accuracy would falsify the claim that TomoSAR data enables the classification.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the TomoSense dataset, including the 3D tomographic SAR imagery and the LiDAR point cloud used in all experiments."}],"review_version":1}