{"id":"ddd08ec8-d9b2-40d0-b163-e4966c013760","arxiv_id":"2506.20315","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"The paper reports 16 autonomous forest-inventory missions with an ANYmal quadruped in three countries, with mission autonomy above 80% of distance traveled and a 2 cm typical DBH accuracy figure inherited from the authors' prior work.","lead":"A team from Oxford and ETH Zurich built an autonomous quadruped robot system that maps forest plots and estimates tree sizes without a human walking the plot. The paper reports field trials across three European countries and distills lessons for making legged robots work in dense, unstructured forests.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 2 cm DBH accuracy claim is inherited from prior work [20] and is not demonstrated on the ANYmal D/Frontier campaigns in this paper; Section V-B only says the authors 'expect' it and mentions an unquantified caliper comparison at Stein am Rhein.","rationale":"The reader's weakest assumption identifies the same load-bearing gap: the DBH accuracy headline is not supported by new quantitative validation in this manuscript. I agree with that assessment. The navigation and autonomy results are genuinely useful and are supported by mission logs, intervention statistics, and repeated tree counts across campaigns; those parts of the paper justify a conditional acceptance. The area claim is also problematic (the abstract says '1 ha under 30 min' while Table 1 lists the largest operator-defined plot as 125 m x 30 m, and Section IV-E calls a 0.93 ha plot the largest mission), but the DBH accuracy issue is more central to the paper's stated purpose of forest inventory. The paper explicitly flags in Section V-B that ground-truth TLS was not available and that accuracy is expected from prior work, so this is not an artifact of the review pipeline; it is an admitted absence of evidence. A conditional acceptance with a request to report the caliper comparison or soften the abstract is the right call, which matches the reader's verdict.","tokens_in":25300,"tokens_out":2086,"duration_ms":25041,"concrete_test":"Quantify the Stein am Rhein caliper comparison: for each mission SaR-01 to SaR-04, compute per-tree DBH error against the manual caliper measurements, reporting N, mean absolute error, RMSE, and bias broken out by tree species; if the combined MAE is within 2 cm, the abstract claim stands for that site. Separately, re-run the same error analysis on the prior [20] dataset and check whether the ANYmal D/Frontier configuration is consistent with it; if caliper records are unavailable, the 2 cm claim should be removed from the abstract or explicitly labeled as expected from prior work.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's stated purpose is forest inventory, and the abstract's headline quality number is 'typical DBH accuracy of 2 cm.' Section V-B says ground-truth TLS was not available for many test sites, that accuracy was studied in related prior work [20], and that the authors therefore 'expect' the system to achieve 2 cm average DBH accuracy, with an unspecified 'additional confirmation' from manual caliper measurements at Stein am Rhein. No per-tree comparison, error distribution, sample size, or bias analysis is reported for any mission in this paper. This matters because the platform changed from the prior work to ANYmal D with the Frontier payload, the LiDAR changed (Hesai QT64), the locomotion controller changed between campaigns, and forest types and seasons differ. The DBH pipeline may be identical in principle, but sensor placement, motion profile, scanning geometry, and tree species all affect stem reconstruction, so the prior-work error bound does not automatically transfer. If the true DBH error is larger than 2 cm, the central claim about producing useful forest inventories is unsupported, even though the navigation and mapping achievements stand on their own.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript presents an integrated autonomy system for forest inventory using ANYmal C and ANYmal D quadruped robots. The system combines LiDAR-inertial odometry, pose-graph SLAM, dense mapping, local terrain mapping, hierarchical mission/local planning, and an online tree segmentation and trait-estimation pipeline. The evaluation covers 16 missions across five campaigns in Finland, the UK, and Switzerland between May 2023 and July 2024, and reports autonomy metrics (MDBI/MTBI), tree counts, scanned areas, a preliminary relocalization study, and a set of five lessons and challenges. The abstract claims that plots up to 1 ha can be surveyed in under 30 min with typical DBH accuracy of 2 cm; the conclusion repeats the 1 ha claim.","tokens_in":25541,"tokens_out":10739,"duration_ms":108717,"significance":"If the headline quantitative claims are substantiated, this would be a significant field-systems contribution: one of the most extensive demonstrations of autonomous under-canopy forest inventory with legged robots, across multiple countries, forest types, and seasons. The strengths of the paper are the detailed system description, the explicit intervention-based autonomy evaluation, the inclusion of failure cases (bog entrapment, dense undergrowth), and an honest discussion of open challenges such as tree height estimation and species identification. The field campaign data and the lessons learned give the paper value beyond the specific experiments. However, the two headline claims in the abstract and conclusion — 1 ha plots in under 30 min and 2 cm DBH accuracy — are not currently supported by the evidence reported in the manuscript, which is the main barrier to acceptance.","major_comments":[{"comment":"The canonical '1 ha in under 30 min' claim rests on an arithmetic inconsistency. Section IV-E states that Dea-01, the largest mission, was a '125 m × 30 m survey area, which corresponded to a 0.93 ha plot.' The product is 3,750 m² = 0.375 ha; 0.93 ha must instead be the effective scanning footprint computed with the 15 m effective range introduced in Section V-B. Table 2's 'Area covered' column therefore conflates operator-defined plot area with scanned coverage area, and the abstract's 'plots up to 1 ha' and the conclusion's 'forest inventories up to 1 ha' are not supported by the size of any surveyed plot. Please correct the arithmetic, label the coverage metric explicitly as an effective scanned area, and adjust the headline claims accordingly.","section":"IV-E, Table 2, and Abstract/Conclusion"},{"comment":"The second headline result, 'typical DBH accuracy of 2 cm,' is not demonstrated on the data collected in these campaigns. The text states that accuracy was studied in related prior work [20], that the authors 'expect' 2 cm average accuracy, and that this was 'additionally confirmed' by unquantified manual caliper measurements at Stein am Rhein. Since the platform (ANYmal D with Frontier payload), the LiDAR (Hesai QT64), the locomotion controller, and the forest types/seasons differ from those in [20], the prior error bound cannot be assumed to transfer without direct evidence. No per-tree comparison, sample size, error distribution, or bias analysis is reported for any mission in this paper, and the claimed consistency of repeated inventories is also left unquantified despite the range of tree counts in Table 2 (e.g., WyJ-01: 28 vs. WyJ-04/05: 46/52; SaR-01: 66 vs. SaR-04: 36). This is load-bearing for the paper's inventory-quality claim. Please add a direct validation (e.g., DBH residuals against calipers or TLS for at least one campaign, with detection recall/precision) or explicitly reframe the 2 cm figure as expected/prior-work accuracy.","section":"V-B, and Abstract/Conclusion"}],"minor_comments":[{"comment":"Using the tabulated MTBI values and N = interventions + 1, the reconstructed autonomous mission time is approximately 89.95% of the total mission time, slightly below the '90% of the mission time' stated in the text; please clarify whether the raw logs give a higher value or adjust the wording to 'nearly 90%.'","section":"V-A, Eq. (2)-(3), Table 2"},{"comment":"The effective-range-based 'Area covered' metric is not defined precisely; please state whether it is a union of 15 m disks, a corridor around the robot path, or a point-cloud footprint, since the headline area numbers depend on this definition.","section":"V-B"},{"comment":"The text's 'relocalization rate of 20%' does not match the tabulated per-sequence rates (21.9%, 13.3%, 35.5%; pooled rate ≈23%); please explain how 20% is derived, and rename the ambiguous 'Relocalizations' column heading.","section":"V-C, Table 3"},{"comment":"Several user-defined parameters influence the reported results — the weights in Eq. (1), the 2–4 m waypoint handover distance, the 90° minimum angular coverage for stem fitting, the 10 s intervention aggregation window, and the 15 m effective LiDAR range; a brief robustness or sensitivity note for these parameters would help readers understand how tuned the system is.","section":"III-D3, III-E4, V-B"},{"comment":"There are several typos: 'Stein am Rheim' should be 'Stein am Rhein,' 'Much for the forest' should be 'Much of the forest,' and 'modeling as as a series' should be 'modeling as a series.'","section":"IV-A, IV-G, III-E4"},{"comment":"The abstract says 'up to 1 ha plot under 30 min,' while Section VI-A says 'autonomous coverage up to 1 ha in 20 min'; these should be reconciled, and the phrasing '1 ha plot' should be corrected to 'plots up to 1 ha.'","section":"Abstract and Lesson 1"},{"comment":"The statement that interventions follow a Poisson distribution in time and distance is based on visual inspection of the histograms; if this is a formal statistical claim, a goodness-of-fit test and per-campaign sample sizes should be reported, otherwise the wording should be softened.","section":"V-A, Figures 9-10"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the journal's scope and the field deployment effort is substantial. My main concern is that the editorial framing, rather than the underlying engineering, produces unsupported quantitative claims: the ha arithmetic is internally inconsistent, and the 2 cm DBH figure is inherited from prior work without validation on the new platform. I believe these can be fixed in a single revision by correcting the area metric, adding at least one direct DBH validation or explicitly softening the headline, and reconciling the autonomy-time wording."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this one. First, it is a genuine field-robotics systems paper: sixteen missions, five campaigns, three countries, across seasons, with a working integrated stack on ANYmal C and D. The autonomy metrics (MDBI/MTBI), the consistent tree counts across repeat missions, and the five lessons are useful, credible content. The paper is honest about its own limitations, especially in Lesson 5, and it does not pretend the system is production-ready. That is real value.\n\nSecond, the headline claims do not survive contact with the paper's own numbers. The abstract says \"survey plots up to 1 ha under 30 min,\" but the largest operator-defined plot is 125 m x 30 m, which is 0.375 ha. The \"0.93 ha\" in Table 2 is the estimated scanned area using a 15 m effective LiDAR range, not the plot area. Those are different things. The authors should either change the headline to something like \"up to 0.93 ha scanned per mission\" or report scan coverage relative to plot area. This is an arithmetic/definitional fix, not a deep flaw, but as written the claim is misleading.\n\nThe DBH accuracy claim is softer. Section V-B says the 2 cm figure comes from prior work [20] and that the authors \"expect\" it to transfer, with an unquantified caliper check at Stein am Rhein. The platform changed (ANYmal D, Hesai QT64), the locomotion controller changed, and forest types differed, so the prior-work error bound does not automatically carry over. Since the paper's stated purpose is forest inventory, not just navigation, this is a load-bearing weakness. A per-tree comparison with caliper or TLS data for at least one 2024 campaign would fix it. Without that, the 2 cm number in the abstract should be flagged as inherited, not demonstrated.\n\nThe relocalization experiment is clearly preliminary and teleoperated; the authors mostly say so, though the word \"reliable\" in the discussion overstates a 13-35% relocalization rate with odometry gaps up to 28 m. Minor.\n\nWho is this for? Field robotics researchers, forestry robotics people, and anyone planning long-duration outdoor legged autonomy. The lessons section alone justifies a read. The paper deserves a serious referee; with the headline fixed and the DBH claim either validated or downgraded, it would be a solid contribution. I would take it into review.\n\nRecommendation: send it out, but have reviewers push on the area arithmetic and the DBH transfer argument. Those are the two spots that need revision before publication.","headline":"A solid field-deployment systems paper whose real contribution is the 16-mission dataset and honest lessons, but whose headline numbers (1 ha under 30 min, 2 cm DBH) overstate what is actually demonstrated.","tokens_in":26119,"tokens_out":956,"would_cite":true,"duration_ms":10831,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An autonomous legged robot can survey a 1-hectare forest plot in under 30 minutes and record tree diameters to about 2 cm.","keywords":["Autonomous Robots","Environmental Monitoring","Forestry","Legged Robots","Simultaneous Localization and Mapping (SLAM)","Forest inventory","Diameter at Breast Height (DBH)","Field robotics"],"falsifier":"Re-measure a sample of the trees in the Forest of Dean, Wytham Woods, and Stein am Rhein plots with calipers or terrestrial laser scanning, and compare per-tree diameter-at-breast-height values from the robot's inventory against those ground-truth values; the central claim holds only if the mean absolute error is at or below 2 cm, and the comparison would need to be published to be verifiable.","tokens_in":25063,"feed_emoji":"🌲","tokens_out":11823,"duration_ms":109630,"temperature":0.7,"pith_summary":"This paper argues that a four-legged robot can carry out a useful forest inventory on its own: given a plot boundary, the robot walks a lawn-mower pattern under the canopy, builds a 3D map as it goes, and outputs a spreadsheet of tree positions, trunk diameters, and heights. The evidence comes from 16 missions over 18 months in Finland, the UK, and Switzerland, using the ANYmal platform with the tree-analysis software running on board. The headline result is that plots up to 1 hectare can be covered in under 30 minutes, with trunk diameters reported at a typical accuracy of about 2 cm. The paper is explicit that this accuracy is expected from earlier validation of the tree-analysis approach rather than re-measured against ground truth in these particular missions. The stated purpose is to test whether legged platforms are mature enough to complement drones and handheld scanners for ground-level forest data, and to spell out where they still fall short.","feed_headline":"Legged robot inventories 1 ha of forest alone in under 30 minutes","feed_subtitle":"A quadruped autonomously walks the forest floor to log tree diameters and locations, complementing manual caliper crews.","key_machinery":"The load-bearing mechanism is the coupling between pose-graph SLAM and an online tree-analysis pipeline. A boustrophedon (lawn-mower) survey pattern with enforced spacing between path segments deliberately creates loop closures, keeping the map consistent; a LiDAR-inertial odometry system drives the pose graph, and dense local clouds (called data payloads) are accumulated about every 20 m of travel. The inventory pipeline removes ground with cloth-simulation filtering, segments tree stems by Voronoi clustering and cylinder fitting, then attaches each stem observation to the nearest SLAM graph node so that multi-view clouds are fused in one frame. Trait estimation fits oblique cone frustums to the fused stem points, requires at least 90 degrees of angular coverage before estimating, and derives diameter at breast height and height from the frustum stack. The same terrain representation feeds a reactive local planner that scores traversability and triggers mission re-planning when a waypoint is unreachable.","core_discovery":"The central claim is that a rugged legged robot, carrying a wide-field-of-view LiDAR and running the full autonomy stack on board, can be sent into an unmapped forest plot and return a usable forest inventory without a person walking the plot. The system couples LiDAR-inertial odometry with pose-graph SLAM and loop closures, builds dense data-payload clouds about every 20 metres of travel, and fuses those clouds through the SLAM graph across viewpoints so tree trunks are reconstructed from several sides before diameter at breast height, height, and position are estimated. The evidence is 16 autonomous missions in conifer, mixed, and deciduous forests across three countries, including a 0.93 ha oak plot covered in about 21 minutes with 97 trees detected; across missions the robot was autonomous for over 80% of the distance and 90% of the mission time. The authors frame the contribution as a feasibility demonstration and a set of five lessons about hardware, state estimation, navigation, forestry use, and assessment of such systems, and they report the 2 cm DBH figure as an expectation inherited from prior validation rather than as a freshly measured result.","pith_inferences":["Inference: the paper's own data support repeatability (similar tree counts across repeated runs in the same plot) rather than accuracy, so the 2 cm DBH claim should be re-tested on the actual platform and forest types before being quoted to foresters.","Inference: a head-to-head comparison with TLS, handheld MLS, and under-canopy drones on the same plots, measured in cost per hectare and soil impact as well as accuracy, would make the legged platform's niche explicit rather than assumed.","Inference: making the mission planner inventory-aware — re-planning to close gaps in angular coverage around detected stems instead of following a fixed lawn-mower pattern — is a natural next step that could improve DBH accuracy without lengthening missions.","Inference: the paper's cost observation implies adoption may hinge more on whether a legged robot can replace a TLS crew at comparable capital cost than on navigation performance alone."],"forward_implications":["Foresters could get ground-level inventories of up to 3 ha from a single battery charge, since a 0.93 ha plot was covered in 21 minutes and the platform supports roughly 90-minute missions.","Because the robot can relocalize against a prior map, the same plot can be re-visited to build longitudinal records of tree growth and change.","The autonomy metrics — over 80% of distance and 90% of mission time without safety-operator intervention — indicate that basic navigation is no longer the binding constraint; dense undergrowth and local-planning edge cases are.","The online inventory output (tree positions, DBH, height) is produced during the mission, letting the operator monitor coverage and tree traits live and intervene early if the survey is going wrong.","The system can use a prior map made by a human-carried scanner to localize itself, which is the key enabler for repeat monitoring missions after an initial survey."],"supporting_citations":[{"why":"Supplies the online tree-reconstruction and inventory pipeline, including the earlier validation from which the 2 cm DBH accuracy figure is inherited.","marker":"[20]"},{"why":"Earlier public report of the same project whose results and campaigns this paper extends with two additional deployments.","marker":"[17]"},{"why":"Provides the LiDAR-inertial odometry (VILENS) that drives the pose graph and localisation during missions.","marker":"[57]"},{"why":"Provides the pose-graph SLAM with loop-closure detection that keeps the global map consistent across the plot.","marker":"[60]"},{"why":"Provides the reinforcement-learning perceptive locomotion controller used to execute velocity commands on the ANYmal.","marker":"[68]"},{"why":"Supplies the Voronoi-based tree segmentation and cylinder-fitting method adapted for stem detection in the inventory pipeline.","marker":"[70]"},{"why":"Supplies the method used to derive diameter at breast height from the reconstructed stem curve.","marker":"[71]"},{"why":"Defines the scanning protocols and validation practices that motivate the lawn-mower survey pattern and ground-level measurement requirement.","marker":"[5]"},{"why":"Provides the LiDAR place-recognition and relocalization system evaluated in the prior-map experiment.","marker":"[74]"}],"fun_headline_variants":["Legged robot autonomously maps 1-hectare forest plot in under 30 min","Quadruped robot does forest inventory solo: 1 ha under 30 min","ANYmal robot autonomously surveys forest, logs tree data in 30 min","Under-canopy forest inventory by autonomous legged robot: 1 ha/30 min"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 2 cm tree-diameter accuracy is inherited from earlier validation work on a different setup and is described as expected rather than measured against ground truth in the forests surveyed here, so the headline measurement claim stands or falls on that transfer.","fun_headline_variants_meta":{"raw":{"variants":["Legged robot autonomously maps 1-hectare forest plot in under 30 min","Quadruped robot does forest inventory solo: 1 ha under 30 min","ANYmal robot autonomously surveys forest, logs tree data in 30 min","Under-canopy forest inventory by autonomous legged robot: 1 ha/30 min"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000619,"raw_usage":{"total_tokens":2917,"prompt_tokens":1036,"completion_tokens":1881,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":652,"completion_tokens_details":{"reasoning_tokens":1792}},"tokens_in":652,"tokens_out":1881,"duration_ms":16195,"temperature":1.0,"reasoning_tokens":1792,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:51:25.856624+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-measure a sample of the trees in the Forest of Dean, Wytham Woods, and Stein am Rhein plots with calipers or terrestrial laser scanning, and compare per-tree diameter-at-breast-height values from the robot's inventory against those ground-truth values; the central claim holds only if the mean absolute error is at or below 2 cm, and the comparison would need to be published to be verifiable.","supporting_citations":[{"cited_title":"Online Tree Reconstruction and Forest Inventory on a Mobile Robotic System,","cited_arxiv_id":null,"evidence_quote":"Supplies the online tree-reconstruction and inventory pipeline, including the earlier validation from which the 2 cm DBH accuracy figure is inherited."},{"cited_title":"Autonomous forest inventory with legged robots: System design and field deployment,","cited_arxiv_id":null,"evidence_quote":"Earlier public report of the same project whose results and campaigns this paper extends with two additional deployments."},{"cited_title":"Vilens: Visual, inertial, lidar, and leg odometry for all-terrain legged robots,","cited_arxiv_id":null,"evidence_quote":"Provides the LiDAR-inertial odometry (VILENS) that drives the pose graph and localisation during missions."},{"cited_title":"Online LiDAR- SLAM for Legged Robots with Robust Registration and Deep-Learned Loop Closure,","cited_arxiv_id":null,"evidence_quote":"Provides the pose-graph SLAM with loop-closure detection that keeps the global map consistent across the plot."},{"cited_title":"Learning robust perceptive locomotion for quadrupedal robots in the wild,","cited_arxiv_id":null,"evidence_quote":"Provides the reinforcement-learning perceptive locomotion controller used to execute velocity commands on the ANYmal."},{"cited_title":"Auto- matic dendrometry: Tree detection, tree height and diameter estimation using terrestrial laser scanning,","cited_arxiv_id":null,"evidence_quote":"Supplies the Voronoi-based tree segmentation and cylinder-fitting method adapted for stem detection in the inventory pipeline."},{"cited_title":"Accurate derivation of stem curve and volume using backpack mobile laser scanning,","cited_arxiv_id":null,"evidence_quote":"Supplies the method used to derive diameter at breast height from the reconstructed stem curve."},{"cited_title":"Evaluation and Deployment of LiDAR-based Place Recognition in Dense Forests","cited_arxiv_id":"2403.14326","evidence_quote":"Provides the LiDAR place-recognition and relocalization system evaluated in the prior-map experiment."}],"review_version":1}