{"id":"63e1d0f9-0011-40b1-be67-dbc386d8aac8","arxiv_id":"2608.01072","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Fine-tuning UniRig on L-system-generated plant meshes restores branching behavior and enables plant skeletal reconstruction, including foliage, without architectural changes.","lead":"PlantRig adapts UniRig, an autoregressive rigging model built for animated characters, to reconstruct plant skeletons from 3D meshes. Multi-round fine-tuning on procedurally generated trees restores branching topology and extends to leafy plants and some real scans.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Generalization claim rests on qualitative inspection; the reported '90 percent accuracy against the original loss' is undefined and no real-data ground-truth metrics or baselines are provided.","rationale":"The reader's verdict is CONDITIONAL, and the reader's rationale already identifies the absence of quantitative metrics and baseline comparisons as a core problem. My stress-test focuses on the precise logical link: the paper's strongest sentence in the abstract ('the resulting model generalized well across diverse plant forms') is supported only by qualitative figures plus an undefined '90 percent accuracy against the original loss' statement. The paper itself admits in Section 6 that the accuracy is based on visual inspection rather than standardized quantitative metrics, and no error bars or comparisons to existing plant skeletonization methods are given. This is a genuine gap in evidence, but it is the same concern the reader already identified, so the verdict should remain CONDITIONAL rather than moving to REJECT or UNVERDICTED. I do not see a separate internal inconsistency that would invalidate the method's plausibility; the weakness is in the strength of the evidence for the generalization claim, not in the logical coherence of the pipeline. A quantitative evaluation on the existing real-data meshes, with baselines and a clear definition of the reported 90% figure, would directly test whether the central claim is supported.","tokens_in":16088,"tokens_out":2050,"duration_ms":19789,"concrete_test":"Run a quantitative evaluation on the same GaussianPlant meshes used in Figures 17-18: obtain reference skeletons by expert manual labeling or by running an existing skeletonization baseline (e.g., Chaudhury-Godin or Smart-Tree), then compute branch-level F1, graph edit distance, and root-to-leaf path similarity for the fine-tuned PlantRig model, the zero-shot UniRig baseline, and the baseline method. Also report the exact loss value and definition behind the '90 percent accuracy against the original loss' claim. If the fine-tuned model does not outperform the baseline by a clear margin on these metrics, the generalization claim should be downgraded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that fine-tuning on synthetic L-system data closes the domain gap well enough for accurate plant skeletal reconstruction. The load-bearing step is the evaluation of that claim: the only real-plant results (Figures 17-18) are visual, with no ground-truth skeleton, no baseline comparison, and no quantitative score. The paper's headline number, 'about 90 percent accuracy against the original loss,' is undefined: loss is not accuracy, and no loss curve, normalization, or task is specified. Section 6 explicitly concedes that the reported reconstruction accuracy is based on visual inspection rather than standardized quantitative metrics. Because the entire generalization argument depends on this evaluation, the evidence is currently consistent with the model having memorized the synthetic archetypes and failing on real scans in ways not visible in two rendered examples. This is not an internal inconsistency, but it means the strongest claim in the abstract is not yet supported by the reported data.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether UniRig, an autoregressive rigging model originally trained on articulated characters, can be adapted to reconstruct plant skeletal structures. The authors first document three failure modes of the base model on synthetic L-system-generated plants: branch-token suppression leading to near-linear skeletons, arbitrary root placement, and limited sensitivity of the frozen mesh encoder to plant geometry. They then apply multi-round fine-tuning on procedurally generated datasets spanning eleven archetypes, with and without foliage and with simulated measurement noise, and report qualitative improvements in branch topology, root placement, and generalization to real GaussianPlant meshes. The central claim is that targeted fine-tuning substantially closes the domain gap between character-rigging priors and plant skeletal structure, without leaf-specific architectural changes.","tokens_in":16352,"tokens_out":3496,"duration_ms":31743,"significance":"If rigorously supported, the result would be a useful demonstration that autoregressive rigging models can transfer beyond articulated objects to hierarchical botanical structures, with practical implications for automated plant rigging, phenotyping, and digital twins. The paper also contributes a detailed diagnosis of branch-token suppression, a procedural plant generation toolkit, and an honest account of negative results from inference-time interventions. However, the load-bearing claim of generalization is supported almost entirely by qualitative figure inspection; the only quantitative assertion, 'about 90 percent accuracy against the original loss,' is undefined, and no standard skeleton metrics, error bars, or baseline comparisons are reported. The paper's own Section 6 concedes that reconstruction accuracy is based on visual inspection rather than standardized quantitative metrics.","major_comments":[{"comment":"The paper explicitly states that 'the reported reconstruction accuracy is currently based on visual inspection rather than standardized quantitative metrics,' yet the abstract and Section 5.1 claim 'about 90 percent accuracy against the original loss.' This quantity is never defined: loss is not accuracy, and no loss curve, normalization, or task is specified. Because the generalization claim is the central contribution, the paper must replace this undefined number with formally defined metrics (e.g., graph edit distance, branch correspondence accuracy, root-to-leaf path similarity) computed on held-out synthetic and real data, and report these separately for each archetype.","section":"Section 6, 'Discussion'"},{"comment":"The real-plant evaluation is limited to two qualitative examples (lavender and a twig) with no ground-truth skeleton, no comparison to existing plant skeletonization methods such as Smart-Tree or Chaudhury-Godin, and no quantitative scores. The claim that the model 'generalizes well to real plants' and 'generalizes well outside of its learned space' is load-bearing but is not supported by the reported evidence, which is consistent with the model having memorized synthetic archetypes and failing on real scans in ways not visible in two rendered examples. The authors should provide quantitative evaluation on a larger set of real scans, ideally with manually annotated or otherwise obtained ground-truth skeletons, and report error bars across multiple plants and species.","section":"Section 5.1, Figures 17 and 18; Section 5.2, Figures 24 and 25"},{"comment":"The training and test synthetic data are generated by the same procedural L-system framework, and the only out-of-distribution test set is a small number of GaussianPlant meshes. The real-to-sim gap is acknowledged but never quantified, and no distribution-shift statistics (e.g., mesh noise amplitude, surface-regularity measures, or reconstruction error distributions) are provided. The paper should quantify the gap between synthetic and real meshes, and validate transfer on a more diverse real dataset; otherwise the conclusion that the model 'did not merely memorize the synthetic archetypes' (Section 7) is not established.","section":"Sections 3.1 and 3.2, data methodology"},{"comment":"The paper reports that BranchBoostLogitsProcessor and root-forcing failed, but it does not specify the exact boost amounts, sampling parameters, or seeds used. Since these negative results are used to justify the turn to fine-tuning, the experimental configuration should be reported in enough detail (ideally in a table or appendix) to allow reproduction and to rule out the possibility that the failures were caused by arbitrary hyperparameter choices.","section":"Section 4.1, branching and root interventions"}],"minor_comments":[{"comment":"The phrase 'not be a sinecural task' contains a typo; it should be 'sinecure.'","section":"Section 3.1"},{"comment":"The word 'interpetability' should be 'interpretability,' and the phrase 'incredibly debauched skeleton' is informal and should be replaced with a more technical description.","section":"Section 5.1"},{"comment":"The phrase 'out most robust' should be 'our most robust.'","section":"Section 5.2"},{"comment":"The paper references UniRig's 'VocabSwitchingLogitsProcessor' but does not specify the exact model checkpoint, training dataset, or hyperparameters used for the base model; this information is needed for reproducibility.","section":"Section 4.1"},{"comment":"The affiliation list shows '2Computer Science, University of Osaka' twice; the entry for author Yang Yang should have a unique affiliation number. Reference [9] lists the last author as 'O. F.' instead of 'F. Okura,' and several references lack DOIs or arXiv IDs; please standardize the reference format.","section":"Author affiliations and references"},{"comment":"The captions do not describe what errors are visible in the real-plant reconstructions; adding annotations or close-up views would make the qualitative claims much easier to assess.","section":"Figures 17 and 18"}],"recommendation":"major_revision","confidential_remarks":"The methodological narrative is interesting and the qualitative results are suggestive, but the central generalization claim currently rests on undefined metrics and visual inspection. I recommend major revision rather than rejection because the required quantitative evaluation is within the scope of the manuscript: the authors have synthetic ground truth and access to real scans, so adding graph-edit-distance or branch-correspondence metrics with baselines is feasible. The main risk is that the reported '90 percent accuracy' may not survive formal measurement, but that risk should be tested rather than assumed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, the paper has a genuinely useful diagnostic finding: UniRig collapses plant branching because top-k sampling (k=5) systematically excludes the branch token (id 256) from the candidate pool, and the frozen 3DShape2VecSet encoder is too insensitive to plant geometry. That is concrete, checkable, and worth building on. Second, the paper's central claim—that multi-round fine-tuning on L-system data closes the domain gap to real plants—is plausible but not yet supported. The real-plant evidence is two rendered examples (Figures 17-18) and the sentence 'about 90 percent accuracy against the original loss' in Section 5.1. That number is undefined: loss is not accuracy, no loss curve, normalization, or task is specified. Section 6 concedes the evaluation is visual inspection.\n\nWhat the paper does well: it documents failures honestly. The BranchIsolation and BranchBoost experiments show why logit manipulation and root-forcing do not work. The multi-round fine-tuning recipe over 11 L-system archetypes, with a mesh autoencoder that mimics scan noise, is a reasonable domain-adaptation pipeline. If the authors release the generation and autoencoding scripts, that is a useful artifact on its own. The citation pattern looks fine: relevant prior work is cited, and the GaussianPlant data reuse is transparent.\n\nThe soft spots are proportional to the claim. The graph edit distance, branch correspondence, and root-to-leaf path similarity metrics the paper itself lists in Section 6 are exactly what is missing. No comparison to Smart-Tree, PlantPose, or geometric skeletonization methods. The synthetic train/test data come from the same procedural generator family; held-out seeds are good but do not break that correlation. Real-leafy generalization is only shown on synthetic autoencoded meshes, not scanned plants. And the writing is rough in places, with typos, but that is a minor issue.\n\nBottom line: this deserves a serious referee. The branch-token suppression diagnosis is original and useful, and the fine-tuning narrative is coherent. But the strong claims in the abstract cannot be accepted until quantitative metrics on real data are added and the '90%' claim is defined or removed. If the authors deliver that, this becomes a solid capability paper. As it stands, treat it as a promising, under-validated domain adaptation study.","headline":"A credible diagnostic study with a real finding about branch-token suppression, but the generalization claim rests on qualitative inspection and an undefined accuracy number; worth a serious referee but needs quantitative evaluation.","tokens_in":16764,"tokens_out":3309,"would_cite":false,"duration_ms":28696,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Retrained character-rigging AI rebuilds plant skeletons from 3D scans.","keywords":["autoregressive rigging","plant skeleton reconstruction","L-systems","fine-tuning","branch topology","Skeleton Tree Tokenization","domain adaptation","3D mesh reconstruction"],"falsifier":"Run the final checkpoint on a held-out set of real scanned plants spanning several species and measure branch correspondence with graph edit distance; if the roughly 90 percent visual accuracy reported here falls well below that level, or if unseen branching habits such as bamboo or candelabra collapse into chains, the central generalization claim is refuted.","tokens_in":15873,"feed_emoji":"🌿","tokens_out":3596,"duration_ms":28919,"temperature":0.7,"pith_summary":"This paper tries to establish that an autoregressive rigging model trained on articulated characters—UniRig—can be turned into a plant skeletal reconstruction system through multi-round fine-tuning on procedurally generated L-system meshes, with no leaf-specific architectural changes. The authors argue that the character-to-plant gap is mostly a training-data and encoder-adaptation problem, not an architectural incompatibility. If correct, automated rigging of branches and foliage becomes practical for phenotyping, agricultural digital twins, and biomechanical simulation, where manual rigging is currently a bottleneck. The paper documents both the base model's failure modes and the iterative fine-tuning recipe that overcomes them.","feed_headline":"Retrained character-rigging AI rebuilds plant skeletons","feed_subtitle":"UniRig, fine-tuned on synthetic L-system trees, recovers branching topology from 3D scans.","key_machinery":"The load-bearing mechanism is UniRig's Skeleton Tree Tokenization (STT): a stack-based depth-first traversal that emits each joint's coordinates once, inserts a dedicated <branch> token exactly when the traversal backtracks to a new parent, and visits children in canonical (z, y, x) order, with coordinates discretized into 256 bins per axis. A GPT-style OPT-125M decoder predicts this token sequence autoregressively, conditioned on a geometric prefix from a frozen 3DShape2Vecset encoder. STT turns plant hierarchy into a sequential prediction problem directly; the <branch> token is the pivot, and the paper's diagnostic processors show that when sampling excludes it, branching disappears even though the tokenizer can represent it losslessly.","core_discovery":"On the paper's own terms, the central discovery is that UniRig's collapse of branching plants into near-linear chains is caused by sampling-level suppression of the branch token—constrained top-k sampling can exclude token 256 from the candidate pool entirely—compounded by a frozen 3DShape2Vecset mesh encoder that is insensitive to plant structural variation. After two fine-tuning rounds, first 15,000 branch-only meshes across eleven archetypes and then 17,600 leafy meshes across the same archetypes, the model recovers accurate branching topology on synthetic and real scanned plants, achieving roughly 90 percent visual accuracy against the original loss on real data and generalizing to foliage despite the zero-thickness, mesh-normal-dependent geometry of leaves.","pith_inferences":["Going beyond the paper, the tokenization view suggests the same stack-based <branch>-token recipe could transfer to other recursive branching systems—river networks, vascular systems, or lightning—where a rooted tree is the ground truth.","The conditional token hierarchy the authors sketch for leaves (petiole implies blade, not conversely) could be formalized as a grammar constraint, which would make plant rigging extensible to flowers and fruit with the same sequential machinery.","A quantitative re-evaluation using graph edit distance or branch correspondence would likely be needed to confirm the reported 90 percent figure; the paper's own limitation statement says current accuracy is based on visual inspection."],"forward_implications":["Branch-only plant skeletons can be produced automatically from noisy 3D scans after fine-tuning, without handcrafted geometric optimization rules.","The same fine-tuned model generalizes to full plants with leaves, so complete plant rigs—branches plus foliage—are within reach without architectural changes.","The failure analysis implies that similar autoregressive rigging models should diagnose sampling-level token suppression and encoder sensitivity before redesigning architectures.","Future work can treat the encoder as trainable: jointly fine-tuning it with the decoder may further improve sensitivity to fine-grained structural variation."],"supporting_citations":[{"why":"Supplies the base autoregressive rigging model and Skeleton Tree Tokenization that the paper adapts and fine-tunes.","marker":"[6]"},{"why":"Serves as the comparative baseline that preserves topology better but over-segments branches and produces an unstable output space.","marker":"[7]"},{"why":"Provides real scanned plant meshes and point clouds used for evaluation and for modeling realistic measurement noise.","marker":"[9]"},{"why":"Establishes the L-system procedural generation foundation from which the synthetic training data with ground-truth skeletons are built.","marker":"[8]"},{"why":"Frames skeletal estimation as tree-constrained graph generation, the topology-aware alternative this work builds context against.","marker":"[5]"},{"why":"Represents the classical stochastic-optimization plant skeleton extraction approach that motivates learned alternatives.","marker":"[2]"}],"fun_headline_variants":["Fine-tuned rigging AI regrows plant branches","Character-rigging AI learns plant skeletons","Retrained AI reconstructs plant branching","Branch collapse fixed: AI rebuilds plant topology","Rigging model generalized to leaves and vines"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the manually filtered L-system synthetic dataset captures enough of real plant morphology that a model fine-tuned on it transfers to real scanned plants; the paper's only real-data evidence is qualitative inspection of a few GaussianPlant meshes.","fun_headline_variants_meta":{"raw":{"variants":["Fine-tuned rigging AI regrows plant branches","Character-rigging AI learns plant skeletons","Retrained AI reconstructs plant branching","Branch collapse fixed: AI rebuilds plant topology","Rigging model generalized to leaves and vines"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000189,"raw_usage":{"total_tokens":1351,"prompt_tokens":979,"completion_tokens":372,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":595,"completion_tokens_details":{"reasoning_tokens":303}},"tokens_in":595,"tokens_out":372,"duration_ms":3686,"temperature":1.0,"reasoning_tokens":303,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:12:22.270040+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the final checkpoint on a held-out set of real scanned plants spanning several species and measure branch correspondence with graph edit distance; if the roughly 90 percent visual accuracy reported here falls well below that level, or if unseen branching habits such as bamboo or candelabra collapse into chains, the central generalization claim is refuted.","supporting_citations":[{"cited_title":"Although UniRig was not designed with botanical data in mind, the experimental results demonstrate that it transfers remarkably well to plant meshes","cited_arxiv_id":null,"evidence_quote":"Supplies the base autoregressive rigging model and Skeleton Tree Tokenization that the paper adapts and fine-tunes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Serves as the comparative baseline that preserves topology better but over-segments branches and produces an unstable output space."},{"cited_title":"Skeletonization of plant point cloud data using stochastic optimization framework,","cited_arxiv_id":null,"evidence_quote":"Provides real scanned plant meshes and point clouds used for evaluation and for modeling realistic measurement noise."},{"cited_title":"3d functional-structural plant modelling for agricultural digital twins: A domain analysis,","cited_arxiv_id":null,"evidence_quote":"Establishes the L-system procedural generation foundation from which the synthetic training data with ground-truth skeletons are built."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Frames skeletal estimation as tree-constrained graph generation, the topology-aware alternative this work builds context against."},{"cited_title":"Plant Skeletal Extraction Recovering skeletal representations from three-dimensional plant data has traditionally been formulated as a geometric optimization problem","cited_arxiv_id":null,"evidence_quote":"Represents the classical stochastic-optimization plant skeleton extraction approach that motivates learned alternatives."}],"review_version":1}