{"id":"ef71430b-9285-484b-929c-f83504122b50","arxiv_id":"2605.26675","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"CART random forests are analyzed as controlled stochastic processes, separating subsampling and split policy effects, with explicit MSE derivations for linear models showing local stabilization but potential global suboptimality.","lead":"The paper introduces a stochastic-control theory for CART random forests by modeling them as sequential allocation over random opportunity sets. This could help researchers better understand and optimize ensemble methods using operations research tools.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Foundational mapping from CART split rule to masked-action allocation policy must hold exactly for the controlled process and its terminal law to match real forests.","rationale":"The reader's weakest assumption is precisely the load-bearing step identified above; the abstract-only review already flags it, and the full-text claims of stabilization and explicit expansion rest on the same equivalence. No other internal inconsistency is visible from the given material.","tokens_in":1744,"tokens_out":353,"duration_ms":31499,"concrete_test":"For a fixed subsampling rate m and linear regression data, compute the exact one-step transition matrix on the informative split-count state under the true CART rule (exhaustive search over the masked features) and compare it entrywise to the transition matrix implied by the paper's masked-action policy; a total variation distance >0.05 on any state would show the representation is inexact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claims (local stabilization via contraction of informative split-count imbalances, concentration of terminal geometry, and explicit MSE risk expansion under the linear model) are derived from modeling the feature-subsampled CART process as a controlled stochastic process whose transitions are induced by the masked-action policy over random opportunity sets. If the CART argmax split selection within each random mask does not correspond precisely to the allocation rule used in the control formulation (for example, because node-specific feature availability or the deterministic nature of the split criterion introduces selection probabilities not captured by the mask), then the induced Markov chain on split-count states and its terminal distribution will differ from actual CART trees. In that case the contraction property and the subsequent risk expansion would apply only to the abstract control model, not to the algorithm analyzed.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper develops a stochastic-control framework for feature-subsampled CART random forests (termed CART-ROSA), interpreting random feature subsets as random feasible action sets and the CART split rule as a masked-action allocation policy. This induces a controlled stochastic process over informative split-count states whose terminal law governs single-tree error and cross-tree interactions in forest MSE. The central claims are that the CART policy is locally stabilizing (contracting imbalances in informative split allocations and concentrating terminal tree geometry) and, under the linear model, yields an explicit MSE risk expansion. The approach separates the informative-opportunity rate from within-mask contraction strength.","tokens_in":1879,"tokens_out":409,"duration_ms":35159,"significance":"If the modeling holds, the work supplies a novel operations-research lens on ensemble methods that renders tractable a theoretical gap in understanding CART forests. The local stabilization result and explicit linear-model risk expansion constitute concrete, falsifiable contributions that could inform both analysis and design of random forests by quantifying how subsampling and split policy interact at the system level.","major_comments":[{"comment":"The foundational mapping (abstract and modeling sections) from the deterministic CART argmax split selection within each random feature mask to the masked-action allocation policy must be shown to induce precisely the same Markov chain on split-count states as real CART trees. Node-specific feature availability and the deterministic nature of the split criterion may introduce selection probabilities not captured by the mask alone; if the correspondence is inexact, the contraction property and subsequent MSE expansion apply only to the abstract control model rather than the algorithm analyzed.","section":"Modeling of the allocation policy (abstract and § on CART-ROSA formulation)"}],"minor_comments":[{"comment":"Clarify whether the invented term 'CART-ROSA' is intended as a new nomenclature or merely descriptive; ensure consistent usage across the manuscript.","section":null}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the careful and constructive review. The major comment raises an important point about the exactness of the modeling correspondence, which we address below by committing to a formal verification in revision.","responses":[{"response":"We agree that a rigorous demonstration of equivalence is required. The CART-ROSA formulation defines the random feature mask at each node exactly as in standard feature-subsampled CART, with the allocation policy selecting the single feature in the mask that maximizes the CART splitting criterion evaluated at the current node. Because masks are drawn independently at every node and the state tracks only the cumulative split counts (which determine the relevant history for subsequent opportunities under the linear-model assumptions), the induced transitions on the split-count process match those of the actual algorithm. The deterministic argmax is applied conditionally on the realized mask and current node, but the state definition ensures the overall process remains Markov with the same kernel. To eliminate any ambiguity, the revised manuscript will add a formal proposition in the modeling section proving that the transition kernel on split-count states is identical to that generated by real CART trees, including explicit verification that node-specific criterion evaluations do not introduce additional dependence beyond the mask and state.","revision_made":"yes","referee_comment":"[Modeling of the allocation policy (abstract and § on CART-ROSA formulation)] The foundational mapping (abstract and modeling sections) from the deterministic CART argmax split selection within each random feature mask to the masked-action allocation policy must be shown to induce precisely the same Markov chain on split-count states as real CART trees. Node-specific feature availability and the deterministic nature of the split criterion may introduce selection probabilities not captured by the mask alone; if the correspondence is inexact, the contraction property and subsequent MSE expansion apply only to the abstract control model rather than the algorithm analyzed."}],"tokens_in":1385,"tokens_out":388,"duration_ms":36337,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The punchline is that this work models CART forests as a controlled process where random feature subsets are opportunity sets and the split rule is a masked allocation policy. From that they derive that the policy locally contracts imbalances in informative splits and concentrates terminal geometry, plus an explicit MSE risk expansion when the response is linear.\n\nWhat is new is the separation of the informative-opportunity rate (from subsampling) and the contraction strength (from the within-mask rule), plus the claim that the resulting Markov chain on split-count states is locally stabilizing. The linear-model risk expansion is also presented as a direct consequence. That framing is not in the usual algorithmic or bias-variance literature on random forests.\n\nThe paper does a clean job laying out the two levers and showing how the control view makes the terminal law tractable. The derivations appear to be carried through without obvious circularity.\n\nThe soft spot is the exact correspondence between the CART argmax inside each random mask and the allocation rule used in the controlled process. If node-specific feature availability or the deterministic nature of the split criterion produces selection behavior not captured by the mask, the contraction property and the risk expansion apply only to the abstract model, not to actual CART trees. The stress-test note flags this correctly; the abstract states the claims for CART forests, so the derivations need to confirm the mapping holds without extra assumptions.\n\nThis is for readers who work on theoretical analysis of tree ensembles or who want new levers for ensemble design. It is not for someone looking for immediate algorithmic improvements or non-linear settings. The work shows clear thinking and honest engagement with the modeling gap, so it deserves a serious referee even if the policy equivalence requires tightening.","headline":"The paper gives a stochastic-control model of feature-subsampled CART that yields an explicit MSE expansion under linear models, but the central claims rest on whether the masked-action policy exactly reproduces real split selection.","tokens_in":2358,"tokens_out":425,"would_cite":false,"duration_ms":16328,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"CART forests can be modeled as stochastic control over random opportunity sets, under which the split policy contracts local imbalances in split allocations and yields an explicit MSE expansion for linear models.","keywords":["CART random forests","stochastic control","feature subsampling","ensemble MSE","split allocation","opportunity sets","risk expansion","masked policy"],"falsifier":"A direct simulation of the split-count process under repeated CART splits on random feature subsets showing no contraction in allocation imbalances would falsify the local stabilization claim.","tokens_in":2633,"feed_emoji":"🌲","tokens_out":631,"duration_ms":24748,"temperature":0.7,"pith_summary":"The paper develops a stochastic-control representation of feature-subsampled CART random forests by treating random feature subsets as random feasible action sets and the CART split rule as a masked-action allocation policy. This induces a controlled stochastic process on informative split-count states whose terminal law governs both single-tree error and cross-tree interaction terms in the forest MSE. The representation separates the informative-opportunity rate from feature subsampling and the contraction strength from the split policy. It establishes that the CART policy contracts imbalances in informative split allocations and concentrates terminal tree geometry, though it may be globally suboptimal for the forest objective. Specializing to the linear model produces an explicit MSE risk expansion.","feed_headline":"CART forests contract split imbalances via opportunity-set control","feed_subtitle":"Stochastic-control model of masked allocation separates subsampling rate from split policy and derives explicit linear-model MSE","key_machinery":"The masked-action allocation policy induced by the CART split rule over random feasible feature sets, which drives a controlled stochastic process on split-count states.","core_discovery":"By recasting feature-subsampled CART as sequential allocation over random opportunity sets, the terminal law of the split-count process determines both single-tree and interaction terms in ensemble MSE; the CART policy contracts imbalances in informative splits and concentrates tree geometry, and the linear-model case admits an explicit risk expansion.","pith_inferences":["Alternative split policies could be designed to achieve better global performance while retaining the same subsampling mechanism.","The split-count process representation may predict ensemble behavior from low-dimensional simulations before full training runs.","The local-versus-global distinction could clarify why certain ensemble variants outperform others despite similar individual trees."],"forward_implications":["The informative-opportunity rate induced by subsampling and the contraction strength from the split policy can be tuned as separate levers.","Local stabilization implies that terminal tree geometry concentrates across realizations.","Global suboptimality of the CART policy for the forest objective suggests room for alternative allocation policies.","The explicit MSE expansion for linear models permits precise comparison of risk under different subsampling rates."],"fun_headline_variants":["Stochastic control recasts CART as opportunity-set allocation","Split-count process determines single-tree and interaction MSE","CART policy concentrates tree geometry in forest risk model","Explicit MSE expansion derived for linear-model CART forests","Opportunity rate and split policy separate in CART-ROSA"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Interpreting random feature subsets as random feasible action sets and the CART split rule as a masked-action allocation policy accurately captures the mechanics of tree construction.","fun_headline_variants_meta":{"raw":{"variants":["Stochastic control recasts CART as opportunity-set allocation","Split-count process determines single-tree and interaction MSE","CART policy concentrates tree geometry in forest risk model","Explicit MSE expansion derived for linear-model CART forests","Opportunity rate and split policy separate in CART-ROSA"]},"model":"grok-4.3","cost_usd":0.004469,"raw_usage":{"total_tokens":2228,"prompt_tokens":666,"num_sources_used":0,"completion_tokens":73,"cost_in_usd_ticks":44687000,"prompt_tokens_details":{"text_tokens":666,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1489,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":666,"tokens_out":73,"duration_ms":18185,"temperature":1.0,"reasoning_tokens":1489,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T16:15:39.568671+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A direct simulation of the split-count process under repeated CART splits on random feature subsets showing no contraction in allocation imbalances would falsify the local stabilization claim.","supporting_citations":[],"review_version":1}