{"id":"dbae2b9f-838d-47ea-a309-b9c81282a968","arxiv_id":"2608.10050","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"From 85,078 company-month accounting logs, CAR-PL and a T-Learner tie on the two growth KPIs while a conservative value model leads on Quick Ratio, and CAR-PL spreads its recommendations most broadly.","lead":"A team at Intuit trained machine-learning policies that recommend one of 34 business-change actions, such as increasing ad spending or switching vendors, to small businesses from historical accounting logs. On firms the models never saw, the best policy depends on the target metric, and the top growth-KPI methods are not statistically distinguishable.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The shared scorer's one-hot outcome-model probes (Eqs. 9-10) are never validated against observed multi-hot data, so the KPI-level leader rankings may be artifacts of off-support extrapolation.","rationale":"The paper is careful and internally consistent: company-disjoint evaluation, matched contrasts, firm-clustered inference, a frozen shared scorer, bounded corrections, and multiple robustness checks all support the reported point estimates as properties of the scoring rule. The reader's weakest-assumption diagnosis is correct and is the most load-bearing concern: Eq. (10) is dominated by an outcome-model contrast evaluated at one-hot and zero treatment probes, configurations that may not occur in the multi-hot logs. The robustness replacement of the zero reference with the modal co-action vector does not address the one-hot side, and the outcome-model diagnostics validate only factual multi-hot inputs. Since the residual term is bounded and outcome-model-only scoring preserves the leaders, the rankings hinge on how the neural outcome model extrapolates to these off-support probes. This concern is acknowledged in Section 7.6 as a limitation, but it is not resolved by the reported experiments. It does not invalidate the internal logic; it limits what the findings can support, which is exactly the reader's CONDITIONAL verdict. I agree with the reader and see no reason to change the verdict. A secondary concern about co-action confounding in CAR-PL's focal R-learner exists, but the paper's central claim is explicitly about the shared model-assisted score, and the one-hot extrapolation issue is the more direct threat to that claim.","tokens_in":12449,"tokens_out":9290,"duration_ms":99450,"concrete_test":"Per category a, count training rows whose treatment vector equals u_a and rows with Ta=1 and at most two co-active categories. For categories with fewer than 100 such rows, recompute the Eq. (9) contrast using a model fit only on factual rows with total co-active count at most 2, or an explicitly additive outcome model, and compare the resulting 34 category contrasts with the headline contrasts. If the Spearman rank correlation falls materially below the reported zero/modal correlations (0.884-0.971) or any KPI-level point-estimate leader changes, the rankings are not robust to the one-hot extrapolation and the central claim should be weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The comparison rests on the outcome-model reference contrast in Eq. (9), evaluated at the one-hot vector u_a and the zero vector. The training treatment distribution is multi-hot, and after KPI-specific filters it is plausible that for several categories few or no training rows have exactly one active category; the neural outcome model is therefore extrapolating at the exact probes that supply the first term of Eq. (10). The paper's robustness check replaces the zero reference with the modal co-action vector but leaves the one-hot probe unexamined, and the outcome-model diagnostics in Table 6 are computed on factual multi-hot test rows, not on one-hot configurations. Because the residual-correction term is bounded and largely inactive (Table 7), and because outcome-model-only scoring reproduces the same leaders or top pair (Table 9), the KPI-level rankings are effectively determined by these unvalidated off-support contrasts. This makes the central claim—objective-specific ranking from multi-hot logs—dependent on an untested extrapolation rather than on observed policy behavior. This is not an internal logical error; it is a missing calibration check that the paper itself flags as prospective validation in Section 7.6.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper formulates the problem of ranking a single business-change category from multi-hot observational accounting logs as observational policy ranking. For each of three KPIs (Gross Profit, Revenue, Quick Ratio), it trains a separate policy: CAR-PL, an action-wise R-learner with support shrinkage; a T-Learner; a conservative contextual value model; plus zero-shot LLM, modal, and random references. All policies are evaluated on company-disjoint held-out firms with a shared frozen model-assisted score (Eq. 10). The paper reports KPI-specific point-estimate leaders, matched company-clustered confidence intervals, robustness to score components and treatment references, and selection-concentration analyses.","tokens_in":12607,"tokens_out":7125,"duration_ms":73108,"significance":"If the scoring instrument is accepted, the paper provides a credible industrial benchmark with careful evaluation design: company-disjoint splits, a frozen policy-agnostic scorer, firm-clustered bootstrapping, matched comparisons, and multiple robustness checks. CAR-PL's support-shrunk residual learning and broad catalog use are a useful methodological data point, and the paper is honest about the distinction between a model-assisted comparison score and an estimate of deployed policy value. However, because the ranking is mediated by a model-assisted score evaluated at standardized treatment probes, the significance of the central claim depends on validating those probes against observed multi-hot data.","major_comments":[{"comment":"The leading term of the shared score in Eq. (10) is the outcome-model contrast m_hat(w, u_a) - m_hat(w, 0) from Eq. (9), evaluated at the one-hot vector u_a and the zero vector. The training treatment distribution is multi-hot, and the paper does not report how often, if ever, each category appears as the sole active category; for categories with few or no exact-one-hot rows, the neural outcome model is extrapolating at precisely the probe that determines the score's first term. This is load-bearing because Table 7 shows the residual correction is saturated on only 2.6-12.9% of rows, and Table 9 shows outcome-model-only scoring retains the same leaders or top pair, so the rankings are effectively determined by these unvalidated contrasts. The existing modal-co-action robustness (Table 9) addresses only the reference side, not the one-hot side. Please add (i) exact-one-hot support counts per category, (ii) outcome-model calibration or rank correlation on exact-one-hot test rows if any exist, and (iii) a sensitivity analysis that replaces the one-hot probe with an observed co-action anchor, such as the average model prediction over training bundles containing category a. Without this, the headline KPI-level ranking claim rests on an untested extrapolation.","section":"Section 5.2, Eqs. (9)-(10)"},{"comment":"The target in Eq. (2) is defined as a focal-action contrast that averages over the observed co-action environment, but the score's first term compares a one-hot treatment vector to the zero vector. These are different estimands unless one assumes either that bundled co-actions are handled by the outcome model in a way that makes the one-hot contrast the relevant ranking instrument, or that the extrapolation to one-hot configurations is accurate. The paper should state precisely which estimand Eq. (10) is intended to score and why the one-hot-versus-zero contrast is the right ranking criterion for a policy that emits a single interpretable category from multi-hot logs. This clarification is needed for the central claim that the reported scores rank policies by objective-specific guidance quality.","section":"Section 3.4 and Section 5.2"}],"minor_comments":[{"comment":"The abstract reports 85,078 observations from 7,505 firms, while Table 1 lists KPI-specific unique firm counts between 7,072 and 7,191; the text should clarify that 7,505 is the pre-filter corpus size and that KPI-specific filtering removes some firms.","section":"Abstract and Table 1"},{"comment":"The outcome-model diagnostics are computed on factual multi-hot test rows; please label them explicitly as factual diagnostics in the table and text, and distinguish them from any new diagnostics at the standardized probe configurations.","section":"Table 6"},{"comment":"The clustered bootstrap resamples firms but does not address potential cross-firm calendar-time correlation from overlapping outcome windows; a sentence noting this limitation, or a two-way clustering check, would strengthen the inference section.","section":"Section 5.3"},{"comment":"The zero/modal Spearman correlation is reported without specifying whether it is computed on training, validation, or test rows; please state the row population used for this robustness statistic.","section":"Table 9"},{"comment":"The limitations paragraph mentions standardized treatment references but should explicitly flag the absence of an exact-one-hot support check as a limitation, rather than only listing the generic dependence on nuisance models.","section":"Section 7.6"}],"recommendation":"major_revision","confidential_remarks":"The paper is an industry study with proprietary data, and the evaluation design is unusually careful. The main obstacle is the unvalidated one-hot probe in the shared scorer; if the authors supply the requested support counts and calibration sensitivity analysis, I would be willing to accept a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: this paper earns its industrial claims. It sets up a clean prospective timeline, company-disjoint splits, firm-clustered bootstrap inference, a frozen policy-agnostic scorer, and a robustness grid. The result is a credible comparison of four policy families on multi-hot accounting logs, and CAR-PL's support-shrunk R-learning is genuinely new as a way to stabilize one-of-34 argmax selection. The headline pattern—CAR-PL highest on Gross Profit, T-Learner highest on Revenue, contextual value model highest on Quick Ratio, with no statistical separation between CAR-PL and T-Learner on either growth KPI—is consistent with the reported intervals and robustness checks. That is more than many applied papers deliver.\n\nThe real soft spot is the scorer's off-support probes. Eq. (9) evaluates the outcome model at one-hot and zero treatment vectors, but the training distribution is multi-hot. The paper never validates the model's extrapolation at those exact points. The residual correction is bounded and mostly inactive, and the outcome-model-only and modal-reference robustness checks do not touch the one-hot probe. The KPI rankings are therefore largely determined by that unvalidated contrast. To the authors' credit, they disclose this and name prospective validation as the next step. It is a missing calibration check, not a hidden error, but it means the claim that objective-specific ranking is feasible is conditional on the scoring model's extrapolation being trustworthy.\n\nOther soft spots are proportionate and mostly external: proprietary data, withheld detector thresholds, and no code make independent replication impossible. The observational setting carries unmeasured-confounder risk, though the company-disjoint held-out evaluation at least tests generalization within the observed environment. There are several tuned constants—support multiplier, clipping bounds, CQL alpha—but the paper reports a sensitivity grid for Gross Profit and shows the outcome-model-only leaders are stable, which mitigates the concern.\n\nWho this is for: people building decision support from operational logs, and researchers working on multi-action off-policy evaluation or policy learning with uneven support. It deserves a serious referee. The right outcome is likely a major revision that either validates the off-support scoring or softens the feasibility claim to match what the evidence supports.","headline":"Serious industrial policy-learning paper with a disciplined evaluation scaffold; the main caveat is an unvalidated off-support scoring probe, not an internal error.","tokens_in":13271,"tokens_out":1631,"would_cite":true,"duration_ms":18020,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"CAR-PL matches top policy for SMB growth guidance while spreading advice across the catalog","keywords":["observational policy ranking","multi-action logs","R-learner","support shrinkage","financial guidance","SMB","model-assisted evaluation"],"falsifier":"Retrain the frozen outcome reference model with a different architecture (e.g., a transformer instead of the shared-trunk MLP) or with different pre-decision features, and re-score all policies on the same test rows; if the KPI-level point-estimate leaders change materially, the reported rankings are artifacts of the scoring model. Alternatively, run a prospective randomized study that assigns the top-scoring categories to firms and checks whether realized KPI changes match the model-assisted score ordering.","tokens_in":12150,"feed_emoji":"📊","tokens_out":1466,"duration_ms":15084,"temperature":0.7,"pith_summary":"This paper asks whether useful financial guidance for small businesses can be learned from historical accounting logs, where business changes are self-selected and often happen together, instead of from randomized experiments. It introduces CAR-PL, a support-shrunk action-wise R-learner that learns one effect head per category while regularizing rare-category noise, and compares it with several baselines under a shared frozen scorer on company-disjoint held-out firms. The main claim is that CAR-PL ties with the T-Learner for the highest point estimates on Gross Profit and Revenue while selecting 33–34 categories, and that the contextual value model leads on Quick Ratio. The authors argue this supports objective-specific ranking of financial guidance from multi-action logs, and that recommendation concentration should be considered alongside aggregate scores.","feed_headline":"CAR-PL ties top policy while diversifying SMB advice","feed_subtitle":"A support-shrunk R-learner matches the best growth-KPI score while selecting 33–34 categories, making coverage a key metric.","key_machinery":"The central object is the Covariate-Adjusted Residual Policy Learning (CAR-PL) method: after cross-fitting a reward nuisance and per-category propensity models, it forms residuals for outcome and treatment, fits one weighted gradient-boosted effect head per category, and applies a support multiplier $\\hat\\delta^{\\mathrm{shr}}_a(x) = \\hat\\delta_a(x)\\, n_a/(n_a+\\lambda_{\\mathrm{sup}})$ before the argmax over 34 categories. This support shrinkage stabilizes selection of rare categories. The comparison itself rests on a shared model-assisted focal-action score $\\hat s_i(a)$ that combines an outcome-model reference contrast with a bounded propensity-weighted residual correction, evaluated at the one-hot and zero treatment vectors.","core_discovery":"The paper's central claim is that a support-shrunk action-wise R-learner (CAR-PL) matches the strongest baseline in aggregate model-assisted score on both growth KPIs while spreading recommendations across the catalog. On Gross Profit, CAR-PL has the highest point estimate (0.084); on Revenue, the T-Learner has the highest (0.085), with CAR-PL at 0.079; the paired differences are not statistically separated on either KPI. The contextual value model has the highest Quick Ratio point estimate (0.062). The paper also finds that CAR-PL and the T-Learner agree on only 12.8–14.2% of company-month states on the growth KPIs, showing that comparable aggregate scores can arise from markedly different state-to-category mappings. The robustness analyses (outcome-model-only scoring and alternative treatment references) retain the same KPI-level point-estimate leader or top pair.","pith_inferences":["If CAR-PL's broader catalog use is not merely a byproduct of the shrinkage multiplier but reflects genuinely different learned mappings, then downstream analyses of which states get which categories could reveal interpretable policy structure across business domains, a testable extension the paper does not pursue.","The stability of category rankings when replacing the all-zero treatment reference with the modal training co-action vector suggests that the shared scorer's outcome model is not strongly distorted by the off-support probes, but this robustness result is limited to the reported KPI-level leaders; per-state scoring instabilities could still exist.","A natural next step the paper leaves implicit is prospective validation: deploying the top-scoring policies in a randomized or A/B setting to measure whether the model-assisted scores predict actual KPI improvements from the recommended category.","The method's reliance on the shared scorer means that policy rankings could change if the outcome model were retrained with a different architecture or feature set; a sensitivity analysis across several outcome models would tell how much of the leaderboard is an artifact of the measurement instrument."],"forward_implications":["If CAR-PL's support shrinkage is what enables broad coverage without losing aggregate score, then recommendation diversity becomes a reportable dimension in offline policy evaluation, not just the mean policy value.","The company-disjoint evaluation design, with firm-level clustered bootstraps for matched comparisons, provides a template for comparing policy families in settings where randomized recommendation labels are unavailable.","CAR-PL's parallel training of independent effect heads suggests the method scales to much larger category catalogs than 34, making it a practical option for industrial guidance systems.","The result that the zero-shot LLM concentrates on one category while outcome-fitted policies remain competitive supports the paper's claim that generic business priors are not a substitute for state-specific policy learning.","Objective-specific ranking matters: the KPI-dependent leaders mean that a single policy is not uniformly best, so deployment decisions should be made per financial objective."],"supporting_citations":[{"why":"Supplies the observational policy-learning foundation for binary treatment that CAR-PL extends to multiple actions.","marker":"[2]"},{"why":"Provides the multi-action doubly robust policy-learning method that grounds the comparison and the focal-action target.","marker":"[3]"},{"why":"Provides the R-learning residualization construction that CAR-PL's action-wise heads build on.","marker":"[6]"},{"why":"The conservative Q-learning penalty that the contextual value baseline adapts.","marker":"[10]"},{"why":"Supplies the propensity-score identification conditions assumed for the focal conditional effect.","marker":"[16]"},{"why":"The augmented inverse-propensity weighting construction on which the shared scoring rule is based.","marker":"[20]"},{"why":"Provides the off-policy evaluation perspective that motivates the bounded residual correction in the scorer.","marker":"[22]"}],"fun_headline_variants":["CAR-PL matches top growth scores with broader SMB advice","Diverse SMB guidance ties top growth-KPI performance","Support-shrunk policy learner ties leader, diversifies advice","CAR-PL: broader SMB coverage, tied top growth scores","Wider SMB advice mix matches top growth-KPI policy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison depends on the shared scorer evaluating categories through one-hot and zero treatment probes that never occur in the real multi-hot logs; if the outcome model extrapolates poorly to those off-support configurations, the KPI leaderboard could be an artifact of the measurement instrument rather than a property of the policies.","fun_headline_variants_meta":{"raw":{"variants":["CAR-PL matches top growth scores with broader SMB advice","Diverse SMB guidance ties top growth-KPI performance","Support-shrunk policy learner ties leader, diversifies advice","CAR-PL: broader SMB coverage, tied top growth scores","Wider SMB advice mix matches top growth-KPI policy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000227,"raw_usage":{"total_tokens":1512,"prompt_tokens":1029,"completion_tokens":483,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":645,"completion_tokens_details":{"reasoning_tokens":397}},"tokens_in":645,"tokens_out":483,"duration_ms":4900,"temperature":1.0,"reasoning_tokens":397,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:14:38.498848+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain the frozen outcome reference model with a different architecture (e.g., a transformer instead of the shared-trunk MLP) or with different pre-decision features, and re-score all policies on the same test rows; if the KPI-level point-estimate leaders change materially, the reported rankings are artifacts of the scoring model. Alternatively, run a prospective randomized study that assigns the top-scoring categories to firms and checks whether realized KPI changes match the model-assisted score ordering.","supporting_citations":[{"cited_title":"Policy learning with observational data.Econometrica, 89(1):133–161,","cited_arxiv_id":null,"evidence_quote":"Supplies the observational policy-learning foundation for binary treatment that CAR-PL extends to multiple actions."},{"cited_title":"Conservative Q-learning for offline re- inforcement learning","cited_arxiv_id":null,"evidence_quote":"The conservative Q-learning penalty that the contextual value baseline adapts."},{"cited_title":"Rosenbaum and Donald B","cited_arxiv_id":null,"evidence_quote":"Supplies the propensity-score identification conditions assumed for the focal conditional effect."},{"cited_title":"Thomas and Emma Brunskill","cited_arxiv_id":null,"evidence_quote":"Provides the off-policy evaluation perspective that motivates the bounded residual correction in the scorer."}],"review_version":1}