{"id":"b9f0bfc1-b50c-4413-bc94-715e6fecbd77","arxiv_id":"2608.06842","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"One utility function over a learned representation recovers hidden choices across risk, time, loss, valuation, and social domains and preserves the correlation structure of the Econographics battery.","lead":"A single economic model, built on top of a frozen tabular AI model, can predict a person's choices in one domain (risk, time, losses, valuation, or social choice) from their choices in other domains. The model keeps most of the AI's predictive gain while also explaining how behavioural measures move together across people.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Leave-domain-out transfer is confounded by target-domain labels remaining in the frozen encoder's context table; the common-utility claim needs a label-free context test.","rationale":"Section 5.2's headline transfer result is the main evidence for the paper's claim that one estimated utility map unifies valuation across domains. For that claim to be true, the coefficient vector (alpha, beta) fitted without target-domain outcomes must itself transport predictive information. The design, however, leaves 100 labelled target-task rows inside D_t when computing h_it in Eq. (3), so the representation at the time the ridge index is applied is already target-domain-informed. The leave-domain-out test therefore cannot distinguish a genuinely common valuation map from a common summarizer of a representation that already contains exact-domain labels. The paper is transparent about this limitation, but the abstract and conclusion phrase the finding more strongly than the design supports. The missing experiment is feasible: reuse the zero-label context construction of Appendix C.1 and refit the index on the other seven domains. If the improvement collapses, the central claim should be rephrased; if it survives, the concern is resolved. The random-projection and pretraining-corpus issues are secondary: the projection is a fixed design choice with a simple sensitivity check, and the corpus is external to the present design. The reader's CONDITIONAL verdict already anticipates the label-in-context problem, so my read does not change the verdict; it sharpens the condition that should be met before acceptance.","tokens_in":28599,"tokens_out":5708,"duration_ms":63063,"concrete_test":"Re-run the leave-domain-out row of Table 4 with the context table D_t constructed under the Appendix C.1 zero-label protocol, so that no target-domain labels appear in the encoder's context, while keeping all splits, seeds, the projection, and the other-domain estimation sample fixed. If within-two accuracy falls from 63.16% toward the 40.42% median and normalized-error retention drops from 78% toward 0%, the transfer claim fails; if performance and the 66-entry ORIV rank alignment stay near 63%/78% and 0.935, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5.2 and Table 4 present the leave-domain-out utility index as evidence that 'the same valuation direction learned from the other seven domains therefore predicts the eighth.' But the representation h_it in Eq. (3) is computed with a context table D_t that still contains 100 labelled target-task outcomes, as the paper states in Section 3.2, Appendix B.1, and the Table 7 notes. Only the ridge coefficients in Eq. (10) are fitted without target-domain outcomes. Consequently, the test does not isolate cross-domain transfer of valuation: the frozen encoder's in-context inference from exact-target labels can supply the target-specific information, and the common index may be a calibrated projection of that label-driven state rather than the carrier of transfer. The paper discloses this — 'the excluded domain test isolates transfer of the valuation rule' and 'not zero shot learning of utility' — but the abstract and conclusion nevertheless assert that 'one estimated utility map unifies valuation across domains' and 'predicts domains excluded from utility estimation.' The 78% normalized-error retention and 0.935 rank alignment are therefore consistent with a weaker reading: a shared coefficient vector summarizing a representation already conditioned on target-domain examples. This is load-bearing because it determines where the predictive content actually lives.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a unified random-utility model built on a frozen tabular foundation model (TabFM). A frozen encoder maps each decision maker's profile, visible choices, and labelled context into a common learned state; a single linear index and one absolute-distance utility function over normalized switch plans are then applied to every task in a 26-task multiple-price-list battery from Chapman et al. (2023). The paper reports three main results: (i) the frozen model predicts a held-out economic domain better than a population median, and the gain disappears when visible choices are shuffled across decision makers; (ii) one common utility index, estimated without target-domain outcomes, retains most of the foundation model's error reduction and reproduces the ordering of the Chapman correlation matrix; and (iii) internal representation geometry separates economic domains in a way consistent with a unified predictive model. The design includes an outcome-blind SHA-256 domain assignment, a locked confirmatory protocol, participant cross-fits, bootstrap intervals, and destructive interventions. The paper is careful to distinguish predictive transfer from zero-shot utility learning, but the main transfer claim is weakened by the fact that exact-target labels remain in the context table used to compute the frozen representation.","tokens_in":28838,"tokens_out":4109,"duration_ms":46732,"significance":"If the central claim survives scrutiny, the paper is significant: it offers a concrete bridge between large pretrained tabular models and classical random utility, and it proposes an operational test of whether one valuation map can serve multiple economic domains. The confirmatory design is unusually strong for this literature: the protocol was locked before test outcomes were loaded, the domain assignment is outcome-blind, the comparisons include matched nonlinear learners, and the paper openly discloses the limits of the zero-label exercises. The replication culture is exemplary, including hashed checkpoints, participant-level bootstrap inference, and a Lean-verified formal statement of the architecture. The main weakness is interpretive: the leave-domain-out test does not isolate cross-domain valuation transfer because target-domain labels remain in the context table, and the single random projection used for the common utility index is not subjected to sensitivity analysis. These issues are load-bearing for the paper's headline claim, but they are fixable by reframing or additional analyses.","major_comments":[{"comment":"The leave-domain-out test does not isolate cross-domain valuation transfer: the context table D_t used to compute the representation h_it in Eq. (3) still contains 100 labelled target-task rows, and only the ridge coefficients in Eq. (10) are estimated without target-domain outcomes. The paper acknowledges this in the text and in Table 7 notes, but the abstract and conclusion assert that the model 'predicts domains excluded from utility estimation' and that 'one estimated utility map unifies valuation across domains.' The reported 78% retention of the normalized-error improvement and 0.935 rank alignment are therefore consistent with a weaker reading in which the common index is a calibrated projection of a representation already conditioned on exact-target examples. Please either add a genuinely label-free context condition to the main analysis, or revise the central claim to state that the valuation map transfers conditional on target-task context labels and treat the zero-label exercises as exploratory. The exploratory Table 9, where TabFM is statistically tied with Extra Trees in the domain holdout, should be discussed in the main text because it bears directly on this distinction.","section":"Section 5.2, Table 4, Eq. (3), Section 3.2"},{"comment":"The common utility index is estimated on a single random Gaussian projection Pi with k=128 and a fixed seed (20,260,806). No sensitivity analysis is reported for the projection seed or dimension. Because the index is linear in the projected state, the substantive claims about one utility map rest on the assumption that this particular projection preserves the utility-relevant geometry of the 2,048-dimensional representation. Please report results across several projections/dimensions or replace the random projection with a deterministic reduction (e.g., PCA) to demonstrate that the headline numbers are not an artifact of one draw.","section":"Section 3.1, Section 3.2, Eq. (4), Eq. (10)"},{"comment":"The correlation-reproduction exercise uses task-masked states, so other tasks in the same domain remain visible in the representation used to construct the utility forecasts. The paper notes this in Appendix D.3, but Section 5.3 presents the 0.939 rank alignment as evidence that 'one utility map reproduces joint behaviour' without carrying the caveat into the main statement. Since the same-domain choices could plausibly supply much of the dependence structure, please either add a whole-domain-masked extraction for the correlation exercise or qualify the Section 5.3 claim to indicate that the representation has access to same-domain choices.","section":"Section 5.3, Table 12, Appendix D.3"}],"minor_comments":[{"comment":"The abstract says 'The foundation model improves on the training-sample median'; this should specify that the held-out whole-domain comparison in Section 5.1 is against a training-sample median, whereas the complete-sample comparisons in Section 5.4 use a context median from the opposite participant fold.","section":"Abstract"},{"comment":"The caption refers to 'the same hierarchical task order' but does not define what that order is; please state the ordering rule or refer to the appendix where it is defined.","section":"Section 6, Figure 3 caption"},{"comment":"The statement that 'the two cross terms vanish by iterated expectations' should spell out the conditioning set explicitly; the decomposition is only valid if the expectations are conditioned on the appropriate information set (e.g., H_i and the task characteristics).","section":"Eq. (9), Section 3.1"},{"comment":"The complete-sample whole-domain mask shows that TabFM improves within-two accuracy in only six of eight domains, with time discounting and distributional choices favouring the median under that loss. The main text reports the aggregate advantage but not these negative cells; they should be mentioned in Section 5.6 rather than only in the appendix.","section":"Section 5.6 and Appendix C.4"},{"comment":"The note says 'validation and test outcomes are not loaded' for the 20,800 train-only predictions; this is clear, but the row label 'Leave one domain out utility index' could be clarified to indicate that only the valuation coefficients, not the representation, exclude target-domain outcomes.","section":"Table 4 note"}],"recommendation":"major_revision","confidential_remarks":"The paper is transparent about many of its limitations, and the replication package appears unusually thorough. The main risk is that the abstract and conclusion overstate the transfer result relative to what the leave-domain-out design can establish. The editor may want to ensure that the revised version either adds a label-free context analysis or reframes the central claim so that the reader is not left with the impression that the utility map alone carries cross-domain transfer. The exploratory zero-label results in Table 9 should be considered a key piece of evidence for this decision rather than a peripheral appendix."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is a careful, unusually self-aware empirical paper, and the core prediction result is real. The new combination is using a frozen tabular foundation model's representation as the domain for a single random-utility valuation rule, then testing transfer across eight behavioral domains in the Chapman et al. data. The confirmatory design is better than most of what crosses my desk: outcome-blind SHA-256 domain assignment, a locked protocol chosen before the test, participant cross-fits, bootstrap intervals, destructive shuffle interventions, and a public replication deposit. The fact that shuffling decision-maker identities kills the gain is strong evidence that the information lives in the coherent behavioral history. I believe the headline predictive result.\n\nThe common-utility compression also holds up as a compression. One linear index plus absolute-distance utility and logit shocks retains most of TabFM's error reduction and reproduces the ordering of the 66 Econographics correlations. That is a genuine result, and the external replication on the Self-Regulation Ontology is a nice bonus. The paper deserves a serious referee.\n\nSoft spots, in proportion. The leave-domain-out transfer test is weaker than the abstract makes it sound. The target domain's 100 labeled outcomes stay in the frozen encoder's context table; only the ridge coefficients are fitted without target-domain outcomes. So the test shows that one coefficient vector estimated on other domains works when the representation already contains target labels. It does not show that the utility map carries cross-domain information on its own. The paper discloses this in Section 3.2 and Appendix B.1, and its own zero-label experiment in Appendix C.1 confirms that without exact-target labels TabFM ties or loses to Extra Trees on normalized error. The abstract and conclusion overstate when they say the model 'predicts domains excluded from utility estimation' without flagging the in-context labels. That is the main fix: either run the leave-domain-out test with target labels removed from context, or rephrase the claim.\n\nSecond, the random projection to k=128 coordinates is fixed with one seed and no sensitivity analysis. Minor, but a one-line robustness check would settle it. Third, the 'structural' utility model is modal and uncalibrated; the paper says so, and that is fine for point prediction, but readers should not take the choice probabilities as quantitative.\n\nBottom line: the predictive and compression results are credible; the 'unity of economic behaviour' claim needs the labels-in-context caveat visible in the abstract. This paper is for people doing behavioral measurement, preference heterogeneity, and anyone using foundation models in economics. I would send it to review, with a request to fix the leave-domain-out design and the abstract. It will be a useful paper either way.","headline":"Careful, unusually honest empirical paper whose predictive result is credible, but the headline 'unity' claim overstates what the leave-domain-out test can show because target labels stay in the encoder's context.","tokens_in":29374,"tokens_out":3788,"would_cite":false,"duration_ms":40112,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The same estimated utility map predicts behaviour across risk, time, loss, and social domains.","keywords":["random utility","bounded rationality","discrete choice","tabular foundation models","joint prediction","multiple price lists","cross-domain transfer","behavioral heterogeneity"],"falsifier":"Run the leave-domain-out exercise a second time with every target-domain label also removed from the encoder's context table; if the common-index accuracy falls to the task-median level once those 100 in-context labels are gone, the claim that one estimated utility map transfers across domains would be refuted. A weaker but still informative check is to vary the fixed projection seed and compare the 85 percent retention across many seeds.","tokens_in":28308,"feed_emoji":"🎲","tokens_out":7040,"duration_ms":64846,"temperature":0.7,"pith_summary":"This paper tries to show that one estimated utility function can organize economic behaviour across risk, time, loss, valuation, ambiguity, and social-allocation decisions. Using a frozen tabular foundation model (TabFM) to encode each decision problem and the person's behavioural history, the author fits a single random-utility model with one common valuation index, one absolute-distance utility over normalized response plans, and one extreme-value shock law. The model retains 85 percent of the foundation model's prediction-error improvement over the population median (78 percent when the target domain's outcomes are excluded from the index fit) and reproduces the ordering of the 66 Econographics correlations with rank alignment 0.939. If true, it means a single utility map on a learned representation can serve all measured economic domains, and behavioural co-movement across domains is mostly a consequence of variation in that one map.","feed_headline":"One utility map retains 85% of cross-domain prediction power","feed_subtitle":"A single random-utility rule on a learned representation predicts hidden risk, time, loss, and social choices.","key_machinery":"The load-bearing object is the learned common choice domain $\\mathcal{M} \\equiv \\mathbb{R}^p \\times [0,1]$. The frozen encoder maps the decision problem and behavioural context to a state $h_{it} \\in \\mathbb{R}^{2048}$; each feasible response plan $r$ is normalized to a share $s_t(r) \\in [0,1]$; together they form $e_{it}(r) = (h_{it}, s_t(r))$. A single estimated index $g_{it} = \\alpha + \\beta^\\top \\Pi h_{it}$, with $\\Pi$ a fixed random projection to $k=128$ coordinates, is transformed by the logistic map to $\\mu_{it}$, an ideal normalized response. One systematic utility function $V_\\theta(e) = -|s_t(r) - \\mu_{it}|$ ranks every feasible plan in every domain, and one common type I extreme value shock law converts those utilities into stochastic choice via the conditional logit formula. The work this machinery does: it separates representation (learned by the frozen model) from valuation (estimated once and shared), so that cross-domain transfer can be tested by hiding a whole domain and asking whether the same $V$ and $\\theta$ predict it.","core_discovery":"The central discovery is that the same valuation direction learned from seven economic domains predicts the eighth. Concretely: a fixed random projection and one coefficient vector map each of TabFM's 2,048-dimensional contextual states into a scalar index $\\mu_{it}$; each feasible switch plan is valued by its absolute distance from that index, $V(e) = -|s - \\mu|$; and independent type I extreme value shocks turn those valuations into conditional-logit choice probabilities. The same index, utility, and shock law are applied to all 26 multiple-price-list tasks. With one index and one utility rule, the model places 65.83 percent of choices within two rows versus 40.42 percent for the task population median; when every outcome from the target domain is excluded while fitting the index, accuracy remains 63.16 percent. The same estimated map also reproduces the ordering of the 66 correlations among 12 Econographics constructs (rank alignment 0.939, and 0.935 in the leave-domain-out fit). The paper's claim is therefore that random utility over a learned common representation is not merely a formal unification but an economically meaningful compression that transfers across domains.","pith_inferences":["One testable extension the paper leaves implicit: the same common-index compression could be applied to demographic and cognitive features alone, and the drop in correlation-geometry alignment would quantify how much of the 'unity' is behavioural as opposed to demographic.","If projection seed and dimension sensitivity were checked, the ridge-plus-projection step could become a minimal structural summary of the foundation model's economic content; high sensitivity would mean the common index is not yet a stable economic object.","The covariance decomposition shows that residual and cross covariances are not negligible, so a fully orthogonal common-utility explanation would require a richer shock structure; that is a natural next model comparison.","Because the zero-label results show TabFM and a locally fitted trees model essentially tied, the economic value of the foundation representation may lie in portability across tasks rather than accuracy within a single task; a test of this would compare the common index built on TabFM states versus one built on a classical learned representation."],"forward_implications":["If the central claim is right, a single estimated utility map—not 26 domain-specific decoders—captures most of the predictable structure in this 26-task behavioural battery.","The leave-domain-out result implies that the valuation rule transfers to domains it never saw in training, which is the minimal form of portability needed for out-of-sample welfare or menu evaluation.","The 0.939 rank alignment with the observed Econographics correlation matrix implies that the coordination of preferences across risk, time, loss, and social domains is substantially mediated by one systematic utility component, with the random shock component playing a smaller role.","The prediction ordering generalizes to other behavioural batteries: the same pattern appears in the 37-task Self-Regulation Ontology replication.","A universal structural model would require replacing the estimated index with an in-context valuation learner that updates $\\theta$ from context without refitting, which the paper identifies as the remaining gap."],"supporting_citations":[{"why":"Supplies the Econographics experimental battery of 26 multiple price list screens and the 12 constructed behavioural measures whose correlation structure is the target of the dependence reproduction.","marker":"Chapman et al. (2023)"},{"why":"Provides the public replication deposit from which raw switch responses, profiles, and survey weights are taken.","marker":"Chapman et al. (2022)"},{"why":"Introduces TabFM, the frozen tabular foundation model whose contextual states are the learned representation used throughout.","marker":"Kong and Das (2026)"},{"why":"Pins the released TabFM checkpoint and inference code that are frozen and audited in the empirical analysis.","marker":"Google Research (2026a)"},{"why":"Supplies the conditional logit random-utility form that converts common utility values into stochastic choice probabilities.","marker":"McFadden (1974)"},{"why":"Gives the obviously related instrumental variables measurement-error correction used to construct the 66-entry correlation matrix that the common-utility model must reproduce.","marker":"Gillen et al. (2019)"},{"why":"Defines centred kernel alignment (CKA), the measure of task representation geometry that shows domain structure in the frozen encoder.","marker":"Kornblith et al. (2019)"},{"why":"Provides the ordered random-utility decomposition that motivates placing empirical content in the map from learned states to utilities.","marker":"Apesteguia and Ballester (2025)"}],"fun_headline_variants":["One utility rule predicts hidden domain choices","Single utility map transfers across all 7 domains","Unified utility model beats median on hidden data","Learned utility function explains cross-domain behavior","Same utility rule predicts eighth domain choices"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The cross-domain transfer claim depends on the assumption that the 100 labelled target-task rows still present in the frozen encoder's context table do not by themselves let the model infer the hidden domain's local response mapping; if they do, the common utility index is a convenient summarizer rather than the carrier of transfer.","fun_headline_variants_meta":{"raw":{"variants":["One utility rule predicts hidden domain choices","Single utility map transfers across all 7 domains","Unified utility model beats median on hidden data","Learned utility function explains cross-domain behavior","Same utility rule predicts eighth domain choices"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000158,"raw_usage":{"total_tokens":1219,"prompt_tokens":936,"completion_tokens":283,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":552,"completion_tokens_details":{"reasoning_tokens":217}},"tokens_in":552,"tokens_out":283,"duration_ms":3377,"temperature":1.0,"reasoning_tokens":217,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T19:48:05.344733+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the leave-domain-out exercise a second time with every target-domain label also removed from the encoder's context table; if the common-index accuracy falls to the task-median level once those 100 in-context labels are gone, the claim that one estimated utility map transfers across domains would be refuted. A weaker but still informative check is to vary the fixed projection seed and compare the 85 percent retention across many seeds.","supporting_citations":[{"cited_title":"2022 , version =","cited_arxiv_id":null,"evidence_quote":"Provides the public replication deposit from which raw switch responses, profiles, and survey weights are taken."},{"cited_title":"2026 , howpublished =","cited_arxiv_id":null,"evidence_quote":"Introduces TabFM, the frozen tabular foundation model whose contextual states are the learned representation used throughout."},{"cited_title":"Proceedings of the 36th International Conference on Machine Learning , series =","cited_arxiv_id":null,"evidence_quote":"Defines centred kernel alignment (CKA), the measure of task representation geometry that shows domain structure in the frozen encoder."},{"cited_title":", title =","cited_arxiv_id":null,"evidence_quote":"Provides the ordered random-utility decomposition that motivates placing empirical content in the map from learned states to utilities."}],"review_version":1}