{"id":"68443dfc-d042-4955-ab90-09f41d6d1a1b","arxiv_id":"2602.16990","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Conv-FinRe is a new benchmark built from real market data and human trajectories that tests LLMs on generating utility-grounded stock rankings over fixed horizons while distinguishing rational analysis from behavioral mimicry or momentum.","lead":"The paper introduces Conv-FinRe, a conversational and longitudinal benchmark that evaluates LLMs on stock recommendations using utility grounded in investor risk preferences rather than just matching observed user behavior. A smart generalist might read it to see how AI evaluation in finance can shift from imitating noisy short-term choices toward supporting better long-term decision quality.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Normative references may embed selection bias from human trajectories used to infer risk preferences","rationale":"The reader's weakest_assumption directly identifies the same construction-bias risk. Full-text access would let us inspect the exact inference procedure for risk preferences, but the load-bearing issue remains whether that procedure is independent of the descriptive trajectories. This moves the verdict from UNVERDICTED to CONDITIONAL pending the proposed check.","tokens_in":1728,"tokens_out":311,"duration_ms":61388,"concrete_test":"Re-derive normative rankings for a 20% random subsample of investors using only the onboarding interview responses and an independent mean-variance utility function with elicited risk aversion parameters, excluding all subsequent human choice data; compare against the original paper rankings—if Kendall-tau distance exceeds 0.25 on average, the claimed separation is sensitive to construction choices.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim requires that multi-view references cleanly separate normative utility (derived from investor-specific risk preferences) from descriptive behavior and market momentum. If risk preferences are inferred or calibrated using the same human decision trajectories that define observed behavior, the separation becomes vulnerable to circularity or selection artifacts; trajectories collected under specific market conditions or from a non-representative cohort would systematically tilt the normative baseline. The abstract states the benchmark is built from real market data and human decision trajectories, but any dependence between preference grounding and observed actions undermines the diagnostic power for LLM rational analysis vs. noise mimicry.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces Conv-FinRe, a conversational and longitudinal benchmark for stock recommendation that moves beyond behavior imitation. It constructs the dataset from real market data and human decision trajectories, instantiates advisory dialogues, and supplies multi-view references: one capturing descriptive user choices and another grounding normative utility in investor-specific risk preferences. Evaluation of state-of-the-art LLMs reveals a persistent tension in which utility-aligned models often diverge from observed user actions while behaviorally aligned models overfit short-term noise. The dataset and codebase are released publicly.","tokens_in":1849,"tokens_out":459,"duration_ms":63301,"significance":"If the multi-view references prove robust, the benchmark would provide a valuable tool for diagnosing whether LLMs perform rational analysis, mimic user noise, or follow market momentum in financial settings. This addresses a clear gap in recommendation evaluation by prioritizing long-term utility over pure behavioral matching. Explicit credit is due for the public release of the dataset on Hugging Face and the codebase on GitHub, which supports reproducibility.","major_comments":[{"comment":"Benchmark construction (likely §3): the paper must specify the exact procedure for inferring investor-specific risk preferences from human decision trajectories and demonstrate that this normative reference is constructed independently of the descriptive behavior trajectories. If the same trajectories are used for both, the claimed separation between normative utility and observed behavior risks circularity or selection bias, directly undermining the diagnostic power for rational analysis versus noise mimicry.","section":"§3"},{"comment":"Evaluation and results (likely §5): the reported tension between utility-based ranking and behavioral alignment lacks details on the precise metrics employed, any statistical significance testing, and controls for market volatility or cohort selection. Without these, it is unclear whether the observed performance gap is robust or an artifact of the reference construction.","section":"§5"}],"minor_comments":[{"comment":"The abstract refers to 'step-wise market context' and 'advisory dialogues'; including one or two concrete example conversations in the main text would help readers understand the longitudinal and conversational structure.","section":"§4"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive and detailed feedback. The comments identify important areas where additional clarity and rigor will strengthen the manuscript. We address each major comment below and indicate the planned revisions.","responses":[{"response":"We agree that explicit specification of the inference procedure and a clear demonstration of independence are necessary to substantiate the multi-view reference design. The current manuscript describes the high-level construction from real market data and human trajectories but does not provide the algorithmic details. In the revision we will add a dedicated subsection to §3 that (i) states the exact procedure: risk aversion parameters are estimated via maximum-likelihood fitting of a CRRA utility model to responses from a dedicated risk-elicitation questionnaire administered at onboarding, using only those survey answers; (ii) shows that the normative reference is computed solely from these elicited parameters and the subsequent market context, while the descriptive reference uses the actual longitudinal choice sequences; and (iii) includes a short validation that the two references are not mechanically identical (e.g., correlation between inferred risk aversion and raw choices is moderate and consistent with rational behavior rather than tautological). This separation uses disjoint data sources and will be illustrated with pseudocode.","revision_made":"yes","referee_comment":"[§3] Benchmark construction (likely §3): the paper must specify the exact procedure for inferring investor-specific risk preferences from human decision trajectories and demonstrate that this normative reference is constructed independently of the descriptive behavior trajectories. If the same trajectories are used for both, the claimed separation between normative utility and observed behavior risks circularity or selection bias, directly undermining the diagnostic power for rational analysis versus noise mimicry."},{"response":"We acknowledge that the evaluation section would benefit from greater methodological transparency and robustness checks. We will revise §5 to (i) define the metrics explicitly (Kendall tau and NDCG for utility alignment; top-k accuracy and behavioral correlation for descriptive matching); (ii) report statistical significance via paired t-tests and Wilcoxon signed-rank tests on the per-model differences, with p-values and effect sizes; and (iii) add controlled analyses that stratify results by market-volatility regimes (high vs. low VIX periods) and by participant cohorts (novice vs. experienced investors). These new tables and figures will be included in the revised manuscript to demonstrate that the reported tension persists across these controls.","revision_made":"yes","referee_comment":"[§5] Evaluation and results (likely §5): the reported tension between utility-based ranking and behavioral alignment lacks details on the precise metrics employed, any statistical significance testing, and controls for market volatility or cohort selection. Without these, it is unclear whether the observed performance gap is robust or an artifact of the reference construction."}],"tokens_in":1402,"tokens_out":583,"duration_ms":44019,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main point is that this paper releases Conv-FinRe, a conversational longitudinal benchmark built from real market data and human trajectories that gives separate references for what users actually chose versus what aligns with their stated risk preferences over a fixed horizon. That distinction lets you test whether a model does rational analysis, copies short-term noise, or just follows market momentum. Prior financial rec work mostly stops at behavior matching, so the multi-view setup is the concrete addition here. They instantiate advisory dialogues from onboarding interviews and step-wise context, then run a set of current LLMs and report a clear tension: utility-strong models often diverge from user choices, while behaviorally aligned ones overfit volatility. Releasing the dataset on Hugging Face and the code on GitHub makes it immediately usable for follow-up work.","headline":"Conv-FinRe adds a benchmark that separates utility-grounded references from observed behavior in financial recommendations, but the abstract leaves the construction and validation steps thin.","tokens_in":2366,"tokens_out":235,"would_cite":false,"duration_ms":42300,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"Financial LLM benchmark using inverse-optimized mean-variance utility; no structural overlap with RS cost or forcing chain","alignment":"orthogonal","rationale":"The paper's core machinery (multi-view references, inverse optimization of λ/γ risk parameters in U(s) = μ̃ − λσ̃² − γDrawdown, uNDCG/MRR evaluation separating y_util from y_user/y_mom) operates entirely in applied behavioral finance and LLM diagnostics. It neither invokes nor parallels any RS element: no J-cost functional equation, no φ-ladder, no 8-tick periodicity, no parameter-free derivation of constants, and no recognition-cost reasoning. The domain (conversational stock recommendation benchmarks) lies outside the scope of the RS forcing theorems (reality_from_one_distinction, AbsoluteFloorClosure, Cost.FunctionalEquation, DimensionForcing, etc.). No contradiction arises because the paper makes no foundational physical or logical claims.","tokens_in":50623,"confidence":"high","tokens_out":211,"duration_ms":21803,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Conv-FinRe supplies multi-view references that separate what investors actually choose from what aligns with their own long-term risk preferences.","keywords":["financial recommendation","LLM benchmark","conversational AI","utility grounded evaluation","risk preferences","longitudinal analysis","behavioral vs normative","stock recommendation"],"falsifier":"A direct comparison in which one set of model outputs is scored only against the normative utility references and another only against the descriptive behavior references, then checked for whether the two rankings produce measurably different portfolio outcomes over the stated investment horizon.","tokens_in":2632,"feed_emoji":"📈","tokens_out":660,"duration_ms":30649,"temperature":0.7,"pith_summary":"Most financial recommendation benchmarks judge models only by how closely they copy observed user actions. Conv-FinRe instead supplies step-wise market context, onboarding interviews, and advisory dialogues over a fixed investment horizon, then supplies separate reference rankings that reflect descriptive behavior and normative utility derived from each investor's risk preferences. This setup lets evaluators determine whether an LLM follows rational analysis, reproduces user noise, or tracks market momentum. Experiments with current LLMs reveal a consistent split: models that rank well by utility often diverge from user choices, while models that match user choices tend to overfit short-term volatility. The benchmark is built directly from real market data and recorded human decision trajectories.","feed_headline":"Benchmark separates rational stock picks from user noise","feed_subtitle":"Conv-FinRe supplies separate references for observed choices and risk-preference utility so evaluators can tell which an LLM is following.","key_machinery":"Multi-view references constructed from real market data and human decision trajectories that separately score descriptive user choices and normative utility based on investor risk preferences.","core_discovery":"Conv-FinRe is a conversational and longitudinal benchmark that evaluates LLMs on stock recommendation by providing multi-view references distinguishing descriptive behavior from normative utility grounded in investor-specific risk preferences, enabling diagnosis of whether models follow rational analysis, mimic user noise, or are driven by market momentum.","pith_inferences":["The same multi-view reference approach could be applied to recommendation domains outside finance where short-term user actions conflict with stated long-term goals.","Benchmark results could guide the design of hybrid systems that first elicit risk preferences and then generate recommendations conditioned on those preferences rather than raw history.","Public release of the dataset allows repeated testing of whether newer models close the observed gap between utility alignment and behavioral imitation."],"forward_implications":["Models that achieve high utility-based rankings can be identified even when they diverge from recorded user selections.","Behaviorally aligned models can be flagged when they reproduce short-term noise rather than stable preferences.","Advisory systems can be tuned toward rational decision quality without requiring perfect imitation of every user action.","Longitudinal evaluation becomes possible because the benchmark supplies context across multiple market steps and a fixed horizon."],"fun_headline_variants":["Conv-FinRe tests rational stock utility against user noise","Benchmark diagnoses LLM adherence to risk preferences or noise","Conv-FinRe distinguishes descriptive behavior from normative utility","Longitudinal financial benchmark for LLM utility evaluation"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The references built from market data and human trajectories cleanly isolate normative utility from observed behavior without major construction artifacts or selection biases.","fun_headline_variants_meta":{"raw":{"variants":["Conv-FinRe tests rational stock utility against user noise","Benchmark diagnoses LLM adherence to risk preferences or noise","Conv-FinRe distinguishes descriptive behavior from normative utility","Longitudinal financial benchmark for LLM utility evaluation"]},"model":"grok-4.3","cost_usd":0.009561,"raw_usage":{"total_tokens":4257,"prompt_tokens":650,"num_sources_used":0,"completion_tokens":60,"cost_in_usd_ticks":95612000,"prompt_tokens_details":{"text_tokens":650,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3547,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":650,"tokens_out":60,"duration_ms":42203,"temperature":1.0,"reasoning_tokens":3547,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-21T12:13:15.343135+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A direct comparison in which one set of model outputs is scored only against the normative utility references and another only against the descriptive behavior references, then checked for whether the two rankings produce measurably different portfolio outcomes over the stated investment horizon.","supporting_citations":[],"review_version":1}