{"id":"4acd5abf-794e-4134-ba55-2012d3551823","arxiv_id":"2604.08362","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Introduces OmniBehavior benchmark from real-world data and shows LLMs exhibit hyper-activity, persona homogenization, and utopian bias in behavior simulation.","lead":"This paper introduces OmniBehavior, the first benchmark for LLM user simulation built entirely from real-world long-horizon behavior traces across multiple scenarios. It finds that current models produce homogenized, overly positive simulations that lose individual differences and rare behaviors.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"OmniBehavior real-world traces may embed selection/measurement biases that artifactually produce the reported LLM homogenization and utopian bias","rationale":"The reader correctly isolated the representativeness of the real-world traces as the weakest assumption. This is load-bearing because every quantitative claim about LLM bias is a comparison against those traces; any systematic distortion in the reference data directly undermines the interpretation of the differences as LLM-specific structural bias. No other internal inconsistency is visible from the abstract, and the reader’s low-confidence UNVERDICTED stance already reflects the missing methods detail.","tokens_in":1723,"tokens_out":426,"duration_ms":20944,"concrete_test":"In the methods section describing OmniBehavior construction, extract the participant recruitment protocol, data sources, and any validation of trace completeness or demographic representativeness. Recompute the key bias metrics (hyper-activity rate, persona entropy, utopian-bias score) after re-weighting the reference traces to match external population statistics or after restricting to fully observed long-horizon subsets; if the LLM–human gap shrinks by >20 % under re-weighting, the structural-bias conclusion is sensitive to data-collection assumptions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that LLMs exhibit structural bias toward a positive average person, hyper-activity, persona homogenization, and loss of long-tail behaviors—rests entirely on systematic differences between LLM outputs and the OmniBehavior ground-truth traces. For this difference to indicate an LLM-specific structural bias rather than a data artifact, the real-world traces must constitute an unbiased, representative sample of long-horizon, cross-scenario human decision-making. The abstract states the benchmark is “constructed entirely from real-world data,” but provides no detail on recruitment, logging completeness, demographic coverage, or handling of missing long-tail events. If the underlying data over-samples active users, digitally logged actions, or self-selected participants, the observed “convergence” and “utopian bias” could simply reflect mismatch between the (biased) reference distribution and the LLM prior, not an intrinsic LLM failure mode.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces OmniBehavior, a benchmark for LLM-based user simulation constructed entirely from real-world long-horizon, cross-scenario, and heterogeneous behavioral traces. It argues that existing isolated-scenario datasets suffer from tunnel vision compared to real-world causal chains, shows that state-of-the-art LLMs struggle to simulate these behaviors with performance plateauing despite larger context windows, and identifies a structural bias in LLMs toward a positive average person, manifested as hyper-activity, persona homogenization, and utopian bias that erases individual differences and long-tail behaviors.","tokens_in":1886,"tokens_out":484,"duration_ms":28254,"significance":"If the empirical comparisons and bias findings hold after rigorous validation of the ground-truth data, this work would be significant for advancing user simulation research in NLP and HCI. It provides the first unified real-world benchmark beyond synthetic or narrow scenarios and surfaces concrete failure modes (homogenization, loss of long-tail events) that could guide mitigation strategies in generative behavior modeling.","major_comments":[{"comment":"The manuscript states that OmniBehavior is 'constructed entirely from real-world data' and that systematic differences reveal LLM structural bias, but supplies no details on recruitment, logging completeness, demographic coverage, or handling of missing long-tail events. This is load-bearing for the central claim because the reported convergence to a positive average person and loss of individual differences could arise from selection or measurement biases in the reference traces rather than an intrinsic LLM property.","section":"Dataset construction / OmniBehavior description"},{"comment":"The abstract reports 'extensive evaluations' of LLMs, performance plateauing with context expansion, and a 'fundamental structural bias,' yet provides no metrics (e.g., behavioral divergence, accuracy on action sequences), statistical tests, data scale (number of users/traces), or controls. Without these, the evidence for both the simulation failures and the specific bias patterns (hyper-activity, homogenization, utopian bias) cannot be assessed for robustness.","section":"Evaluation and results"}],"minor_comments":[{"comment":"Clarify the precise definition of 'utopian bias' and 'positive average person' with concrete examples from the traces to avoid ambiguity in interpretation.","section":"Abstract / bias analysis"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive comments on our manuscript. We address each of the major comments in detail below and have prepared revisions to improve the clarity and completeness of the paper.","responses":[{"response":"We agree that providing more details on dataset construction is important for validating our claims. In the revised version of the manuscript, we will expand the relevant section to include information on recruitment procedures, logging completeness, demographic coverage summaries, and methods for handling missing long-tail events. This will help demonstrate that the observed biases are not artifacts of data collection biases. We note that ethical and privacy considerations limit the extent of detail we can provide on individual participants.","revision_made":"yes","referee_comment":"[Dataset construction / OmniBehavior description] The manuscript states that OmniBehavior is 'constructed entirely from real-world data' and that systematic differences reveal LLM structural bias, but supplies no details on recruitment, logging completeness, demographic coverage, or handling of missing long-tail events. This is load-bearing for the central claim because the reported convergence to a positive average person and loss of individual differences could arise from selection or measurement biases in the reference traces rather than an intrinsic LLM property."},{"response":"The full paper contains these metrics and details in the experiments and results sections. To address the concern about accessibility, we will update the abstract to briefly mention key quantitative findings and include a summary of the evaluation metrics, statistical tests, data scale, and controls in the main text or a new table. This revision will make the evidence more readily assessable while preserving the paper's structure.","revision_made":"yes","referee_comment":"[Evaluation and results] The abstract reports 'extensive evaluations' of LLMs, performance plateauing with context expansion, and a 'fundamental structural bias,' yet provides no metrics (e.g., behavioral divergence, accuracy on action sequences), statistical tests, data scale (number of users/traces), or controls. Without these, the evidence for both the simulation failures and the specific bias patterns (hyper-activity, homogenization, utopian bias) cannot be assessed for robustness."}],"tokens_in":1407,"tokens_out":454,"duration_ms":38286,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's core contribution is a benchmark built from actual user traces that span multiple scenarios and long time horizons, then used to test how well current LLMs can replay those traces. It reports that the models settle on an overly positive, high-activity average persona and lose the individual variation and rare behaviors present in the logs. That pattern is presented as a structural limitation rather than a fixable prompt issue. The shift away from synthetic or single-scenario data is the clearest advance; it lets them demonstrate that isolated benchmarks miss the causal chains across contexts that real decisions involve. The plateau in performance with larger context windows is also shown directly against the real traces. Those comparisons are the parts that could matter for people trying to build agent simulators or user models. The main uncertainty sits with the ground-truth data. The abstract claims the traces are entirely real-world, but without specifics on how participants were recruited, how complete the logging was across quiet periods, or how demographic coverage was checked, it is hard to separate LLM shortcomings from possible skew in the reference set. If the logs over-sample active or digitally visible users, the reported homogenization and utopian tilt could partly reflect that mismatch instead of an intrinsic model property. The paper would benefit from explicit controls or sensitivity checks on the trace collection. This work is aimed at researchers building or evaluating behavior simulators for HCI, virtual environments, or personalized systems. Anyone already running LLM-based agents will find the empirical gaps useful to see. It is coherent enough on its own terms to warrant referee time, even though the data-quality questions will likely require revision. I would send it out for review rather than desk reject.","headline":"OmniBehavior gives a real-data benchmark for long-horizon LLM simulation and flags a convergence bias, but the data collection details will decide if the bias claim holds.","tokens_in":2421,"tokens_out":404,"would_cite":false,"duration_ms":28479,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[{"relation":"echoes","rs_module":"IndisputableMonolith/Cost/FunctionalEquation.lean","rs_theorem":"washburn_uniqueness_aczel; Jcost_pos_of_ne_one","paper_passage":"LLMs tend to converge toward a positive average person, exhibiting hyper-activity, persona homogenization, and a utopian bias. This results in the loss of individual differences and long-tail behaviors."},{"relation":"echoes","rs_module":"IndisputableMonolith/Cost.lean","rs_theorem":"Jcost_unit0; Jcost_pos_of_ne_one","paper_passage":"Real human behavior is inherently sparse, with positive interaction rates remaining below 10%. By contrast, all evaluated LLM-based simulators exhibit a hyper-activity bias."}],"headline":"LLM homogenization to 'positive average person' and suppression of long-tail/negative behaviors parallels J-cost minimum at identity (x=1)","alignment":"aligned","rationale":"The paper's core empirical finding (hyper-activity, persona homogenization, utopian/positivity bias, loss of individual differences and long-tail behaviors) is structurally compatible with the RS recognition cost J(x) = ½(x + x⁻¹) − 1 whose unique minimum is exactly at the identity x=1 (J(1)=0) with strictly positive cost elsewhere. This forces convergence toward the calibrated average while penalizing deviations, mirroring the observed 'positivity-and-average' filter. However the paper is purely empirical/ML benchmarking with no cost functions, golden-ratio ladders, or formal forcing; the parallel is observational rather than isomorphic.","tokens_in":56975,"confidence":"low","tokens_out":368,"duration_ms":41160,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"LLMs simulating real human behavior converge toward a positive average person and erase individual differences.","keywords":["LLM user simulation","behavior benchmark","structural bias","persona homogenization","long-horizon traces","real-world data","utopian bias"],"falsifier":"Collect a fresh set of long-horizon traces from a demographically different population that explicitly includes documented long-tail decisions, then measure whether LLM outputs still flatten those decisions into positive averages.","tokens_in":2625,"feed_emoji":"📊","tokens_out":625,"duration_ms":36041,"temperature":0.7,"pith_summary":"This paper builds OmniBehavior, a benchmark drawn entirely from real-world traces, to test how well large language models can act as user simulators across long sequences that cross multiple life scenarios. It first shows that prior benchmarks using isolated or synthetic settings miss the causal chains that link decisions over time in actual human lives. When state-of-the-art models are evaluated on the new benchmark, they produce behaviors that are more active, more uniform across people, and more optimistic than the source data. The resulting structural bias removes the variability and infrequent patterns that define real individuals, limiting how faithfully any downstream application can replay or predict human actions.","feed_headline":"LLMs default to positive average personas in behavior simulations","feed_subtitle":"OmniBehavior benchmark shows hyper-activity and loss of individual differences plus long-tail behaviors","key_machinery":"The OmniBehavior benchmark, which assembles long-horizon, cross-scenario, and heterogeneous behavioral patterns directly from real-world data to serve as ground truth for simulation fidelity.","core_discovery":"A systematic comparison between simulated and authentic behaviors uncovers a fundamental structural bias: LLMs tend to converge toward a positive average person, exhibiting hyper-activity, persona homogenization, and a utopian bias. This results in the loss of individual differences and long-tail behaviors.","pith_inferences":["The same convergence may appear in other generative tasks that rely on modeling user preferences or sequences, such as personalized recommendation or dialogue systems.","A practical extension would be to add explicit regularization or retrieval steps that force models to reproduce measured frequencies of rare actions from the source traces.","Testing whether the bias persists when models are given explicit negative or low-activity examples from the same data would clarify whether the issue is data scarcity or architectural.","If the homogenization is confirmed across multiple languages or cultures, it would indicate a training-data skew rather than a language-specific artifact."],"forward_implications":["Isolated-scenario datasets create tunnel vision that hides the cross-scenario causal chains present in real decision-making.","LLM simulation performance plateaus even when context windows are enlarged.","The structural bias produces outputs that systematically omit the low-frequency behaviors observed in authentic traces.","High-fidelity simulation will require explicit mechanisms to preserve individual differences rather than defaulting to an averaged persona."],"fun_headline_variants":["LLMs converge to positive average person in simulations","OmniBehavior exposes hyper-activity in LLM behavior sims","Simulated behaviors lack individual and long-tail differences","LLMs exhibit utopian bias and persona homogenization"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The real-world behavioral traces collected for OmniBehavior accurately and representatively capture authentic long-horizon, cross-scenario human decision-making without significant selection or measurement biases.","fun_headline_variants_meta":{"raw":{"variants":["LLMs converge to positive average person in simulations","OmniBehavior exposes hyper-activity in LLM behavior sims","Simulated behaviors lack individual and long-tail differences","LLMs exhibit utopian bias and persona homogenization"]},"model":"grok-4.3","cost_usd":0.010293,"raw_usage":{"total_tokens":4457,"prompt_tokens":626,"num_sources_used":0,"completion_tokens":58,"cost_in_usd_ticks":102928000,"prompt_tokens_details":{"text_tokens":626,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3773,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":626,"tokens_out":58,"duration_ms":48802,"temperature":1.0,"reasoning_tokens":3773,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-22T10:26:50.614032+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Collect a fresh set of long-horizon traces from a demographically different population that explicitly includes documented long-tail decisions, then measure whether LLM outputs still flatten those decisions into positive averages.","supporting_citations":[],"review_version":2}