{"id":"6935372c-84c4-4ea1-bc00-311051e81198","arxiv_id":"2607.14485","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Step-level human preference data collected via SimPref, then SFT+DPO, improves long-horizon social-simulation behavior of open-weight LLM agents on held-out events.","lead":"The authors built SimPref, an interface that lets humans supervise and choose intermediate decisions — planning, memory, reflection, actions, dialogue — of AI agents in social simulations. They collected 57,239 step-level human preference pairs and used them to fine-tune open-weight language models, finding consistent gains in multi-day held-out social events.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Missing control: SFT on human-preferred outputs may simply distill GPT-4o; the unique contribution of human step-level preference is not isolated.","rationale":"The reader's weakest assumption focuses on LLM-as-a-judge bias in the evaluation metrics. That is a valid concern about the measurement of the outcome. However, the more load-bearing issue for the paper's central claim is internal attribution: the training signal that produces the headline improvements is a blend of (i) GPT-4o output quality and (ii) human selection among GPT-4o candidates. Because 97.9% of accepted outputs are GPT-4o-generated and the authors already collect LLM self-preferences, a simple control could separate these factors. The paper does not include it, and its limitations section does not acknowledge this confound. The LLM-judge concern would remain important even if the attribution were clean, but the missing control threatens the causal claim irrespective of measurement. The existing CONDITIONAL verdict is appropriate; it should perhaps explicitly require this control (or a human evaluation) before acceptance. I do not see grounds to reject, as the dataset and interface have independent value and the held-out generalization design is reasonable. The concern is concrete and testable, so the verdict stays conditional rather than moving to accept or reject.","tokens_in":12174,"tokens_out":6022,"duration_ms":60556,"concrete_test":"Train three additional Q7B (or L8B) models from the same base with identical hyperparameters and data splits: (a) SFT on human-preferred y+ (the current method), (b) SFT on GPT-4o's self-ranked top-1 candidate for the same triggers, and (c) SFT on a randomly sampled GPT-4o candidate per trigger. Then evaluate all three under the exact Table 2 protocol. If (b) or (c) performs within noise of (a) on the five metrics, the improvement is not attributable to human step-level preference. If (a) clearly outperforms both, the central claim survives. This control is directly enabled by the stored LLM self-rankings in §5.2.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that step-level human preference is an effective training signal. But the accepted outputs y+ in the dataset are almost always one of three GPT-4o-generated candidates: only 1,217 of 57,239 pairs (2.1%) are human-customized (§4.2, Table 1). Therefore the SFT stage, which accounts for most of the gains in Table 2, is essentially supervised fine-tuning on GPT-4o outputs filtered by human choice. The paper records the LLM's own self-preference ranking among the three candidates for every trigger (§5.2), yet never trains a control on the LLM's top-1 candidate or on random GPT-4o candidates. Without such a control, the large SFT improvements (e.g., Q7B-SFT gains of +0.76 to +1.14 across metrics) could be attributable to distillation of a stronger model rather than to human selection. The DPO stage, which is the only component uniquely tied to human contrastive preferences, yields modest and uneven gains over SFT (e.g., L8B temporal adherence drops from 2.75 to 2.56; Q7B requirement consistency drops from 2.93 to 2.89). The +MI ablation does not address this confound, as it holds training labels fixed and only changes the retrieval scoring. Thus the paper's central attribution of improvements to human step-level supervision is not yet established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces SimPref, an interactive simulation interface for collecting step-level human preference supervision over the intermediate decisions of GA-style generative agents. The authors construct a dataset of 57,239 preference pairs spanning six agent modules across 30 social events, then train open-weight LLMs (Qwen2.5-7B/14B, Llama-3.1-8B) with SFT on human-accepted outputs and DPO on accepted--rejected pairs. They report consistent improvements over five whole-trajectory metrics on 10 held-out events, a shift in behavioral time allocation toward a human reference (KL 0.610 to 0.084), and qualitative case studies. The central claim is that step-level human supervision is an effective training signal for improving both local decision quality and long-horizon agent behavior.","tokens_in":12435,"tokens_out":4615,"duration_ms":44690,"significance":"If the central claim is established, this would be a useful contribution: it is the first step-level human preference dataset for GA-style social simulation, and the interface and training pipeline are valuable resources. The +MI ablation and the KL-to-human behavioral analysis are thoughtful attempts to isolate mechanisms and long-horizon effects. However, the attribution of the gains to human step-level supervision is currently not supported by the experimental design, and the evaluation relies entirely on LLM-as-a-judge scoring. The paper's own limitations section acknowledges the judge-bias and missing inter-annotator agreement, but those limitations are load-bearing for the headline results.","major_comments":[{"comment":"The central attribution is confounded by model distillation. Only 1,217 of 57,239 pairs (2.1%) are human-customized; the SFT stage trains almost entirely on GPT-4o-generated outputs that a human selected. SFT accounts for most of the Table 2 gains (e.g., Q7B-SFT +0.76 to +1.14 across metrics), yet no control is trained on the LLM's own top-1 candidate or on random GPT-4o candidates. The +MI ablation in §6.2 does not address this confound because it changes only retrieval scoring while holding training labels fixed. A control isolating human selection (e.g., SFT on GPT-4o top-1 vs. SFT on human-selected candidates) is needed to support the claim that step-level human supervision, rather than distillation of a stronger model, drives the improvements.","section":"§4.2, Table 1, §6.2"},{"comment":"Only mean Likert scores are reported, without standard deviations, confidence intervals, or significance tests. With 10 held-out events and 3 episodes per method, the claimed consistent improvements—especially the modest, uneven DPO gains (e.g., L8B temporal adherence drops from 2.75 to 2.56; Q7B requirement consistency drops from 2.93 to 2.89)—may be within noise. Per-event score distributions, paired tests, or effect sizes are needed to substantiate the 'consistent improvement' and the DPO contribution claims.","section":"Table 2, §6.2"},{"comment":"All quantitative claims rest on GPT-5.2 and DeepSeek-v3.2 Likert judgments, and the human reference trajectories are scored by the same judges. The paper acknowledges evaluator bias in §7, but because the evaluation metrics (requirement consistency, role fulfillment, temporal adherence) closely mirror the annotation criteria, there is a nontrivial risk that the LLM judges reward stylistic artifacts of SFT/DPO rather than true behavioral fidelity. A human evaluation of a subset of trajectories, or at least a judge-agreement and robustness analysis, is necessary to ground the headline improvements.","section":"§6.1, §7"}],"minor_comments":[{"comment":"There are several typos/OCR artifacts: 'diﬀicult' in the Introduction, 'suﬀicient' in §6.2, and 'A veraged' in §6.2. Please proofread.","section":"§6.2"},{"comment":"The Top-1 Match numbers are compared against a naive 33% baseline, but for the Importance Score module the output space is smaller and chance agreement is higher. The paper notes this but does not quantify the adjusted baseline; please report a chance-corrected agreement or otherwise handle the output-space size.","section":"Table 1, §5.2"},{"comment":"The absence of inter-annotator agreement is acknowledged in §7, but given that the entire dataset rests on single-annotator judgments, a small overlap sample with agreement statistics would substantially strengthen the dataset's credibility.","section":"§5.1, §7"},{"comment":"The KL divergence is computed over normalized frequencies of six behavioral categories for Q14B only. Please report the exact category definitions, the coding procedure, and the raw counts, since KL estimates over six bins can be unstable when counts are small.","section":"§6.3"}],"recommendation":"major_revision","confidential_remarks":"The missing-control issue is decisive for the central claim. I am not asking for a full rejection because the data resource and pipeline are valuable and the claim may be salvageable, but the current experiments do not isolate human step-level preference from GPT-4o distillation. If the proposed control shows no difference from top-1 GPT-4o SFT, the paper should be reframed as human-filtered distillation; if it shows a clear difference, the main claim would be established. The LLM-as-judge evaluation is a secondary but real concern; a modest human evaluation or judge-calibration study would help. The authors' candid limitations section is appreciated but does not fix these load-bearing gaps."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, what's actually new: the SimPref interface and the 57K-pair step-level preference dataset are a real first for GA-style social simulation. The collection design — candidate outputs from GPT-4o, human selection/customization, recording the LLM's own ranking — is thoughtful, and the dataset composition table is useful. The held-out evaluation across 10 events, 3 open-weight models, multiple metrics, plus the behavioral time-allocation analysis and KL-to-human drop, is more than most dataset papers do. The +MI control, showing that improving the importance estimator alone doesn't get the gains, is a good check. I agree with the reader that this is a conditional accept-level paper, not a reject.\n\nBut there is a load-bearing gap that the stress-test correctly identifies: 97.9% of the accepted outputs are GPT-4o-generated, chosen by humans. The SFT stage, which produces most of the Table 2 gains, is therefore essentially fine-tuning on GPT-4o outputs filtered by human choice. Without a control trained on the LLM's own top-1 candidates or on unselected GPT-4o outputs, you cannot attribute those gains to human step-level preference. The DPO stage, the only component uniquely tied to human contrastive preferences, gives modest and uneven improvements. So the central claim in the abstract and conclusion is, strictly speaking, not yet established.\n\nSecondary concerns: the evaluation rests on LLM-as-a-judge Likert scores for all five trajectory metrics, with no error bars, no significance tests, and no human evaluation of the final models. The paper acknowledges the evaluator-bias limitation, which is honest, but it doesn't fix it. The behavioral distribution analysis is a more objective check and the KL drop is striking, though I'd want to know exactly how the six categories were assigned. Also, no inter-annotator agreement, and the dataset/code are not yet available, so replication is impossible right now.\n\nThe weaknesses are fixable. A control on GPT-4o top-1 SFT, confidence intervals or bootstrap, a small human eval, and releasing the artifacts would make the claim much stronger. As is, the paper is a solid dataset contribution with an over-interpreted headline. I'd send it to review, require the control, and expect a revised version that narrows the claim appropriately.\n\nWho it's for: anyone working on LLM agents, social simulation, or preference data collection. The dataset itself will likely be useful to the community once released.","headline":"A genuinely useful dataset and interface, but the paper's headline claim that human step-level preferences drive the gains is not yet isolated from GPT-4o distillation.","tokens_in":12990,"tokens_out":2718,"would_cite":true,"duration_ms":26557,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that collecting human preferences over the intermediate decision steps of generative agents—rather than only over final trajectories—provides an effective training signal, improving both local decision quality and long-hor","keywords":["social simulation","generative agents","step-level preference learning","human-in-the-loop annotation","direct preference optimization","long-horizon behavior","memory importance","LLM-as-a-judge evaluation"],"falsifier":"Run a blind human evaluation of full trajectories from base versus trained agents on held-out events and check whether human raters consistently prefer the trained agents on the same five dimensions; if they do not, or if objective event-completion markers such as guests actually attending the event and information propagating correctly show no improvement, the claim that step-level supervision improves long-horizon behavior fails.","tokens_in":12017,"feed_emoji":"🤖","tokens_out":3142,"duration_ms":34672,"temperature":0.7,"pith_summary":"The paper introduces an interactive simulation interface through which human annotators supervise the moment-by-moment decisions of generative agents: planning, memory retrieval, reflection, action, and dialogue. Using this interface, the authors collect 57,239 step-level preference pairs across 30 social events, then fine-tune open-weight language models with supervised learning and direct preference optimization on those pairs. On 10 held-out events with different scripts and goals, the trained models consistently improve all five whole-trajectory metrics over their base versions. The trained agents also spend time across behavioral categories much closer to a human reference, with the divergence dropping from 0.610 to 0.084. The paper's central claim is that fine-grained human preferences over intermediate decisions are an effective training signal for producing more socially competent long-horizon agents.","feed_headline":"Step-level human preferences make simulated agents behave more humanly","feed_subtitle":"57K fine-grained choices improve every trajectory metric and shrink the time-allocation gap to humans eightfold.","key_machinery":"The central mechanism is the step-level preference tuple (x, k, y+, y−), where x is the agent's decision context—partial observations, explicit goal, retrieved memories, and local state—k is the triggered module, y+ is the human-preferred output, and y− is a rejected alternative. The interaction interface exposes these decision contexts to annotators and records their choices across six modules. A module-conditioned backbone language model is then trained first by supervised fine-tuning on accepted outputs and second by direct preference optimization on the preference pairs, aligning local decisions with human judgments of decision competence.","core_discovery":"Step-level human supervision is an effective training signal for generative agents in social simulations. For each triggered decision module, the model is shown the same partial observations, goals, and retrieved memories the agent sees, and generates three candidate outputs; a human annotator chooses the most competent one or writes a custom alternative. Training on these accepted outputs (and contrasting them against rejected ones) markedly improves open-weight models across location adherence, temporal adherence, role fulfillment, requirement consistency, and interaction quality, with most of the gain coming from supervised fine-tuning and smaller, dimension-dependent gains from direct pr","pith_inferences":["Editorial extension: the dataset's low human–LLM agreement on action and importance-score modules (around 30% top-1 match, near random) suggests that current LLMs systematically misjudge feasibility and information grounding; this points to a concrete failure mode that step-level preference data could expose in other agent architectures.","Editorial extension: an implicit but testable consequence is that step-level preference learning will transfer across event structures that share the same module decomposition; the paper only tests scenario-level transfer within one architecture, so cross-architecture transfer is an open question.","Editorial extension: because annotation cost is concentrated in a small number of annotators and single-label per step, the method's value depends on whether decision-competence judgments are more reproducible across annotators than subjective taste, which the paper does not measure.","Editorial extension: one could design a cheaper data-collection pipeline using model-generated critiques to mimic the step-level preference oracle, and measure how much of the observed gain survives without human annotators."],"forward_implications":["If step-level supervision is effective, future agent-alignment pipelines can shift annotation effort from whole trajectories to the internal decisions that generate behavior, yielding more actionable signal per annotation.","Because most gains come from supervised fine-tuning on accepted outputs, fixing poor local decisions may matter more than adding general reasoning capability when building socially competent agents.","Memory-importance recalibration alone is insufficient; improving long-horizon social behavior requires coordinated improvement across multiple decision modules.","Step-level preference learning moves agents' long-horizon time allocation toward human behavioral distributions, suggesting that local supervision does not induce myopic or collapsing behavior over multi-day simulations.","Combining step-level preferences with trajectory-level supervision is a promising direction for jointly optimizing internal decisions and global social outcomes."],"fun_headline_variants":["Step-level human preferences make agents act more humanly in simulations","57K fine-grained choices train agents to better follow social norms","Step-level preference learning improves agent coordination and quality","Training on step-level human picks cuts agent time-gap to humans 8x","Step-level supervision beats full-trajectory labels for social agents"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The central claim collapses if the LLM judges used to score trajectories reward fluent or stylistically polished outputs rather than genuine behavioral fidelity, since every reported improvement is measured by those judges.","fun_headline_variants_meta":{"raw":{"variants":["Step-level human preferences make agents act more humanly in simulations","57K fine-grained choices train agents to better follow social norms","Step-level preference learning improves agent coordination and quality","Training on step-level human picks cuts agent time-gap to humans 8x","Step-level supervision beats full-trajectory labels for social agents"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000223,"raw_usage":{"total_tokens":1244,"prompt_tokens":644,"completion_tokens":600,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":388,"completion_tokens_details":{"reasoning_tokens":513}},"tokens_in":388,"tokens_out":600,"duration_ms":6870,"temperature":1.0,"reasoning_tokens":513,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T01:56:04.159504+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a blind human evaluation of full trajectories from base versus trained agents on held-out events and check whether human raters consistently prefer the trained agents on the same five dimensions; if they do not, or if objective event-completion markers such as guests actually attending the event and information propagating correctly show no improvement, the claim that step-level supervision improves long-horizon behavior fails.","supporting_citations":[],"review_version":1}