{"id":"f24d62ca-eccb-4bc9-92e7-8b770e0e853b","arxiv_id":"2506.14448","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"LLMs improve only slightly and unstably from test-time experience on semantic reasoning games, while humans learn much faster.","lead":"This paper tests whether LLMs get better at reasoning games when they use lessons from their own past attempts, and compares them with human players. It finds small, inconsistent gains for LLMs and much faster, steadier learning in humans.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Twenty Questions judge is the same LLM under test; measured test-time gains may reflect self-consistency with the judge, not learning against a neutral environment.","rationale":"The reader's weakest assumption is exactly this self-simulation confound; I agree. The concern is load-bearing but not yet demonstrated, so a conditional verdict with a required control oracle is appropriate. The reader already marked CONDITIONAL, so no adjustment is needed. Secondary issues (AIME negative results, no confidence intervals) reinforce but do not replace this primary concern.","tokens_in":14050,"tokens_out":3655,"duration_ms":38181,"concrete_test":"Re-run the Twenty Questions evaluation (Table 2 rows and Figure 3) with the LLM judge replaced by a deterministic oracle built from the fixed 157-word list and an external semantic hierarchy (e.g., WordNet or a hand-specified category tree) that answers Yes/No/Invalid based on the true target word, keeping the same prompts, policy-update pipeline, and N=32/test rounds. Compare w/ Exp. Policy vs w/o Policy gains against the original self-judged setup. If the gains vanish or reverse, the measured test-time learning is at least partly self-consistency with the model's own judge; if gains remain within noise, the concern does not land.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim hinges on the Twenty Questions results (Table 2, Figure 3, Section 4.4). In Section 4.1 the authors state: 'the environment is simulated using the same model under evaluation to ensure alignment in question understanding and knowledge base.' The judge that produces Yes/No/Invalid is therefore the same LLM whose test-time learning is being measured, run at temperature 1. A policy distilled from five rounds of experience can then improve performance by learning the stochastic judge's own category ordering and lexical preferences, rather than by discovering objective structure of the 157-word set. This makes 'learning from experience' partially 'learning to be self-consistent with one's own judge.' The human comparison in Section 4.4 uses the same simulated environment, so the human-model gap is also contaminated by how well each human's questions align with the LLM judge's semantics. Unless the judge is an independent ground-truth oracle or a different, fixed model, the improvements in Table 2 and the cumulative curves cannot be attributed to test-time learning about the task rather than self-consistency.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper advocates test-time learning as a distinct evaluation axis for LLMs and proposes semantic games—Twenty Questions and Who is Undercover—together with AIME 2025 as testbeds. It introduces a lightweight framework that compares model performance without policy, with a rule-based policy, with a policy distilled from five rounds of the model's own experience, and with a human-authored policy, and it adds an incremental setting that tracks cumulative reward over 50 rounds. On this basis the authors report that experience-derived policies yield measurable improvements in the two semantic games, that cumulative gains are model-dependent (Claude being the clearest beneficiary), and that eight human participants learn faster and more stably than the best model. The paper concludes that LLMs have measurable but limited test-time learning ability relative to humans.","tokens_in":14226,"tokens_out":6871,"duration_ms":69406,"significance":"If the central finding held, the paper would add a useful evaluation axis beyond static benchmarks and provide a repeatable protocol for comparing human and machine learning curves. The design has several strengths: it uses a policy-based representation of experience rather than unbounded context, it includes an external human baseline, and it ships code and data. The contribution is conditional, however, because the main quantitative claims rest on a self-simulated environment and on small improvements without reported variance. The framing contrast with humans is valuable but needs stronger experimental controls before the general claim can be accepted.","major_comments":[{"comment":"Section 4.1 states that in Twenty Questions \"the environment is simulated using the same model under evaluation,\" and Section 3.1 says the environment responds with Yes/No/Invalid. Since the judge is the same LLM being measured, a policy distilled from five prior rounds can improve by learning the stochastic judge's category ordering and lexical preferences (including its notion of \"Invalid\") rather than by discovering the objective structure of the 157-word set. The human comparison in Section 4.4 uses this same simulated judge, so the human-model learning-speed gap is also contaminated by how well each participant's questions align with the LLM judge's semantics. Please either replace the self-simulation with a fixed independent judge (e.g., a different, frozen model or a rule-based oracle over the word set) or provide a control showing that the learned policies also improve against ground-truth category labels. Without this, Findings 1 and 4 are not established for the Twenty Questions testbed.","section":"Sections 3.1, 4.1, 4.2, 4.4"},{"comment":"Finding 1 claims \"measurable improvements across models and tasks,\" but the AIME 2025 rows in Table 2 show negative or zero improvement for all three models (GPT-4o -30.57%, Claude 3.5 Sonnet 0.00%, DeepSeek-V3 -10.62%). This is a direct contradiction of the cross-task wording. The paper should either restrict Finding 1 to the two semantic games, or provide a substantive explanation of why AIME is a valid test-time-learning testbed and why the negative result is consistent with the framework.","section":"Section 4.2, Table 2 (AIME 2025 rows)"},{"comment":"No uncertainty quantification is reported for the central comparisons. In Twenty Questions the reported experience-policy improvements are 3.97–6.33 percentage points; with M=32 test cases, sampling noise can be material. In Who is Undercover, 32-game win rates have binomial standard errors up to roughly 9 percentage points, so the 10–25 point gains need confidence intervals or paired tests to be described as \"measurable.\" The human study comprises eight participants with no statistical test of the human-model gap, and Figure 4 shows two groups without error bars. Please report per-run values, bootstrap intervals, or significance tests for the improvements and for the human comparison.","section":"Tables 2–3, Figures 3–4"}],"minor_comments":[{"comment":"The definition of r_exp(t) for t=1 appears to mix baseline and experience rewards in the numerator; please clarify the intended initialization and why it differs from t>1.","section":"Section 3.2.2, Eq. (1)"},{"comment":"The figure caption refers to an \"upper figure\" and a \"lower figure,\" but the panels are not labeled in the image; add (a)/(b) labels and describe the split criterion (performance variance) in the caption.","section":"Section 4.4, Figure 4"},{"comment":"The paper reports a pilot study comparing full-history and policy-based representations and says the latter underperformed, but no pilot results are shown; please include the numbers or state that they are provided in supplementary material.","section":"Section 3.2.1"},{"comment":"The human policy in Appendix B.1 contains a typo (\"electoricity\" instead of \"electricity\"); please copyedit the appendix, which also has inconsistent formatting in the policy lists.","section":"Appendix B.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is topical for the journal, but the self-simulation issue is the key technical risk. The AIME negative result should be handled honestly in revision; if the authors can add an independent judge control, the paper could become a solid contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper gives the community something real: a measurement axis for test-time learning, with a clean pipeline and a genuine human comparison. That is worth having, and the authors deserve credit for framing the question clearly and for not overselling the cumulative curves, where instability is visible.\n\nWhat is actually new is applying the adaptive memory pipeline (Suzgun et al.) to open-ended semantic games, and adding a human learning baseline. The finding that non-thinking models improve modestly with experience-based policies, while thinking models do not, is interesting even if the latter is consistent with known few-shot degradation in R1. The AIME contrast is a nice touch, even with negative results.\n\nThe paper is also honest. The Limitations section admits the narrow environment set, the small human sample, and the absence of test-time training. The neutral role-name fix for the undercover game is a thoughtful detail.\n\nThe soft spots are real and load-bearing. In Twenty Questions, the judge is the same LLM under test, run at temperature 1. That is stated plainly in Section 4.1, so it is not hidden. But it means the measured improvement could be the model learning its own judge's biases and category ordering, not the objective structure of the word set. The human baseline in Section 4.4 uses the same simulated judge, so the human-model gap also reflects how well each human's questions align with that judge's semantics. This confound weakens the central claim that LLMs show test-time learning on this task.\n\nThe AIME results contradict the 'across models and tasks' framing in Finding 1: two models get worse, one is flat. The claim should be qualified to the semantic games. Also, there is no uncertainty quantification anywhere; with M=32, three cumulative samples, and eight humans, point estimates overstate precision. The rule-based policy baseline hurts all models in Twenty Questions, which is itself a finding, but it suggests the policy pipeline is sensitive to prompt style and needs more discussion.\n\nNone of this kills the paper's value as a proposal for a new evaluation axis. The confound is addressable with an independent judge or a fixed different model, and adding confidence intervals would tighten the claims. As is, the central comparison needs a caveat.\n\nThis paper deserves peer review. The evaluation idea is important enough, and the empirical work is honest enough, that a good referee can push the authors to fix the controls and soften the claims. I would not cite the current version's numbers as evidence for test-time learning ability, but I would cite the framing and the game-based testbeds. Bring it to reading group if your group cares about LLM evaluation or agentic learning; there is a good discussion to be had about what counts as learning when the environment is the model itself.","headline":"A useful evaluation axis for test-time learning, but the main semantic-game result is confounded by using the same LLM as judge; still worth a serious referee.","tokens_in":14772,"tokens_out":1961,"would_cite":true,"duration_ms":23935,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLMs can improve from experience, but they learn more slowly and less steadily than humans.","keywords":["test-time learning","LLM evaluation","semantic games","Twenty Questions","Who is Undercover","human comparison","experience-based policy","cumulative learning"],"falsifier":"Run the same Twenty Questions protocol with an independent answerer (a different model, or a human who knows the target word) instead of the questioner's own LLM, keeping the policy-distillation pipeline fixed; if the experience-derived policy no longer improves scores, the measured learning is self-consistency, not general improvement. A second check is the cumulative curve: a model whose experience-enabled reward stays at or below its no-experience baseline across fifty rounds would refute the claim that gains accumulate.","tokens_in":13819,"feed_emoji":"🧠","tokens_out":7629,"duration_ms":72924,"temperature":0.7,"pith_summary":"The paper argues that how quickly a model improves from its own experience at test time is a distinct and under-measured axis of intelligence, separate from how well it performs on static benchmarks. To measure it, the authors build an evaluation around semantic games—Twenty Questions and Who is Undercover, plus the AIME math exam—in which the model plays rounds, reflects on its interactions and rewards, and distills that reflection into a short policy that guides the next round. Under this protocol, current LLMs do show measurable gains from experience-derived policies, but the gains are uneven: they shrink or reverse as experience accumulates, and they are outweighed by policies written by human players. The authors read this as evidence that LLMs have real but immature test-time learning ability, and that the gap between self-derived and human-authored strategies marks concrete headroom for improvement.","feed_headline":"LLMs improve from experience, yet learn slower than humans","feed_subtitle":"In a Twenty Questions benchmark, models do improve from their own past rounds, yet humans climb faster and more steadily.","key_machinery":"The load-bearing mechanism is the policy-based distillation of experience: after each block of games, the model is given its past dialogue interactions, rewards, and its own reflections, and is prompted to write a short test-time policy (roughly 250 tokens) that it will follow in the next block. This compresses the full history into an executable strategy, making repeated experience cheap enough for a cumulative setting, where policies are updated as rounds grow. The evaluation framework contrasts four experience representations—no policy, a rule-only policy, an experience-derived policy, and a human-authored policy—and runs them in two regimes: a fixed five-round experience setting and an incremental cumulative setting over fifty rounds. The semantic games themselves are chosen as testbeds because they are open-ended, resistant to saturation, and require discovering latent strategies rather than recalling memorized answers.","core_discovery":"The paper's central claim is that large language models can improve their performance on experience-based, reasoning-intensive tasks through test-time experience, but not as well as humans. Concretely, when a model is given a five-round history of its own questions, answers, and rewards and is asked to write a policy from that history, its subsequent performance improves on both Twenty Questions and Who is Undercover—this is Finding 1. Yet the improvement is not robust: in the cumulative setting, only Claude sustains a consistent advantage from accumulating experience, while GPT-4o's gains appear late and DeepSeek-V3 degrades after five rounds (Finding 3). Eight human players, given the same Twenty Questions setup, improve faster and more steadily than the best model, approaching near-perfect binary-questioning performance (Finding 4). The paper also reports that thinking models such as o1 and DeepSeek-R1 show no test-time learning gains from self-derived policies, which it links to R1's known few-shot degradation (Finding 5).","pith_inferences":["A testable extension is to replace the self-simulated answerer with another model or a human oracle; if learning gains vanish, the measured improvement may be calibration to the judge's own biases rather than general learning.","The framework could be moved from games to real-world interactive tasks such as customer support or scientific tool use, where feedback is noisy; the cumulative setting's instability would then become a practical reliability metric.","The policy-distillation step resembles an in-context version of what fine-tuning does in-weight; open-weights models could be tested with explicit test-time training to see whether parameter updates close the gap to human learning curves.","The eight-participant human sample is a floor, not a ceiling; a larger, more diverse sample would let the human learning curve be decomposed into strategy acquisition versus task familiarity effects."],"forward_implications":["Static benchmark rankings will systematically overstate how capable a model is as a learner; two models with equal static scores can differ sharply in whether and how fast they improve from experience.","Test-time learning should be treated as a separate evaluation axis, with both a fixed-experience score and a cumulative-learning curve, because the two can disagree.","Because human-authored policies consistently beat self-derived ones, current models have headroom: better strategy induction, not just more compute at inference, is what stands between them and human-level learning.","The absence of test-time gains in thinking models suggests that long chain-of-thought reasoning optimized for single problems does not automatically transfer to accumulating and applying experience across episodes.","Cumulative learning trajectories discriminate among models, so reporting only average gains can hide the instability that matters for deployment in interactive settings."],"supporting_citations":[{"why":"Supplies the Twenty Questions task framing and the multi-turn reinforcement learning benchmark lineage the study builds on.","marker":"Abdulhai et al., 2023"},{"why":"Provides the 157-word candidate set and the fixed game setup used in the Twenty Questions environment.","marker":"Zhou et al., 2024"},{"why":"Provides the fixed-number-of-experience evaluation setup and the in-context reinforcement learning paradigm this work adapts.","marker":"Laskin et al."},{"why":"Supplies the memory-management pipeline used for the incremental, cumulative experience setting.","marker":"Suzgun et al., 2025"},{"why":"Provides the Who is Undercover multi-agent game used as a second semantic game testbed.","marker":"Xu et al., 2023"},{"why":"Defines the AIME 2025 math benchmark used as the third testbed for test-time learning.","marker":"MAA, 2025"},{"why":"Documents DeepSeek-R1's few-shot degradation, which the paper invokes to explain why thinking models fail to gain from experience.","marker":"Guo et al., 2025"},{"why":"Identifies o1 as a thinking model whose behavior under self-derived policies is tested in the thinking-model experiments.","marker":"Jaech et al., 2024"}],"fun_headline_variants":["LLMs improve with experience, but humans improve more","Experience helps LLMs, but humans learn faster","LLMs show test-time learning, but not like humans","Human comparison: LLMs learn slower from experience","LLMs can learn from experience, yet trail humans"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole comparison assumes that using the same LLM to simulate the game environment produces a neutral, trustworthy source of feedback; if the model is instead just getting better at predicting its own judge's behavior, the measured learning may not transfer to real interactions.","fun_headline_variants_meta":{"raw":{"variants":["LLMs improve with experience, but humans improve more","Experience helps LLMs, but humans learn faster","LLMs show test-time learning, but not like humans","Human comparison: LLMs learn slower from experience","LLMs can learn from experience, yet trail humans"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000204,"raw_usage":{"total_tokens":1392,"prompt_tokens":948,"completion_tokens":444,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":564,"completion_tokens_details":{"reasoning_tokens":369}},"tokens_in":564,"tokens_out":444,"duration_ms":4764,"temperature":1.0,"reasoning_tokens":369,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:17:29.019893+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same Twenty Questions protocol with an independent answerer (a different model, or a human who knows the target word) instead of the questioner's own LLM, keeping the policy-distillation pipeline fixed; if the experience-derived policy no longer improves scores, the measured learning is self-consistency, not general improvement. A second check is the cumulative curve: a model whose experience-enabled reward stays at or below its no-experience baseline across fifty rounds would refute the claim that gains accumulate.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the AIME 2025 math benchmark used as the third testbed for test-time learning."}],"review_version":1}