{"id":"9ba7592c-4f96-40fd-bcd0-f9be3eeaf08b","arxiv_id":"2502.02456","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A parameter-free computational model of learning predicted the main effects of two human A/B tutoring experiments, a first for the Model Human Learner concept.","lead":"A computer model that simulates how people learn predicted the key outcomes of two real experiments on math tutoring, without ever seeing the human results. This suggests instructional designers could someday test teaching strategies in simulation before running expensive human trials.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Study 2's claimed prediction is not fully specified: the hard-problem subset and half-pretraining are post hoc, and the paper never reports whether the constrained advantage survives without them, so the central 'first evidence' claim is not yet established.","rationale":"The paper is a serious proof-of-concept and the two demonstrations are real: in Study 1, a zero-knowledge agent reproduces the interleaving advantage on the posttest; in Study 2, agents reproduce the constrained advantage on hard problems at strikingly similar absolute levels. The reader's conditional verdict is appropriate. My stress-test focus is narrower than proxy faithfulness. The most load-bearing premise of the conclusion is not 'Trestle is exactly like a human' but 'Trestle's predictions are genuine predictions of the experimental manipulation.' In Study 2 this premise is currently unverifiable because the analysis contains two post hoc decisions—restricting to hard problems and pre-training half the agents—whose separate effects are never reported. The sentence 'the agents all have identical initial conditions' is inconsistent with the footnote that half were pretrained. Because the paper does not report the easy-problem results or the no-pretraining results, a reader cannot tell whether the constrained advantage is a property of the task structure or an artifact of those choices. This does not mean the result is fake; it means the paper has not yet demonstrated the central claim at the strength claimed. The concrete test settles it by requiring all four analysis variants. If the effect is robust, the condition for the claim is satisfied and the conclusion stands. If not, the conclusion should be weakened to a conditional demonstration. Hence I keep the reader's CONDITIONAL verdict; no further adjustment needed.","tokens_in":1080,"tokens_out":1105,"duration_ms":62285,"concrete_test":"Re-run the Study 2 simulations under four fully crossed conditions: (a) analyze all 32 problems, (b) analyze only the 16 hard problems, (c) no easy-problem pretraining for any agent, and (d) easy-problem pretraining for all agents (not half). Report the condition odds ratio and accuracy for each combination, e.g., using the same mixed-effect logistic regression as in the paper. If the constrained advantage is robust across all four combinations and the human hard-problem effect is matched, the central claim survives. If it appears only in the hard-only, pretraining-mixed analysis actually reported, then the claimed prediction is an artifact of post hoc selection, and the conclusion should be reworded to 'the model can reproduce the effect on hard problems under the reported settings.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that Trestle 'successfully predict[s] the main effects observed in multiple human A/B experiments' and that these are parameter-free predictions 'based entirely on the structure of the task and the sequence of the items' (Conclusion; Study 2 Discussion). Study 2 is the more quantitative of the two validations, but its reported prediction is the product of two post hoc choices. The Simulation and Analysis Method section states that agents were trained on all 32 problems yet 'only analyzed their performance on the hard problems ... because my model does not take into account prior knowledge,' and Footnote 4 adds: 'Half of the simulated agents received prior training on 16 easy problems.' The paper never reports model accuracy for the easy problems, never compares the no-pretraining half with the pretrained half, and does not state whether the constrained-vs-unconstrained effect is present in both halves. It also says the agent regression had no random effect 'because the agents all have identical initial conditions,' which cannot be literally true if half were pretrained. Since the hard-problem filter and the pretraining were introduced with knowledge of the human results, the model's 'prediction' is not yet shown to be a genuine out-of-sample prediction; it could depend on the selected subset or the pretraining split. Study 1 further concedes that agreement is only qualitative (Discussion: 'the model only qualitatively predicts the experimental effects'), so the abstract's word 'accurately' is not supported. Together these points mean the load-bearing assertion—first successful Model Human Learner demonstration—is plausible but unverified as stated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the concept of a Model Human Learner, a mechanistic computational model of learning intended to let instructional designers pre-evaluate A/B experiments before running human studies. The model, Trestle from the Apprentice Learner Architecture, is used to simulate two published human experiments: a fraction arithmetic tutor experiment comparing blocked vs. interleaved practice (Study 1) and a Box-and-Arrows tutor experiment comparing constrained vs. unconstrained problem designs (Study 2). The paper reports that the simulated agents reproduce the main effects of both experiments: interleaving reduces tutor performance but improves posttest performance in Study 1, and constrained problems produce higher accuracy than unconstrained problems in Study 2. The paper also claims to generate parameter-free learning curves that qualitatively match human trends, and it uses the Box-and-Arrows simulation to argue against Lee et al.'s (2015) explanation that constrained problems help because they make the correct procedure computationally easier.","tokens_in":9602,"tokens_out":4286,"duration_ms":45074,"significance":"If the central claim is sustained, the paper would be an important proof of concept: it would show that a mechanistic, theory-driven simulation can anticipate the direction of instructional-intervention effects without being fitted to the target human data. The strengths of the paper are real: the simulations produce statistically significant odds ratios in the same direction as the human effects in two independent datasets; the learning curves are not fitted to the target data in the usual curve-fitting sense; and the theoretical observation about procedural ambiguity is a useful, testable alternative to Lee et al.'s computational-cost hypothesis. However, the paper's headline claim of 'first evidence' that a model can 'successfully predict' multiple human A/B experiments is currently stronger than the evidence supports, because in Study 2 the prediction depends on post hoc data selection and an ad hoc pretraining split, and in Study 1 the author concedes that agreement is only qualitative.","major_comments":[{"comment":"The claimed prediction for the Box-and-Arrows experiment is not fully specified. The analysis is restricted to hard problems post hoc ('only analyzed their performance on the hard problems, because my model does not take into account prior knowledge'), and half of the simulated agents received pretraining on 16 easy problems (Footnote 4). The paper never reports the model's accuracy on easy problems, never compares the pretrained half with the non-pretrained half, and never states whether the constrained-vs-unconstrained effect survives in both halves. Because both choices were made with knowledge of the human results and the desired outcome, the reported prediction could depend on these choices. To support the 'first evidence' claim, the paper must report the full-problem analysis and the no-pretraining analysis, or justify both choices from theory in advance.","section":"Study 2, Simulation and Analysis Method (incl. Footnote 4)"},{"comment":"The statement that no random effect was included for agents 'because the agents all have identical initial conditions' is internally inconsistent with Footnote 4, which states that half of the simulated agents received prior training on easy problems. If half were pretrained, the agents do not all have identical initial conditions, and the regression should account for this grouping or the pretraining should be removed. At minimum, the author should clarify whether the reported odds ratio comes from the pooled agent sample, the pretrained half, the non-pretrained half, or some combination.","section":"Study 2, Simulation and Analysis Method"},{"comment":"The paper's own Discussion states that the model 'only qualitatively predicts the experimental effects' and 'does not accurately predict the absolute tutor and posttest scores for each student (or their average).' Yet the Abstract says the model can 'accurately predict the outcomes' of the two experiments, and the Conclusion claims 'first evidence' of successful prediction of main effects in multiple A/B experiments. The main-effect directions match, which is nontrivial, but the word 'accurately' is not supported by Study 1, where the agents start with no prior knowledge and have 100% first-problem error versus 53.8% for humans. The claims should be scaled back to 'qualitatively predict the direction of the main effects' unless the model is extended to handle prior knowledge.","section":"Study 1, Discussion and Abstract/Conclusion"},{"comment":"The paper calls the learning curves 'parameter-free predictions based solely on task structure' and later says they 'could have been generated prior to collecting any human data.' However, in Study 1 the agents are given the exact problem sequence each human student received, and in Study 2 the same simulation procedure is used. If those sequences are taken from the human logs, then the predictions are not independent of human data in the way the text suggests. The paper should clarify whether the problem sequences were specified a priori by the experimental design or copied from the human student logs, and the same clarification is needed for the claim that the Box-and-Arrows predictions were generated 'based entirely on the structure of the task and the sequence of the items.'","section":"Study 1 and Study 2, Learning Curves and Prediction Independence"},{"comment":"The 'parameter-free' claim cannot be verified from the text because the model's parameters, if any, are never enumerated. Trestle is described as having utility values, condition refinement, and generalization mechanisms, but the paper does not state which numerical parameters these mechanisms contain, what values were used, or whether any values were tuned on prior data. If the model has no parameters, that should be stated explicitly; if it has parameters, the paper must list them and explain how their values were set, because the claim that predictions are not fitted to human data is load-bearing for the entire paper.","section":"The Computational Model"}],"minor_comments":[{"comment":"The phrase 'how different instructional choices effect student learning' should be 'how different instructional choices affect student learning.'","section":"Study 2, Human Data"},{"comment":"There is a duplicated word in 'because all problems were of of the same type (hard).'","section":"Study 2, Simulation and Analysis Method"},{"comment":"The manuscript uses 'ANOV A' with a stray space where 'ANOVA' is intended.","section":"Throughout"},{"comment":"The pretraining detail is important enough that it should be in the main text, not relegated to a footnote, especially since the analysis depends on it.","section":"Study 2, Footnote 4"},{"comment":"The reference 'Card, S. K., Moran, T. P., & Newell, A. (1986)' lacks publisher and page range information, and the title 'The handbook of human perception' should be italicized.","section":"References"},{"comment":"The captions should state explicitly which panel is the human interface and which is the machine-readable tutor interface, since the text relies on this distinction.","section":"Figure 1 and Figure 4"}],"recommendation":"major_revision","confidential_remarks":"The paper is a single-author proof of concept built substantially on the author's own prior architecture, which is appropriate given the dissertation-based lineage. The main concern is not novelty or self-citation per se, but that the headline claim of 'first evidence' exceeds what the current analyses establish; the load-bearing post hoc choices in Study 2 and the qualitative-only agreement in Study 1 need to be either fixed or explicitly delimited. The journal's scope is suitable, and the concept is worth publishing once the prediction claim is made commensurate with the evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read. The paper does something real: it takes a mechanistic, parameter-free model of skill acquisition and shows that it reproduces the direction of two published human A/B experiments—blocked vs interleaved fraction practice and constrained vs unconstrained box-and-arrows problems. That had not been shown before for this class of model, and the learning curves are genuinely emergent rather than fitted. The surprise result on the box-and-arrows task—agents don't care about computational ease, yet still show the constrained advantage—is a legitimate challenge to Lee et al.'s account, and the procedural-ambiguity explanation is worth taking seriously. The author also uses real datasets and reports regression details, which is more than most proof-of-concept papers do.\n\nThe soft spots are exactly where the reader put them. Study 1's agreement is only qualitative; the model has zero prior knowledge, so its absolute error rates are far from human. The abstract's word 'accurately' is not supported by the body, which says 'qualitatively' in the discussion. That needs to change. Study 2's prediction is weaker than the text claims. The hard-problem subset and the half-pretraining were chosen after seeing the human results, and the paper never reports whether the constrained advantage holds without those choices. The regression note that agents have identical initial conditions is also inconsistent with half being pretrained. These are fixable by reporting robustness checks, but until then the 'first successful demonstration' language outruns the evidence.\n\nThe self-citation and the fact that the model instantiates Lee et al.'s broader theory are not problems by themselves. The target predictions are emergent, and prior work did not validate against human A/B results. The circularity burden is low.\n\nBottom line: this is a solid proof-of-concept, not a paradigm shift. It deserves a serious referee and likely publication after revision. I'd require the author to (1) align abstract claims with qualitative agreement, (2) show the Study 2 effect across the full item set and both pretraining halves, and (3) ship code or scripts. With that, it becomes a useful reference point for anyone doing simulation-based instructional design.","headline":"Genuine proof-of-concept with two real validations, but the 'accurate' claims outrun what the analyses show.","tokens_in":10151,"tokens_out":2164,"would_cite":true,"duration_ms":20992,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A computational model of learning reproduces the main effects of two human A/B teaching experiments without training on human data.","keywords":["Model Human Learner","instructional design","computational models of learning","skill acquisition","A/B experiments","tutoring systems","parameter-free prediction","learning curves"],"falsifier":"A concrete test would be to preregister the model's predictions for a third instructional manipulation, for instance an item-design experiment that varies only the number of candidate procedures consistent with the correct answer, and then run the human study; if the human main effect is absent or reversed, the paper's claim that the model can act as a Model Human Learner is falsified.","tokens_in":9112,"feed_emoji":"🧠","tokens_out":11221,"duration_ms":99414,"temperature":0.7,"pith_summary":"The paper proposes a 'Model Human Learner': a computational model of skill acquisition that instructional designers could run to evaluate teaching interventions before spending money on human studies. It presents what the author identifies as the first evidence that such a model can predict real experimental outcomes, using two existing human A/B experiments as tests. In the fraction-arithmetic experiment, the model reproduced the human result that interleaved practice hurts training performance but improves posttest performance; in the box-and-arrows experiment, it reproduced the human result that constrained item designs aid rule learning. In both cases the predictions were generated without fitting parameters to human data, and the model also produced learning curves that tracked human trends. The author is explicit that the match is qualitative in the fraction study because the simulated students start with zero prior knowledge.","feed_headline":"Simulated learners match humans in two tutoring A/B tests","feed_subtitle":"Task-driven, parameter-free simulations reproduce main effects and trace them to procedural ambiguity, not math ease.","key_machinery":"The central machinery is Trestle, the paper's computational model of skill acquisition. When confronted with a problem, the model matches learned skills against the current state and executes the highest-utility match; when no skill matches, it requests a demonstration, searches for a sequence of basic arithmetic operations that explains the demonstration, generalizes that procedure into a reusable skill, and then refines the skill's conditions and utility from correctness feedback. This cycle converts task structure and problem order directly into simulated learning curves, which is what lets the model generate predictions before any human data are collected.","core_discovery":"The paper's central claim is that a computational model that learns skills from demonstrations and correctness feedback, run through the same problems a human student would receive, reproduces the direction and statistical significance of the main effects in two human A/B experiments. In the blocked-versus-interleaved fraction study, both humans and agents show worse training accuracy and better posttest accuracy under interleaving. In the constrained-versus-unconstrained box-and-arrows study, both humans and agents do better with constrained problems, and when prior knowledge is controlled the agent accuracies are close to the human ones (19.6% versus 19.1% constrained; 10.0% versus 8.8% unconstrained). The predictions come from task structure and item order alone, with no fitted parameters, which the paper offers as evidence that the model captures something about human learning rather than merely fitting outcomes. The paper also uses the model to test why constrained items help, concluding that the benefit comes from reduced procedural ambiguity rather than from whole numbers being easier to compute.","pith_inferences":["If this approach generalizes, instructional design could shift from running many human experiments to first comparing the structure of the candidate rule spaces that different designs create.","The model's zero-prior-knowledge limitation suggests that a practical version will need a way to estimate incoming student competence without using the target experiment's data.","The same simulation protocol could be extended to rule-learning tasks outside arithmetic, such as programming or science tutors, where the space of candidate procedures can be enumerated.","The paper's proposed mechanism for constrained items implies a directly testable design rule: hold arithmetic difficulty constant and vary only the number of candidate procedures, and human learning should follow ambiguity."],"forward_implications":["Instructional designers could use the model to screen candidate interventions in simulation, reserving human A/B tests for designs the model identifies as promising.","Because predictions are parameter-free and depend only on task structure and item order, the model can generate learning-curve forecasts for tasks where no student data exist yet.","The reproduced effects support the rule-search theory of learning on these tasks, since the model instantiates that theory and produces the same main effects as humans.","The results point to a concrete design principle: constrain items so that only one procedure is consistent with the correct answer, rather than making computation easier.","A validated Model Human Learner would give researchers a low-cost way to test competing explanations for why an intervention works."],"supporting_citations":[{"why":"Supplies the fraction-arithmetic tutor dataset and the blocked-versus-interleaved human experiment whose main effects the model is asked to predict.","marker":"(Patel et al., 2016)"},{"why":"Supplies the box-and-arrows tutor dataset, the constrained-versus-unconstrained manipulation, and the rule-search learning theory that the model instantiates and then challenges.","marker":"(Lee et al., 2015)"},{"why":"Provides the full specification of the Trestle skill-acquisition model used in the simulations.","marker":"(MacLellan & Koedinger, 2022)"},{"why":"Establishes the near-zero-parameter prediction approach that lets the model generate learning curves without fitting human data.","marker":"(Weitekamp et al., 2019)"},{"why":"Describes the computational learning models and architecture from which this study's simulations are drawn.","marker":"(MacLellan, 2017)"},{"why":"Provides the authoring tools used to create the machine-readable tutors through which the simulated agents interact.","marker":"(Aleven, McLaren, Sewall, & Koedinger, 2006)"},{"why":"Provides the data repository through which both human datasets were accessed for the study.","marker":"(Koedinger et al., 2010)"}],"fun_headline_variants":["Parameter-free model predicts two human tutoring A/B tests","Simulated learner matches human results without any tuning","Model explains why constrained practice beats unconstrained"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the simulated learner's skill-acquisition mechanisms are a faithful enough stand-in for human learning that the condition effects seen in simulation carry over to real students; if that proxy fails, the predicted outcomes do not transfer.","fun_headline_variants_meta":{"raw":{"variants":["Parameter-free model predicts two human tutoring A/B tests","Simulated learner matches human results without any tuning","Model explains why constrained practice beats unconstrained"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000348,"raw_usage":{"total_tokens":1862,"prompt_tokens":863,"completion_tokens":999,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":479,"completion_tokens_details":{"reasoning_tokens":952}},"tokens_in":479,"tokens_out":999,"duration_ms":11127,"temperature":1.0,"reasoning_tokens":952,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T12:04:35.190961+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test would be to preregister the model's predictions for a third instructional manipulation, for instance an item-design experiment that varies only the number of candidate procedures consistent with the correct answer, and then run the human study; if the human main effect is absent or reversed, the paper's claim that the model can act as a Model Human Learner is falsified.","supporting_citations":[{"cited_title":", Liu, R","cited_arxiv_id":null,"evidence_quote":"Supplies the fraction-arithmetic tutor dataset and the blocked-versus-interleaved human experiment whose main effects the model is asked to predict."},{"cited_title":", Betts, S","cited_arxiv_id":null,"evidence_quote":"Supplies the box-and-arrows tutor dataset, the constrained-versus-unconstrained manipulation, and the rule-search learning theory that the model instantiates and then challenges."},{"cited_title":"\\ Koedinger, K R","cited_arxiv_id":null,"evidence_quote":"Provides the full specification of the Trestle skill-acquisition model used in the simulations."},{"cited_title":", Harpstead, E","cited_arxiv_id":null,"evidence_quote":"Establishes the near-zero-parameter prediction approach that lets the model generate learning curves without fitting human data."},{"cited_title":"APACrefauthors \\ 2017","cited_arxiv_id":null,"evidence_quote":"Describes the computational learning models and architecture from which this study's simulations are drawn."},{"cited_title":", McLaren, B M","cited_arxiv_id":null,"evidence_quote":"Provides the authoring tools used to create the machine-readable tutors through which the simulated agents interact."},{"cited_title":", Baker, R S J d","cited_arxiv_id":null,"evidence_quote":"Provides the data repository through which both human datasets were accessed for the study."}],"review_version":1}