{"id":"8f45626e-2bdd-419d-96b4-e981c3fc2cf9","arxiv_id":"2508.08292","paper_version":2,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"A contamination-resistant Putnam math benchmark shows large accuracy drops on programmatically varied unseen problems, suggesting LLMs memorize rather than reason.","lead":"This paper introduces Putnam-AXIOM, a 522 problem math benchmark built from Putnam competition problems, plus 100 unseen 'variation' problems created by changing variables and constants. The strongest model, o1-preview, drops from roughly 42% on the originals to about 22% on the variations, suggesting that LLMs partially memorize benchmark answers.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The equal-difficulty premise for the Variation set is unvalidated; without human or independent difficulty calibration, the observed performance drop cannot be attributed to memorization rather than to distribution shift.","rationale":"The benchmark is credible and the public release of data and code is a real strength. The central inference—that the accuracy drop is caused by memorization—is neither internally inconsistent nor obviously false, but it rests on the assumption that the Variation protocol preserves difficulty. This assumption is empirically testable and, if it fails, the headline conclusion loses its support. The abstract-only review cannot resolve this, so the conditional verdict remains appropriate. No additional concern is more load-bearing than this one.","tokens_in":770,"tokens_out":2174,"duration_ms":28872,"concrete_test":"Run a blinded human calibration study: give a matched group of strong mathematical problem-solvers (e.g., Putnam participants or graduate students) both the Original and Variation items in randomized order, with identical time limits and scoring, and compare accuracy and solution time. Pre-specify equivalence bounds (e.g., accuracy within ±5 percentage points and time within ±20%). If humans perform equivalently on both sets, the equal-difficulty premise is supported and the memorization reading is credible. If human performance drops on Variations, the observed model gap is confounded by difficulty, and the paper must recalibrate its conclusion.","verdict_should_be":"UNCHANGED","load_bearing_attack":"To infer memorization from the Original-to-Variation drop, the variations must be 'equally difficult' in every way except the model's prior exposure. The abstract asserts this but does not provide evidence. Programmatic perturbation of variables and constants is not difficulty-neutral in general: changing values can affect the need for case analysis, the availability of simplifying identities, the size of intermediate arithmetic, the uniqueness of solutions, and the probability that a shortcut works. The variation set is also smaller (100 vs. 522) and generated by a protocol not described in the abstract, so there is no check that the paired problems sample the same difficulty distribution or that the drop is not driven by format changes in the generated prompts. Without a human baseline or an independent difficulty proxy, the 19.6% gap is consistent with both memorization and a difficulty shift. This is the load-bearing weakness: the paper's strongest claim—'these gaps suggest memorization'—depends entirely on an unvalidated equivalence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces Putnam-AXIOM, a benchmark of 522 Putnam Competition problems, and a companion set of 100 programmatically perturbed 'Variation' problems. The central claim is that model accuracy drops sharply on the paired Variations relative to the Originals (o1-preview: 41.9% to 22.3%), and that this gap 'suggests memorization' and demonstrates the need for dynamic, contamination-resilient benchmarks. The abstract also proposes Teacher-Forced Accuracy (TFA), a metric for directly scoring reasoning traces. This review is based on the abstract only, as the full text was not available.","tokens_in":1095,"tokens_out":2074,"duration_ms":29810,"significance":"If the central claim were validated, the benchmark would be a valuable contribution: it is a public, functional-variation benchmark aimed at a genuine problem (training-set contamination), and the TFA metric could be useful for automating proof evaluation. The abstract reports concrete numbers with confidence intervals, and the code/data are promised publicly. However, the significance hinges entirely on the unverified premise that the variations are 'equally difficult' to the originals. Without a human baseline or independent difficulty calibration, the observed performance drop is equally consistent with a difficulty shift induced by the perturbation protocol. The contribution is therefore potentially significant but currently under-evidenced.","major_comments":[{"comment":"The abstract asserts that the variation protocol produces 'equally difficult, unseen instances' and that the accuracy drop on the Variations 'suggests memorization,' but no evidence is provided that the perturbations preserve difficulty. Changing variables and constants can affect the need for case analysis, the size of intermediate arithmetic, the availability of simplifying identities, and the probability that a shortcut succeeds. The 19.6% drop for o1-preview and the similar trend across other models are consistent with a difficulty shift or a format change, not necessarily memorization. To support the memorization interpretation, the paper must validate difficulty equivalence, for example by showing that human solvers achieve comparable accuracy or solving time on both sets, or by matching items on an independent difficulty proxy.","section":"Abstract, 'equally difficult' claim"},{"comment":"The abstract says the benchmark is 'contamination-resilient' and that the performance gap 'suggests memorization,' but no contamination test is described. A direct test would be needed, such as checking n-gram overlap between the training data and the Variation prompts, probing the model with near-duplicates, or comparing behavior on semantically identical problems with different surface forms. Absent such a test, the drop could reflect a distribution shift unrelated to memorization, such as changed notation, prompt formatting, or problem presentation. The claim that the variation protocol 'removes' memorization while preserving reasoning difficulty is the load-bearing premise and cannot be inferred from the performance gap alone.","section":"Abstract, contamination-resilience claim"},{"comment":"The abstract introduces Teacher-Forced Accuracy (TFA) as a metric that 'directly scores reasoning traces and automates natural language proof evaluations,' but gives no description of how the metric works, how it handles partial credit, or how it was validated. Because TFA is presented as a companion contribution, the paper should specify the exact scoring mechanism and compare it against human evaluation or other proof-grading metrics; otherwise it is difficult to assess whether TFA adds value beyond boxed-answer accuracy.","section":"Abstract, TFA metric"}],"minor_comments":[{"comment":"The phrase 'These gaps suggest memorization' is too assertive for an observational performance gap; it should be qualified as 'are consistent with memorization' or similar, given the alternative explanations discussed above.","section":"Abstract, wording"},{"comment":"The abstract reports 522 original problems and 100 variations; the large size difference is not discussed, and it is unclear whether the paired comparison uses all 522 originals or only the 100 that were varied. This should be clarified.","section":"Abstract, dataset sizes"},{"comment":"The phrase 'an unlimited stream of equally difficult, unseen instances' promises a property that cannot be established by finite sampling; the paper should explain how the difficulty preservation guarantee is obtained, rather than asserted.","section":"Abstract, 'unlimited stream'"}],"recommendation":"major_revision","confidential_remarks":"This review is based on the abstract only; the full text might contain additional validation. The main concern is that the paper's headline claim—memorization inferred from an Original-to-Variation performance drop—depends on an unvalidated equal-difficulty premise. This is fixable in revision by adding human baselines, independent difficulty calibration, and explicit contamination tests. If those are present in the full text, the paper could be acceptable; as presented, the evidence is insufficient for the strong conclusion drawn."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this paper gives the community a genuinely useful resource—a 522-problem Putnam benchmark, a 100-problem programmatic variation set, and a lightweight trace-scoring metric (TFA), all with public code and data. The empirical observation is also interesting: o1-preview drops from 41.9% to 22.3% on variants, with most other models showing a similar downward trend. That pattern is worth taking seriously.\n\nThe soft spot is exactly where the stress-test note lands. The abstract's claim that the variations are 'equally difficult' is asserted, not demonstrated. The gap between original and variation could reflect memorization, but it could also reflect a distribution shift—changes in constants or variables can make problems harder (more casework, messier arithmetic, less convenient identities) even for a model with zero prior exposure. The variation set is also smaller (100 vs. 522) and the generation protocol isn't described in the abstract, so we don't know whether the paired problems sample the same difficulty distribution. A human baseline, or an independent difficulty proxy, would settle this; without it, the confidence intervals support the drop but not the memorization interpretation.\n\nThat said, I'm not treating this as a fatal objection. The paper is a benchmark paper, and the benchmark itself—the original Putnam set plus the variation protocol—is valuable regardless of whether the memorization claim survives. The TFA metric is a nice addition, though I'd want to see it validated against human grading. The public code and data are a plus.\n\nOne caveat: I've seen only the abstract, so the full text may already contain difficulty calibration or contamination checks. If it does, the main empirical claim is much stronger. If it doesn't, that's the section that needs the most work.\n\nWho it's for: anyone building or evaluating math LLMs, especially people worried about contamination. It deserves a serious referee; the question is whether the referee should focus on the variation protocol's equivalence rather than the benchmark construction. I'd send it out.","headline":"A useful new Putnam benchmark and variation protocol, but the memorization claim rides on an unvalidated equal-difficulty assumption.","tokens_in":1431,"tokens_out":1790,"would_cite":true,"duration_ms":20071,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a programmatic variation of Putnam problems cuts o1-preview's accuracy from 41.9% to 22.3%, suggesting that high benchmark scores largely reflect memorization rather than mathematical reasoning.","keywords":["mathematical reasoning","LLM evaluation","benchmark contamination","Putnam competition","dynamic benchmark","programmatic variation","teacher-forced accuracy","reasoning trace"],"falsifier":"Have human Putnam contestants solve the 100 variations under exam conditions and compare their accuracy with the originals; a similar human drop would indicate a difficulty shift rather than memorization.","tokens_in":605,"feed_emoji":"📉","tokens_out":4999,"duration_ms":55330,"temperature":0.7,"pith_summary":"The paper claims that high scores on existing math benchmarks overstate how well large language models actually reason, because the models may have memorized the problems. To test this, the authors built Putnam-AXIOM: 522 problems from the William Lowell Putnam Competition, plus 100 variation problems made by programmatically swapping variables and constants. On the original set, the strongest evaluated model, o1-preview, scores 41.9%; on the paired variations it falls to 22.3%, a 19.6-percentage-point (46.8% relative) drop. The paper argues this gap is evidence of memorization and that dynamic benchmarks like this one are needed; it also introduces Teacher-Forced Accuracy, a metric that scores the reasoning trace rather than just the final boxed answer.","feed_headline":"Top LLM's Putnam math score falls 47% when variables change","feed_subtitle":"o1-preview drops from 41.9% to 22.3% on perturbed variants — a sign of memorized answers.","key_machinery":"The load-bearing mechanism is the Variation protocol: a programmatic procedure that rewrites a competition problem by changing its variables, constants, and other surface elements while preserving its mathematical difficulty, generating an unlimited stream of fresh instances that are unlikely to have appeared in training data. Paired with the original problems, these variations isolate memorization from reasoning: if a model's accuracy collapses when only surface details change, the original score was probably memorized. The second piece of machinery is Teacher-Forced Accuracy (TFA), a metric that scores the model's generated reasoning trace token-by-token against a reference trace, extending evaluation beyond whether the final boxed answer is correct.","core_discovery":"Putnam-AXIOM is a benchmark of 522 original Putnam problems; Putnam-AXIOM Variation consists of 100 paired instances generated by programmatically perturbing variables and constants. The central empirical finding is that accuracy on the paired variations is far below accuracy on the originals: o1-preview scores 41.9% on the originals and 22.3% on the variations, a 19.6-point drop (46.8% relative), and all nineteen evaluated models show the same downward trend, with ten having non-overlapping 95% confidence intervals. The authors interpret this gap as suggesting memorization of original problems rather than genuine reasoning, and present the variation protocol as a contamination-resilient evaluation method producing an unlimited stream of equally difficult, unseen instances. They also propose Teacher-Forced Accuracy (TFA), a lightweight metric that evaluates the model's reasoning trace directly, as a complement to final-answer accuracy.","pith_inferences":["If the equal-difficulty assumption holds, the same protocol could serve as a contamination probe for any competition-style benchmark: a large originals-minus-variations gap becomes a measurable memorization index, and publishers could set a threshold above which a model's official score is suspect.","The variation approach could transfer beyond math, for example to code benchmarks (renaming variables, reordering functions) or scientific reasoning, wherever fixed problem sets are at risk of contamination.","A human baseline on the variations would be the most direct check: if human Putnam solvers also drop by a similar amount, the gap is difficulty, not memorization; if they do not, the paper's interpretation is strongly supported.","Teacher-Forced Accuracy inherits the usual risks of reference-based scoring: a model could produce a shorter or differently worded but correct proof and be penalized, so it may need semantic equivalence checks to serve as a standalone metric."],"forward_implications":["Scores on fixed math benchmarks should be read as upper bounds on reasoning ability, since even strong models lose nearly half their accuracy on perturbed variants.","Programmatic variation is a viable, contamination-resilient benchmark design that can generate unlimited test instances without human annotation.","Teacher-Forced Accuracy offers a way to evaluate reasoning traces automatically, not just final answers, making proof-style evaluation scalable.","The 19.6-point gap on o1-preview is a concrete measure of the performance the paper attributes to memorization on this benchmark.","Future model evaluations should include both originals and variations to distinguish reasoning improvements from recall of fixed problems."],"supporting_citations":[],"fun_headline_variants":["o1-preview Putnam score drops 47% when variables are altered","Perturbed Putnam problems cut o1-preview accuracy by 47%","Memorization exposed: LLM math skills fail on modified Putnam","All 19 LLMs do worse on perturbed Putnam problems","Putnam-AXIOM: fresh math problems reveal LLM memorization"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim stands on the assumption that perturbing variables and constants yields problems equally as hard as the originals, so that the accuracy drop isolates memorization rather than a change in difficulty or format.","fun_headline_variants_meta":{"raw":{"variants":["o1-preview Putnam score drops 47% when variables are altered","Perturbed Putnam problems cut o1-preview accuracy by 47%","Memorization exposed: LLM math skills fail on modified Putnam","All 19 LLMs do worse on perturbed Putnam problems","Putnam-AXIOM: fresh math problems reveal LLM memorization"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000623,"raw_usage":{"total_tokens":2906,"prompt_tokens":988,"completion_tokens":1918,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":604,"completion_tokens_details":{"reasoning_tokens":1824}},"tokens_in":604,"tokens_out":1918,"duration_ms":14527,"temperature":1.0,"reasoning_tokens":1824,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T04:11:55.456645+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have human Putnam contestants solve the 100 variations under exam conditions and compare their accuracy with the originals; a similar human drop would indicate a difficulty shift rather than memorization.","supporting_citations":[],"review_version":1}