{"id":"a339171f-6730-4b09-aa5f-41da32e92d73","arxiv_id":"2505.14148","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"MM-Agent, a multi-stage LLM pipeline with a hierarchical modeling method library, is claimed to outperform prior agents and award-winning human solutions on a new 111-problem MCM/ICM-based mathematical modeling benchmark.","lead":"The paper introduces MM-Bench, a 111-problem benchmark built from MCM/ICM contests, and MM-Agent, a four-stage LLM agent that analyzes, models, solves, and writes up mathematical modeling problems. It reports that MM-Agent outscores award-winning human solutions on its own benchmark and helped two teams earn a Finalist award in 2025, but the headline comparison relies on a subjective LLM judge and a small test subset.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 11.88% human-outperformance claim rests on a same-family LLM judge; a blinded re-scoring by independent human experts and another LLM family is needed before the headline can be credited.","rationale":"The reader's weakest assumption—the validity of the four-criteria rubric as implemented by GPT-4o and a small human panel—is indeed the load-bearing point. The paper's headline is a numerical comparison against human experts, and every number in that comparison comes from the same uncalibrated LLM judge. The appendix's own disclosure of low inter-annotator agreement (human–human MR = 0.4813) and the agent/judge same-family overlap make the reported 0.94-point gap insufficiently secured. I agree with the CONDITIONAL verdict: the framework and benchmark are useful, the code is released, the OPTIBENCH experiment provides a small ground-truth check, and the 2025 contest outcomes are suggestive real-world evidence, but the central outperformance claim should not be accepted as stated without an independent re-scoring study. My proposed concrete test—blinded re-scoring by both independent human experts and a different LLM family—directly targets this concern. I do not see a separate internal inconsistency that would change the verdict; the subset-selection ambiguity (32 of 111 problems, selection criteria not specified) is real but secondary, and the evaluation concern already subsumes it. Hence the reader's verdict stands unchanged pending that check.","tokens_in":46178,"tokens_out":3529,"duration_ms":38973,"concrete_test":"Select a blinded, stratified random sample of 10–16 problems from the 32 in Table 1. For each, take one MM-Agent report and one award-winning human team report, strip identifying headers, and have (a) two independent MCM/ICM judges with Finalist/Outstanding experience and (b) a different LLM family (e.g., Claude or Gemini) score pairs with the paper's four criteria, blinded to authorship. Compare the mean difference (MM-Agent − Human) against the reported 0.94-point gap (8.85 vs. 7.91 on 2021–2024). If the independent judges return a gap ≤0.3 points or negative, the '11.88% improvement over human expert solutions' claim is not confirmed. Report inter-annotator reliability (e.g., Cohen's κ or ICC) on the same sample to gauge rubric stability.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim—that MM-Agent exceeds award-winning human solutions by 11.88%—is computed from Table 1, whose scores come from the GPT-4o-based automatic rubric described in §4.1 and Appendix D. The judge prompt explicitly rewards 'innovation', 'goes beyond standard machine learning', and 'novel frameworks' (Figure 12, criteria 3.1/3.2). Because MM-Agent's backbone is GPT-4o and the judge is also GPT-4o, there is a concrete risk of systematic stylistic preference for agent-generated reports, independent of modeling quality. This is not a hypothetical: Appendix E.4 reports model–human agreement of only 0.5068 on AE and 0.5692 on RBA, and human–human agreement on MR is 0.4813, showing the rubric is noisy. A systematic judge bias of even 0.5 points on a 10-point scale would erase the reported gap (8.85 vs. 7.91). The 'human expert solutions' are also not official MCM scores but reports re-scored by this same GPT-4o judge, so the headline gain is entirely constructed from the contested evaluator. The competitive Finalist awards and the OPTIBENCH ground-truth results provide partial independent support, but neither establishes that MM-Agent beats human experts on open-ended modeling; the OPTIBENCH gain is only ~2% accuracy on well-posed optimization tasks. Thus the evaluation rubric is the load-bearing premise, and the paper's own appendix concedes judge-bias risk without providing a countermeasure.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper formalizes LLM-based real-world mathematical modeling as an agentic task, introduces MM-Bench (111 MCM/ICM problems from 2000 to 2025 across 10 domains), and proposes MM-Agent, a four-stage pipeline (problem analysis, mathematical modeling with hierarchical retrieval and actor-critic refinement, computational solving, and report generation) built around a tri-level Hierarchical Mathematical Modeling Library (HMML). On a 32-problem subset of MM-Bench, the paper reports that MM-Agent outperforms baseline agents and award-winning human solutions by 11.88% in overall score under a GPT-4o-based rubric, at roughly 15 minutes and $0.88 per task with GPT-4o, with similar results under DeepSeek-R1. The system also assisted two undergraduate teams in achieving the Finalist Award (top 2.0%) in MCM/ICM 2025. Additional experiments include ablations, cost and runtime analysis, human expert evaluation, and zero-shot results on the ground-truth OPTIBENCH dataset for well-posed optimization problems.","tokens_in":46531,"tokens_out":5391,"duration_ms":46555,"significance":"If the claims held up, this would be a notable advance: the first systematic benchmark for open-ended LLM mathematical modeling, a reusable agent framework with a structured modeling library, and evidence that an autonomous pipeline can approach competition-level modeling. The public code and demo, the use of real competition problems with manual verification of extracted elements, and the ground-truth OPTIBENCH experiments are concrete strengths. However, the headline result, that MM-Agent significantly outperforms award-winning human teams, depends entirely on a same-family LLM judge: GPT-4o generates the agent's reports and GPT-4o scores them. The paper's own Appendix A concedes judge-bias risk, and Appendix E.4 reports model-human agreement as low as 0.5068 on AE and 0.5692 on RBA, with human-human agreement of only 0.4813 on MR. The human evaluation in Appendix E.3 does not even include the human-team condition. The central claim therefore needs independent, blinded human rescoring and cross-family judge validation before it can be credited; the current evidence supports a more modest claim of strong performance on the proposed benchmark.","major_comments":[{"comment":"The 11.88% improvement over human experts is computed from scores assigned by GPT-4o, the same model family that generates MM-Agent's solutions. The Practicality and Scientificity prompt (Figure 12, criteria 3.1/3.2) explicitly rewards 'innovation', 'goes beyond standard machine learning', and 'novel frameworks', so a systematic stylistic self-preference would directly inflate the reported gap. The paper's own Appendix A states that 'bias still exists in both human and LLM annotators', and Appendix E.4 reports model-human agreement of 0.5068 (AE) and 0.5692 (RBA). Given that the overall gap is only 8.85 vs. 7.91 on a 10-point scale, a judge bias of about 0.5 points would erase the claimed advantage. The manuscript must provide counter-evidence: blinded rescoring of both agent and human reports by contest-experienced human judges, scores from at least one different LLM family, and per-criterion confidence intervals or significance tests.","section":"§4.1, Table 1, Appendix D, Figure 12"},{"comment":"The construction of the 32-problem test set is underspecified: the text only says the authors 'select a subset' from the past five years 'ensuring diversity across problem types and domains', without a sampling protocol, a list of selected problems, or stratification details. This undermines reproducibility and weakens the claimed 'temporal consistency' between the 2021-2024 and 2025 splits. In addition, the 'Human Team' baseline scores are not official MCM/ICM scores but re-scores by the same GPT-4o judge; reporting the official award level for each selected problem would provide an independent calibration anchor.","section":"§4.1"},{"comment":"The human expert evaluation covers only Agent Laboratory, DS-Agent, ResearchAgent, and MM-Agent; it does not include human team solutions. Consequently, the statement that MM-Agent 'significantly outperforms human experts' is not corroborated by human evaluation in any direct way. Either include award-winning human reports in the human evaluation, or restrict the human-outperformance claim to the LLM-judged comparison and clearly say so.","section":"Appendix E.3, Figure 6"},{"comment":"Ablation results are presented only as line charts without numerical values, sample sizes, or error bars. With 32 problems and inter-annotator agreement as low as 0.4813 on MR (Table 6), the observed differences between MM-Agent and its ablated variants cannot be distinguished from annotator noise. Report the actual scores, per-condition standard deviations, and paired significance tests (e.g., bootstrap or Wilcoxon signed-rank) for the ablations and for the main comparisons in Table 1.","section":"§4.3, Figure 4, Table 6"},{"comment":"The Finalist Award result is presented as evidence of practical effectiveness, but MM-Agent acted as a copilot assisting two human teams, not as an autonomous agent. This supports a human-AI collaboration claim, not the autonomous outperformance claim made in the abstract. Please separate these claims explicitly, and if possible report what the teams achieved without the aid of MM-Agent, or state that no controlled comparison was performed.","section":"Abstract, §4.2"}],"minor_comments":[{"comment":"The heading 'Problem Undersanding' contains a typo and should be 'Problem Understanding'.","section":"§3.3.1"},{"comment":"The heading 'Evalaution' contains a typo and should be 'Evaluation'.","section":"Appendix D"},{"comment":"The axis labels contain uninterpretable glyph strings such as '/uni00000024/uni00000028/...'; the figure needs proper axis and legend labels.","section":"Figure 4"},{"comment":"The hyperparameters w, K, n_r, n_c, and tasknum are introduced in §3.3.2 but their experimental values are never reported; please provide them for reproducibility.","section":"§4.3, Appendix D"},{"comment":"The claim that this is the first work on LLMs for real-world mathematical modeling should be qualified in light of the cited optimization-modeling papers (NL4OPT, OPTIMUS, OptiBench), which address a related but narrower formulation; a brief comparison would clarify the novelty.","section":"§2"},{"comment":"The text says 'We observe consistently high agreement in the four metrics', but Table 6 reports human-human agreement of only 0.4813 for MR; this apparent contradiction should be resolved.","section":"Appendix E.4"}],"recommendation":"major_revision","confidential_remarks":"This is a potentially useful benchmark and framework, but the headline claim of outperforming human experts is not yet supported by the evidence. The authors were transparent in Appendix A about judge bias, but a concession does not replace a countermeasure. I would encourage the editor to require an independent, blinded human rescoring of agent and human reports, plus a second judge model family, before the paper is accepted. The scope fit is fine for an ML venue, and I would not reject on novelty grounds."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline result—MM-Agent beating award-winning human solutions by 11.88%—should not be quoted as established. It is computed from GPT-4o judging reports produced by a GPT-4o-based agent, and the judge prompt explicitly rewards innovation and novel frameworks, which biases toward the agent's style. The paper's own appendix concedes judge-bias risk without proposing a countermeasure. So the central quantitative claim is load-bearing on an uncalibrated evaluator.\n\nThat said, this is a solid piece of work in the parts that matter for the community. The formalization of open-ended mathematical modeling as a task, MM-Bench (111 MCM/ICM problems across domains and years), and the HMML hierarchical retrieval library are real artifacts. The four-stage pipeline is a reasonable assembly of established techniques (decomposition, actor-critic, program-aided solving, report generation), and the paper is honest about borrowing them. The ablation study is present and shows each module contributes. The OPTIBENCH results, while modest (~2% accuracy gain over GPT-4o), are at least grounded in ground-truth optimization tasks.\n\nThe soft spots are real but mostly evaluative, not architectural. The test set of 32 problems is underspecified—how were they selected, and what is the per-year/per-problem breakdown? There are no error bars or significance tests, and the human evaluation is thinly reported: one figure, no per-problem numbers, and the agreement table shows human-human agreement on Modeling Rigorousness at 0.4813 and model-human agreement on AE at 0.5068. That is not a reliable yardstick for a 0.94-point overall gap (8.85 vs 7.91). The Finalist awards are genuinely useful external evidence that the system helps humans, but they don't establish that the agent alone beats human experts, since the human teams did the final assembly.\n\nThe cleanest fix is what the stress-test note says: blind re-scoring of a sample of reports by independent human experts and a different LLM family, plus reporting score distributions rather than means. That would either make the claim credible or reveal how much of the gap is judge preference.\n\nWho this is for: anyone working on LLM agents for scientific or engineering tasks, or on evaluation methodology for open-ended problem solving. It deserves a serious referee. I'd recommend accept-with-revisions if the authors add external re-scoring and tighten the test-set description; otherwise the headline claim should be softened.","headline":"Useful benchmark and agent framework, but the 'beats human experts' number rests on a same-family LLM judge and won't survive scrutiny until re-scored independently.","tokens_in":47042,"tokens_out":2000,"would_cite":true,"duration_ms":20417,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An autonomous LLM agent, MM-Agent, can take an open-ended real-world mathematical modeling problem from raw description to finished, competition-grade report, outperforming award-winning human teams by 11.88% on the authors' MM-Bench.","keywords":["LLM agents","mathematical modeling","MM-Bench","hierarchical modeling library","MCM/ICM contests","actor-critic optimization","open-ended problem solving","benchmark evaluation"],"falsifier":"Take the 32-problem MM-Bench test set, remove authorship labels from human and MM-Agent reports, and have a fresh panel of at least three contest-experienced judges score them on the four rubric dimensions. If the blinded overall margin is no longer positive and statistically significant, especially on Modeling Rigorousness where the paper reports human-human agreement of only 0.4813, then the claim of beating award-winning human teams is not supported.","tokens_in":46005,"feed_emoji":"🧮","tokens_out":7500,"duration_ms":69812,"temperature":0.7,"pith_summary":"Mathematical modeling, turning an open-ended real-world situation into assumptions, equations, code, and a written analysis, has resisted automation because the problem formulation itself is open. The paper argues that an LLM-based agent can do this end to end by decomposing the work into four stages: problem analysis, model formulation, computational solving, and report generation. To support the claim, the authors build MM-Bench, a benchmark of 111 problems drawn from the MCM/ICM competitions (2000 to 2025), and evaluate MM-Agent against existing LLM agents and award-winning human solutions. On a 32-problem test set, MM-Agent scores above the human benchmark by 11.88% overall while using roughly $0.88 of API cost and about 15 minutes per task on GPT-4o. If the results hold, an autonomous agent can perform real-world mathematical modeling at expert-competition level, shifting the bottleneck from constructing models to reliably evaluating open-ended work.","feed_headline":"AI agent out-scores human contest winners in math modeling","feed_subtitle":"MM-Agent solves open-ended real-world modeling end to end in 15 minutes for under $1 per problem.","key_machinery":"The central object is the four-stage workflow plus the Hierarchical Mathematical Modeling Library (HMML). HMML is a three-level tree of modeling domains (such as operations research and optimization), subdomains (such as programming theory), and method nodes that each store a method name, its core idea, and typical applications. At retrieval time a depth-first traversal scores each node by the embedding similarity between the subtask and the method, blended with its parent's score, and returns the top-K methods; an actor-critic loop then iteratively proposes, critiques, and revises the modeling scheme. The task coordinator additionally builds a directed dependency graph of subtasks and a memory that passes intermediate models, code, and results between stages. This machinery is what lets the agent abstract an unstructured scenario into a formal model rather than merely solving a pre-formulated problem.","core_discovery":"The central claim is that open-ended mathematical modeling can be reduced to a four-stage expert-inspired pipeline that an LLM can execute autonomously. In the authors' telling, the decisive ingredient is not a bigger model but structure: the agent first analyzes the problem and decomposes it into a dependency graph of subtasks; it then retrieves candidate modeling methods from a hierarchical library (HMML) of 98 schemas organized into domains, subdomains, and method nodes; an actor-critic loop refines the modeling scheme; a code-writing module generates and debugs the computation; and a reporting module compiles a structured LaTeX report. On MM-Bench the pipeline outperforms repurposed data-science and research agents and award-winning human teams on all four evaluation dimensions, with 2021–2024 and 2025 results consistent across two backbones, GPT-4o and DeepSeek-R1-671B. The same system, operated under official contest rules as a copilot, helped two undergraduate teams reach the Finalist tier, the top 2.0% of 27,456 teams, in MCM/ICM 2025.","pith_inferences":["Because the rubric is subjective, with human-human agreement as low as 0.4813 on Modeling Rigorousness, a natural next step is a double-blind study in which fresh contest-experienced judges score human and MM-Agent reports without knowing the author; the size of the reported 11.88% gain under those conditions would reveal how much is modeling quality versus style.","The library-retrieval design is not tied to competition problems; the same abstraction-aware retrieval could be tested on other open-ended engineering or policy tasks where the hard step is choosing a modeling paradigm.","If the four-stage structure matters mostly by compensating for weak base-model reasoning, then as LLMs improve the gap between a raw LLM and MM-Agent should narrow; tracking that gap across model generations would test the architecture's specific contribution.","MM-Bench should be refreshed periodically with new contest problems, as the paper itself recommends, so that public competition data does not eventually contaminate LLM pretraining; a public annual update would preserve the benchmark's value for measuring progress."],"forward_implications":["MM-Agent produces complete competition-grade modeling reports in roughly 15 minutes and about $0.88 per task on GPT-4o, so the per-problem cost of expert-level modeling drops below one dollar.","Performance is similar on problems from 2021–2024 and from 2025, which the authors read as evidence that the results come from genuine modeling rather than memorized contest solutions.","Acting as a copilot under official MCM/ICM rules, MM-Agent helped two undergraduate teams reach Finalist, the top 2.0% of 27,456 teams, in MCM/ICM 2025.","On well-defined OPTIBENCH optimization problems, MM-Agent also beats GPT-4o in zero-shot settings and raises code pass rate to 99.3%, indicating the pipeline generalizes beyond open-ended modeling.","Stronger backbones help: MM-Agent on DeepSeek-R1-671B posts overall scores of 8.85 on 2021–2024 and 8.92 on 2025, above its GPT-4o scores, suggesting the architecture composes with model improvements."],"supporting_citations":[{"why":"GPT-4o system card; supplies the main base model used for all MM-Agent and baseline runs.","marker":"[6]"},{"why":"DeepSeek-R1; the stronger second backbone on which MM-Agent reaches its highest scores.","marker":"[7]"},{"why":"DS-Agent; the adapted data-science agent that is the strongest baseline competitor.","marker":"[33]"},{"why":"Agent Laboratory; the scientific-discovery baseline that MM-Agent beats while using less cost and runtime.","marker":"[46]"},{"why":"ResearchAgent; the research-agent baseline adapted for modeling tasks.","marker":"[50]"},{"why":"mGTE; the embedding model that drives similarity retrieval inside HMML.","marker":"[67]"},{"why":"MLE-Solver/MLE-bench; supplies the iterative code generation and debugging used in computational solving.","marker":"[68]"},{"why":"OPTIBENCH; the well-defined optimization benchmark used to test MM-Agent's generalization.","marker":"[4]"},{"why":"COMAP contest rules and evaluation criteria; the source of the four-rubric evaluation and of the MCM/ICM protocol behind the copilot result.","marker":"[66]"}],"fun_headline_variants":[],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's headline result, that an agent beats expert humans, rests on subjective scores from a four-criteria rubric assigned by GPT-4o and a small panel of contest-experienced humans, not on verifiable ground truth; if those judges are systematically friendlier to agent-produced reports, the reported advantage shrinks or disappears.","fun_headline_variants_meta":{"error":"Client error '402 Payment Required' for url 'https://api.deepseek.com/chat/completions'\nFor more information check: https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/402"},"cache_creation_input_tokens":0},"created_at":"2026-08-07T15:38:57.524463+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the 32-problem MM-Bench test set, remove authorship labels from human and MM-Agent reports, and have a fresh panel of at least three contest-experienced judges score them on the four rubric dimensions. If the blinded overall margin is no longer positive and statistically significant, especially on Modeling Rigorousness where the paper reports human-human agreement of only 0.4813, then the claim of beating award-winning human teams is not supported.","supporting_citations":[{"cited_title":"Ds-agent: Automated data science by empowering large language models with case-based reasoning","cited_arxiv_id":null,"evidence_quote":"DS-Agent; the adapted data-science agent that is the strongest baseline competitor."},{"cited_title":"the player who seemed to have the advantage are often attributed to “momentum","cited_arxiv_id":null,"evidence_quote":"MLE-Solver/MLE-bench; supplies the iterative code generation and debugging used in computational solving."},{"cited_title":"Mcm/icm contest rules, registration and instructions, 2025","cited_arxiv_id":null,"evidence_quote":"COMAP contest rules and evaluation criteria; the source of the four-rubric evaluation and of the MCM/ICM protocol behind the copilot result."}],"review_version":1}