{"id":"f7e26633-4947-495d-b7df-abbd3b4bf003","arxiv_id":"2507.11936","paper_version":6,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"This survey organizes deep learning work on geometry problem solving into task, method, benchmark, and evaluation categories, and highlights open challenges.","lead":"This paper surveys deep learning systems that solve geometry problems, covering task types, datasets, model architectures, and evaluation methods. It is a reference for researchers and engineers working on AI mathematical reasoning and multimodal understanding.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Comprehensiveness claim lacks quantitative support: §1 says >310 papers but Figure 5's yearly counts sum to 250, with data only through Apr 2025, and the search protocol is a single keyword plus one snowballing round.","rationale":"The reader's weakest_assumption correctly targets the lightly documented search protocol in §1. My stress test agrees but sharpens it: the paper's own numbers contradict each other. Figure 5's yearly counts sum to 250, not 310+, and the figure is not updated to full-year 2025 or 2026, despite citations to 2026 work. This is not a stylistic nit; it is the quantitative backbone of the comprehensiveness claim. Still, the survey has real organizational value: a sensible taxonomy, dense dataset tables, and an honest limitations section, and the GitHub list provides a checkable artifact. These strengths justify conditional acceptance, matching the reader's verdict. I do not think the contradiction is fatal—surveys routinely have stale figures—but it does mean the 'comprehensive' claim should be softened or the count and figure should be reconciled before unconditional acceptance. The reader's verdict of CONDITIONAL already reflects this level of caution, so I recommend UNCHANGED rather than moving to REJECT or ACCEPT.","tokens_in":44462,"tokens_out":6074,"duration_ms":67656,"concrete_test":"Download the maintained paper list at https://github.com/majianz/dl4gps, count unique entries, and compare with the §1 claim of 'more than 310' and with the sum of Figure 5 yearly counts (currently 250). If the list has ≤250 entries, the stated count is inconsistent; if it has ≥310 while Figure 5 is stale, the figure must be updated or labeled as a subset. Additionally, run an independent query-based search (e.g., Semantic Scholar/arXiv, 2018-2026) with 'geometry problem solving', 'geometric reasoning', 'geometry question answering', 'Euclidean geometry', and measure recall of the survey's included papers; a recall well below 80% would confirm that the single-keyword snowballing protocol missed substantial relevant literature.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is to provide 'a comprehensive and practical reference' of deep learning for GPS (§1). The load-bearing support is a claimed corpus of 'more than 310 academic papers' collected via a single Google Scholar search for the keyword 'geometry' and one round of forward/backward snowballing. This support is internally inconsistent: the appendix's own Figure 5, which plots 'papers on deep learning for geometry problem solving over the years', lists yearly counts 1 (2018), 5 (2019), 3 (2020), 6 (2021), 15 (2022), 25 (2023), 112 (2024), and 83 (Jan-Apr 2025), summing to 250. The text was revised to include 2026-dated references, yet the figure still terminates at April 2025, so the figure does not account for the claimed >310 count. No alternative definition of the subset is given. The Limitations section itself concedes the survey 'may not fully represent the development process of the entire field.' Because the comprehensiveness claim rests on an unverifiable and self-contradictory count, the paper's status as a definitive reference is weaker than stated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper surveys deep-learning approaches to geometry problem solving (GPS). It proposes a three-level task taxonomy (fundamental, core, and composite tasks), reviews the associated datasets, architectures, training-stage and inference-stage methods, and discusses automatic and manual evaluation metrics. It also compares a small set of representative models on three benchmarks and closes with challenges and future directions. The manuscript provides dense dataset and architecture tables in the appendix and points to a community-maintained GitHub list of relevant papers.","tokens_in":44703,"tokens_out":4895,"duration_ms":55360,"significance":"The survey addresses a genuine gap: broader mathematical-reasoning surveys treat geometry only as a subtopic, and prior GPS-specific surveys are narrower in scope. The proposed taxonomy is reasonable, and the appendix tables (Tables 2–5) offer a practical consolidation of datasets and system components that researchers entering the area will likely find useful. The paper is a survey rather than a derivation, so circularity is not at issue; its value rests on the completeness and accuracy of its coverage. The main risk is that the advertised corpus size and the search protocol are not internally consistent, which directly weakens the central claim of being a comprehensive reference. With the coverage claims repaired, the survey would be a solid entry point to the field.","major_comments":[{"comment":"The introduction claims that the survey collected \"more than 310 academic papers\" on deep learning for GPS, but the paper's own Figure 5 reports annual counts of 1, 5, 3, 6, 15, 25, 112, and 83 for 2018 through April 2025, which sum to only 250. The figure does not cover the 2026-dated references that appear elsewhere in the bibliography, and no alternative definition of the corpus is provided. Because the abstract and introduction stake the paper's value on being \"a comprehensive and practical reference,\" this internal inconsistency in the corpus count is load-bearing and must be resolved, either by updating the figure and the count or by qualifying the comprehensiveness claim.","section":"§1, Figure 5 (Appendix)"},{"comment":"The search protocol described in Section 1 (a single Google Scholar keyword \"geometry\" plus one round of forward and backward snowballing) is not sufficient to substantiate the claim of comprehensive coverage, and the manuscript's own Limitations section concedes that the survey \"may not fully represent the development process of the entire field.\" Please provide a more detailed and transparent protocol—database and query strings, snowballing directions, inclusion/exclusion criteria, screening counts, and the relationship between the Figure 5 subset and the claimed >310 papers—and move the coverage caveat into the main text where it qualifies the central claim.","section":"§1, Limitations"},{"comment":"Table 1 reports state-of-the-art and second-best results on Geometry3K, GeoQA, and MathVista, but the numbers are taken from the original papers without specifying the evaluation protocol used in each case (e.g., answer extraction, multiple-choice versus free-form output, model version, or prompt format). The four trends stated in §5.1, such as \"neural-symbolic methods demonstrate superior performance on symbolic-oriented tasks,\" rely on cross-model comparisons between numbers that may not be directly comparable. Please either state the conditions under which each score was obtained or limit the trends to subsets with a verified common evaluation setting.","section":"§5.1, Table 1"}],"minor_comments":[{"comment":"The citation \"Sinha et al.\" appears without a year in §3.1.2 (e.g., \"Trinh et al., 2024; Sinha et al.; Chervonyi et al., 2025\"), and the corresponding reference entry also lacks a year and venue. The citation \"Tey\" in §3.2.2 and §3.3.1 similarly points to a bibliography entry with no year and no publication venue. These entries should be completed or removed.","section":"§3.1.2, §3.3.1, References"},{"comment":"The composite-datasets table contains a row labeled \"MATH()(2024)\" with an apparently lost version tag, and several cells are left blank rather than marked with \"—\" or \"N/A\", which makes the table harder to parse and less reliable as a reference.","section":"Table 3"},{"comment":"The figure caption says \"data for 2025 is up to April,\" but the x-axis label reads \"(Jan. Apr.)\" without the year, and the figure should either be extended to include the 2026 references cited in the text or explicitly state that the count is only through April 2025.","section":"Figure 5"},{"comment":"There are numerous mechanical formatting artifacts, including \"LLaV A\" instead of \"LLaVA\", \"MA VIS\" instead of \"MAVIS\", \"DEBRUP DAS\" in the reference list, and the semicolon after \"Tey\". A careful copyedit would improve the manuscript's usability as a reference.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The manuscript has clearly undergone several revisions, but the corpus-count contradiction and the incomplete citations suggest the revision process has not fully converged. I would ask the authors to reconcile the coverage statistics and the search protocol before the paper is accepted, and to do a systematic pass over the bibliography."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a useful organizational survey, not a breakthrough. Its taxonomy of fundamental/core/composite tasks, the training/inference method decomposition, and the detailed dataset and architecture tables are the real value. If you work on geometry problem solving, you can use this as a starting map.\n\nThe paper does several things well. It constructs a clear three-level task taxonomy and places datasets and systems within it. Tables 2-5 are dense and mostly accurate, and the maintained GitHub list is a practical touch. The limitations section is honest: it admits theorem proving is underrepresented, that solid/analytic geometry are thin, and that rationale annotations are lacking. That kind of self-assessment is rare and helpful.\n\nNow the soft spots. The paper claims to have collected 'more than 310 academic papers' via a single Google Scholar search for 'geometry' plus one snowballing round. The appendix figure 5, which presumably counts the same corpus, sums to 250 papers through April 2025. There is a real inconsistency here. Either the count is inflated or the figure is incomplete; the paper doesn't explain the discrepancy. That matters because the abstract sells the survey as a 'comprehensive and practical reference.' The search protocol is also underdocumented: one keyword and one round of snowballing is a thin basis for a comprehensiveness claim, and the limitations section concedes the survey 'may not fully represent the development process of the entire field.' That concession is good, but it sits awkwardly with the strong framing.\n\nOther issues are minor: a few malformed references (e.g., Joseph Tey with no year; an entry in Table 3 that appears as 'MATH()' with no author), and the performance table reproduces numbers from original papers without standardizing evaluation settings. Those are fixable in revision.\n\nOn balance, the central organizational contribution holds up. The taxonomy is sensible, the tables will save people time, and the discussion of challenges (evaluation saturation, data scarcity, perception bottlenecks) is level-headed. The internal count inconsistency is a flaw, but it doesn't invalidate the survey's mapping of the field.\n\nWho's this for? Someone new to geometry problem solving who wants a structured entry point, or an experienced researcher checking coverage of datasets and methods. It deserves a serious referee. I'd send it to review with a request to reconcile the paper count, document the search protocol, and clean up the references.","headline":"A genuinely useful mapping of the geometry-solving literature, but the paper's own numbers undermine its comprehensiveness claim; the taxonomy and tables carry it.","tokens_in":45201,"tokens_out":2704,"would_cite":true,"duration_ms":29264,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey organizes deep learning for geometry problem solving into a three-tier task taxonomy and a three-axis method taxonomy.","keywords":["geometry problem solving","deep learning","multimodal large language models","theorem proving","geometric numerical calculation","neural-symbolic reasoning","mathematical reasoning benchmarks","survey"],"falsifier":"Searching with related terms such as 'math word problem', 'diagram parsing', or 'geometric reasoning' and counting GPS papers absent from this survey would test the claim that about 310 papers suffice; if dozens of relevant papers are missing, the comprehensiveness claim weakens. A second test is to inspect the reference list for a known relevant paper and see whether the survey discusses it.","tokens_in":44290,"feed_emoji":"📐","tokens_out":7571,"duration_ms":77180,"temperature":0.7,"pith_summary":"Geometry problem solving (GPS) is a multimodal mathematical task that requires reading diagrams, parsing formal statements, and performing proof or calculation. This survey aims to be a comprehensive reference for deep learning work on GPS, collecting more than 310 papers and organizing them into fundamental, core, and composite tasks, with methods classified by architecture, training stage, and inference stage. It also analyzes evaluation metrics, benchmarks representative models on Geometry3K, GeoQA, and MathVista-GPS, and identifies four trends: reinforcement learning helps, neural-symbolic methods stay competitive on symbolic benchmarks, high-quality large-scale data matters, and test-time scaling is promising. If the survey's organization is right, a newcomer can use it to locate the datasets, methods, and open problems of the field in one place.","feed_headline":"Survey maps 310 geometry-AI papers into a three-tier taxonomy","feed_subtitle":"Task, method, and evaluation categories reveal where theorem proving, solid geometry, and reasoning metrics lag.","key_machinery":"The carrying device is the survey's three-level task taxonomy (fundamental, core, composite) combined with a method taxonomy over architecture, training stage, and inference stage. This taxonomy is what lets the survey compare otherwise disparate systems: it assigns each dataset to a task level, places every reviewed method on the architecture/training/inference axes, and turns the field's gaps (few theorem-proving datasets, little solid and analytic geometry, mostly multiple-choice evaluation, scarce reasoning-process annotations) into visible vacancies rather than scattered observations.","core_discovery":"The central claim is that the field has a stable structure that can be described by a task taxonomy and a method taxonomy. Tasks fall into fundamental abilities (diagram understanding, semantic parsing, relation extraction, knowledge prediction), core tasks (theorem proving and numerical calculation), and composite tasks (mathematical reasoning); methods fall into architectures (encoder-decoder with text encoder, diagram encoder, fusion module, decoder, and knowledge module) and strategies at training time (pre-training, supervised fine-tuning, reinforcement learning) and inference time (test-time scaling, knowledge-augmented inference). Within this structure, the paper reports that state-of-the-art results come from large high-quality data and RL, that neural-symbolic solvers still lead the symbolic Geometry3K benchmark, and that diagram perception is the current bottleneck, with adding diagrams sometimes lowering accuracy. It concludes that data, evaluation, and perception gaps define the field's open agenda.","pith_inferences":["Editorial inference: a literature search built on one keyword and one snowballing round likely undercounts work published under names like 'math word problem' or 'diagram parsing', so the 310-paper corpus is probably a lower bound and the taxonomy may need periodic re-screening.","Editorial inference: the fundamental/core/composite structure is general enough that it could be transferred to other diagram-based reasoning domains, such as physics or chart-based question answering.","Editorial inference: because the performance table mixes benchmarks with different question formats, answer grading, and image types, a controlled benchmark that fixes these variables would directly test the reported trends."],"forward_implications":["Researchers can use the taxonomy to choose an underserved niche, since the paper identifies theorem proving, solid and analytic geometry, non-English datasets, and reasoning-process annotations as thin.","Because most benchmarks are multiple-choice, reported accuracy may overstate ability; the paper's call for option-free and harder evaluation implies current leaderboards need re-reading.","If the four performance trends hold, future GPS systems will likely pair large high-quality training data and reinforcement learning with neural-symbolic components rather than rely on a single architecture.","As foundation models improve and existing benchmarks saturate, progress will depend on process-based metrics and efficiency measures, not just answer accuracy."],"supporting_citations":[{"why":"Provides the Geometry3K benchmark and the Inter-GPS neural-symbolic solver that anchors the performance analysis.","marker":"Lu et al. (2021)"},{"why":"Provides the GeoQA dataset, the main numerical-calculation benchmark for GeoQA-test results.","marker":"Chen et al. (2021)"},{"why":"Provides the MATH benchmark and AMPS pre-training data used in the composite-task and pre-training discussion.","marker":"Hendrycks et al. (2021)"},{"why":"Provides the MathVista benchmark whose GPS subset is used for state-of-the-art comparison.","marker":"Lu et al. (2024)"},{"why":"AlphaGeometry, the decoder-only synthetic-data theorem prover that represents the theorem-proving line.","marker":"Trinh et al. (2024)"},{"why":"URSA, the process-reward-model system whose result supports the large-scale-data and RL trends.","marker":"Luo et al. (2025)"},{"why":"LANS, the layout-aware neural solver that achieves the best Geometry3K result reported.","marker":"Li et al. (2024g)"},{"why":"SANS, the spatial-aware neural solver that gives the second-best Geometry3K result.","marker":"Lin et al. (2024)"},{"why":"GeoSense, a knowledge-prediction dataset that grounds the fundamental-task category.","marker":"Xu et al. (2025c)"},{"why":"GNS-260K, a knowledge-prediction and numerical-calculation dataset that grounds the knowledge-prediction task.","marker":"Ning et al. (2025)"}],"fun_headline_variants":["Geometry AI survey: three-tier taxonomy, diagram perception bottleneck","Diagram perception is the bottleneck in geometry AI, survey finds","Neural-symbolic solvers still lead Geometry3K benchmark","RL and large-quality data drive state-of-the-art geometry AI","Geometry problem solving survey: taxonomy, gaps, and RL-driven progress"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's comprehensiveness rests on a literature search that used one keyword and one round of snowballing, and the paper itself says this may not fully represent the field.","fun_headline_variants_meta":{"raw":{"variants":["Geometry AI survey: three-tier taxonomy, diagram perception bottleneck","Diagram perception is the bottleneck in geometry AI, survey finds","Neural-symbolic solvers still lead Geometry3K benchmark","RL and large-quality data drive state-of-the-art geometry AI","Geometry problem solving survey: taxonomy, gaps, and RL-driven progress"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00047,"raw_usage":{"total_tokens":2305,"prompt_tokens":875,"completion_tokens":1430,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":491,"completion_tokens_details":{"reasoning_tokens":1358}},"tokens_in":491,"tokens_out":1430,"duration_ms":11273,"temperature":1.0,"reasoning_tokens":1358,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:57:15.331266+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Searching with related terms such as 'math word problem', 'diagram parsing', or 'geometric reasoning' and counting GPS papers absent from this survey would test the claim that about 310 papers suffice; if dozens of relevant papers are missing, the comprehensiveness claim weakens. A second test is to inspect the reference list for a known relevant paper and see whether the survey discusses it.","supporting_citations":[],"review_version":1}