{"id":"14fa5a06-a672-498c-9f09-05832f25248e","arxiv_id":"2507.02873","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"An LLM-assisted corpus study finds that roughly 3% to 12% of 5,000 arXiv math papers contain clear or borderline appeals to mathematical explanation, with frequency varying by subfield.","lead":"A philosopher used Google's Gemini 2.5 Pro to scan 5,000 arXiv math papers for places where mathematicians explain why results are true rather than just proving them. The result is a dataset of hundreds of examples and a first statistical look at how explanation talk varies across mathematics.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's evidence for Gemini's annotation accuracy is anecdotal; with ~20% low-quality and ~60% borderline cases, the subfield and prevalence results are not robust, so the central claim of 'accurate' corpus work is under-supported.","rationale":"The reader's weakest_assumption identifies exactly the load-bearing concern: Gemini 2.5 Pro's reliability as an annotator is validated only informally, while the paper's own estimates of dataset quality (20% low-quality, 60% borderline) suggest that the annotation labels are substantially noisy. This is not merely a question of adding error bars; it affects the central claim itself, which asserts that current models are capable of 'sophisticated, accurate and interesting corpus work.' If the model's precision or recall is significantly lower than assumed, the quantitative patterns in §3.1 and §3.2 could be artifacts, and the paper's proof-of-concept for LLM-based PMP research is weakened. I considered other concerns: the same model is used for data generation and interpretation, but the paper treats §3.4 as exploratory and does not rest the central capability claim on it; the prompt's reliance on the SEP characterization of explanation is acknowledged and is a design choice rather than a hidden flaw; and the D/C statistic does compare papers, not examples, so the ratio is not obviously conflating units. The strongest critique is therefore the absence of a systematic gold-standard evaluation. The reader's CONDITIONAL verdict is appropriate: the paper is honest and promising, but the empirical claims should be conditional on such validation. My read does not change the verdict, so I set verdict_should_be to UNCHANGED.","tokens_in":18904,"tokens_out":6694,"duration_ms":73357,"concrete_test":"Randomly sample 200 papers stratified by arXiv category from the 5000-paper corpus. Have two independent PMP experts annotate these papers for the presence of mathematical explanation using the same criteria and prompt guidelines from §2, blinded to Gemini's outputs. Measure Gemini's precision and recall against the human labels (with disagreements adjudicated by a third expert), and compute inter-annotator agreement (Cohen's kappa). Then recompute Table 1's D/C coefficients using only examples agreed as high-quality by both experts. If precision is below roughly 80%, recall below roughly 70%, kappa below 0.6, or if any subfield coefficient changes by more than 0.1, the empirical conclusions and the central claim of 'accurate' corpus work are not robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (§1) that Gemini 2.5 Pro can do 'sophisticated, accurate and interesting corpus work' hinges on the reliability of the model's annotations of mathematical explanation. The paper's validation consists of informal spot checks (§2), yet the author estimates the filtered dataset is only ~20% high-quality, ~20% low-quality, and ~60% borderline cases (§2). The quantitative findings that would demonstrate accuracy — the subfield 'coefficient of explanatory richness' (§3.1, Table 1) and the prevalence estimates (§3.2) — are computed from this noisy dataset with no sensitivity analysis or inter-annotator agreement. Since the author admits inclusion of borderline cases 'is in part a matter of taste,' the observed variation (e.g., logic/set theory 1.32 vs. probability/statistics 0.77) could shift substantially under reasonable alternative coding decisions. Furthermore, the prevalence estimate of 'at least ~3%' assumes low false negatives, but recall is never measured; if Gemini misses many clear cases, the contribution to the 'rarely versus routinely' debate is potentially misleading. Thus the load-bearing assumption — that the LLM's outputs are accurate enough to support empirically grounded philosophy — is not yet established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a large-scale corpus study of mathematical explanation in 5,000 arXiv mathematics papers, using Google's Gemini 2.5 Pro as an automated annotator. The author describes a pipeline in which the model is prompted with a long excerpt from the Stanford Encyclopedia of Philosophy article on mathematical explanation, processes papers in batches of 25, produces candidate annotated examples, and then applies a second strict filtering prompt. The resulting dataset contains roughly 1,250 examples from about 735 papers. From this dataset the paper reports subfield-level 'coefficients of explanatory richness' (Table 1), an estimate that at least ~3% of papers contain clear explanation claims and ~12% contain borderline-or-better cases, several targeted queries of the dataset, and an extended experiment in which Gemini adjudicates between philosophical theories of explanation and proposes a novel theory ('Explanatory Resonance Theory'). The paper frames itself as a proof of concept, arguing that current LLMs can perform corpus work that is 'sophisticated, accurate and interesting' on a scale impossible by other means.","tokens_in":19111,"tokens_out":2807,"duration_ms":29636,"significance":"If the central claim holds, the paper would establish a genuinely new methodological avenue for the philosophy of mathematical practice, moving beyond keyword-count corpus methods to concept-sensitive annotation at scale. The paper is transparent about its pipeline: the full prompt is quoted, the filtering instruction is quoted, and the dataset is made available as a 500+-page document. It also contains honest and explicit limitation statements, including the author's own estimate that roughly 20% of the filtered dataset is low-quality and 60% is borderline. These strengths are real. However, the paper's quantitative findings—the subfield coefficients and the prevalence estimates—rest on the reliability of Gemini's annotations, and that reliability is asserted rather than demonstrated. The absence of any gold-standard validation, inter-annotator agreement, or sensitivity analysis is a load-bearing gap, not a cosmetic one. The philosophical-adjudication experiment in §3.4 is also weakened by circularity: the same model that generated and filtered the dataset is asked to assess which theory best fits it.","major_comments":[{"comment":"The central claim that Gemini can do 'sophisticated, accurate and interesting corpus work' (§1) is not backed by systematic validation. The paper's own quality estimate at the end of §2 is that the filtered dataset is roughly 20% high-quality, 20% low-quality, and 60% borderline, with inclusion of borderline cases 'in part a matter of taste.' Yet Table 1 reports coefficients of explanatory richness (e.g., 1.32 for logic and set theory, 0.77 for probability and statistics) as though they are robust measurements. Because the filtering decision is made by the same model that created the examples, and because no gold standard, inter-annotator reliability measure, or human-expert audit is reported, the observed variation across subfields could easily shift under reasonable alternative coding decisions. A sensitivity analysis (e.g., recalculating coefficients under strict-only, borderline-included, and permissive inclusion policies) is needed before these numbers can support any empirical conclusion.","section":"§2 (Methods) and §3.1 (Table 1)"},{"comment":"The estimate that 'at least ~150 out of 5000 research papers (around 3%)' contain clear explanation claims relies on two unexamined assumptions: that the model's false-negative rate is negligible, and that example-level quality percentages transfer to paper-level counts. The paper explicitly says 'I expect it not to have missed large numbers of high-quality cases,' but recall is never measured. Since the filtering prompt instructed the model to exclude at least 50–60% of examples, the false-negative rate could be substantial, and the 'at least' claim would then be misleadingly low. The paper should report a recall-oriented validation (e.g., on a random subset of papers that human experts annotate independently) and present the prevalence estimate with confidence intervals or as a range across coding policies.","section":"§3.2 (Prevalence of explanatory concerns)"},{"comment":"The adjudication experiment is circular in a way that limits its evidentiary value. The same model (Gemini 2.5 Pro) generated the initial candidate examples, applied the filtering prompt, and was then asked to evaluate which philosophical theory of explanation best fits the filtered dataset. Its assessment is therefore shaped by its own earlier filtering decisions and by the SEP excerpt embedded in the original prompt, which already presupposes a particular taxonomy of explanatory concepts. While the discussion is interesting as an illustration of what LLMs can produce, any claim that it 'helps settle debates between rival theories' (§1) requires at least a blind comparison with human expert judgments on the same dataset. As it stands, the model's pluralist/epistemic conclusion is better described as a hypothesis generated by the method than as evidence for that hypothesis.","section":"§3.4 (Gemini as philosophical assistant)"},{"comment":"The filtering prompt includes the instruction 'You MUST exclude at least 50-60% of the original examples' and repeatedly urges the model to be 'ruthless.' This makes the filtering threshold a free parameter that directly influences all downstream quantitative claims, but no justification is given for choosing this threshold, and the paper reports that repeated applications of the filter produced no further changes. This is not merely a technical detail: the 3% and 12% prevalence estimates in §3.2 are computed from the filtered dataset, so the quantitative conclusions are partly determined by an arbitrarily chosen exclusion target. The paper should either justify the threshold empirically or present results as a function of filtering strictness.","section":"§2 (Filtering prompt)"}],"minor_comments":[{"comment":"There is a typo: 'algebra, topology and and combinatorics' should read 'algebra, topology and combinatorics.'","section":"§3.1"},{"comment":"The prompt and filtering instructions are reproduced in full, which is excellent for reproducibility, but the exact model version used for each stage (e.g., 2.5 Pro Experimental vs. Preview) is not always specified in the results; this matters for replication.","section":"§2"},{"comment":"The sentence 'In general, It's difficult to see what such a broad and ill-defined construct might add' has an unnecessarily capitalised 'It's' after the period.","section":"§3.4"},{"comment":"Reference [D'Alessandro 2025] is listed with 'DOI: XXXX,' which should be completed before publication.","section":"References"},{"comment":"The paper states that the dataset is available at a URL, but a persistent identifier (e.g., a DOI or a stable repository link) would be more appropriate for a dataset that is central to the paper's claims.","section":"§2"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nQuick take: this is the first serious attempt to use LLMs for corpus analysis in the philosophy of mathematical practice, and as a proof of concept it mostly works. The author builds a pipeline that runs Gemini 2.5 Pro over 5000 arXiv papers, extracts annotated examples of explanatory discourse, filters them, and produces a dataset of over 1000 references. That dataset alone is a real contribution: it gives philosophers something to work with beyond cherry-picked case studies.\n\nThe paper is unusually honest about its own weaknesses. The author states plainly that after filtering the data is roughly 20% high-quality, 20% low-quality, and 60% borderline. He also acknowledges the word-concept problem and the difficulty of automating judgments that are “in part a matter of taste.” That candor lets the reader see exactly where the empirical claims get shaky.\n\nThe main soft spot is the one the skeptic flags: no systematic validation. The same model that produces the examples also applies the filter and then adjudicates between philosophical theories using the filtered dataset. There is no gold standard, no inter-annotator agreement, and no error bars around the prevalence estimates or the subfield coefficients. The author’s spot checks are plausible but anecdotal. Given the 20/60 split, the subfield comparisons (e.g., logic 1.32 vs. probability 0.77) could easily move under different coding decisions. And recall is never measured, so the “at least ~3%” prevalence claim is really a lower bound that assumes the model doesn’t miss many clear cases—which we don’t know.\n\nThat said, I don’t think these problems sink the paper. The central claim is not that Gemini is a perfect annotator; it’s that LLMs can do useful, large-scale corpus work for PMP. The examples in §2 and §3.3 show that the model frequently identifies genuine explanatory discourse, including cases a keyword search would miss. The discussion of explanations of method and proof technology, and the observation about obstructions as explanatory, are genuinely interesting philosophical leads. The paper also cites the relevant literature fairly, and its self-citations are in service of continuity rather than inflation.\n\nBottom line: this is worth engaging with and worth sending to referees. It should be revised to add some form of external validation—even a small human-coded sample would help—and the quantitative sections should carry explicit caveats. But as a proof of concept with a public dataset, it deserves serious peer review rather than desk rejection.","headline":"First real LLM-corpus study in PMP; the dataset and honest methodology make it worth refereeing, but the quantitative claims need external validation.","tokens_in":19653,"tokens_out":1884,"would_cite":true,"duration_ms":20422,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["00A30","00A35"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper reports that a large language model can scan 5,000 mathematics papers, recognize genuine discussions of explanation, and produce an annotated corpus large enough to ground empirically driven philosophy of mathematics.","keywords":["mathematical explanation","philosophy of mathematical practice","large language models","corpus analysis","word-concept problem","explanatory proof","subfield variation","LLM annotation"],"falsifier":"Take a random sample of about 200 of the 5,000 corpus papers, have a panel of mathematicians and philosophers independently mark every passage that clearly discusses mathematical explanation, and compare their annotations with Gemini's flags on the same papers; if agreement is near chance or the model misses most passages experts agree on, the claim that the model can do accurate corpus work fails.","tokens_in":18669,"feed_emoji":"📐","tokens_out":7172,"duration_ms":75674,"temperature":0.7,"pith_summary":"The paper tries to establish that current large language models can do reliable, large-scale corpus analysis for the philosophy of mathematical practice, focused on when mathematicians discuss explanation. It uses Gemini 2.5 Pro to read 5,000 randomly sampled mathematics papers, prompted with a philosophical survey that defines the target concept, and produces a dataset of hundreds of annotated examples. The point is to move past cherry-picked case studies and ambiguous keyword counting, letting the model reason about meaning rather than just count words like 'explain'. On the empirical side, the paper reports that clear explanation talk appears in at least a few percent of recent papers, that its frequency varies by subfield, and that the model's analysis favors unificationist and ontic theories supplemented with epistemic and pragmatic elements. It concludes that LLM-assisted corpus work is feasible and valuable, while acknowledging that the dataset is imperfect.","feed_headline":"AI scans 5,000 math papers to find where proofs explain why","feed_subtitle":"A large-scale annotated corpus of mathematical explanation lets philosophers test theories without cherry-picked examples.","key_machinery":"The machinery is a large language model used as a semantic annotator: Gemini 2.5 Pro, whose one-million-token context window and integrated chain-of-thought reasoning let it process batches of 25 papers per query. The prompt embeds a roughly 5,000-word excerpt from an encyclopedia survey on mathematical explanation, which defines the target concept and supplies examples, along with instructions to quote sources and avoid hallucination. A Python script written by the model itself automated 200 runs over about 24 hours, and a second, stricter filtering prompt removed 50–60% of candidates, leaving the final annotated dataset. The defining move is replacing keyword counting with a prompt that asks the model to judge whether the concept of explanation is genuinely in play in each passage, regardless of the specific words used.","core_discovery":"The central discovery is that a current frontier LLM, prompted with a nuanced definition of mathematical explanation and run over a 5,000-paper random sample of the mathematics preprint archive, can produce useful annotations at a scale impossible by manual reading, thereby sidestepping the word-concept problem that plagued keyword-counting corpus methods. The author reports that Gemini 2.5 Pro identified roughly 1,250 candidate instances from about 735 distinct papers, with roughly 20% high-quality cases, 20% low-quality cases, and 60% borderline cases by his estimate. On substantive empirical questions, the paper claims that at least about 3% of papers contain clear explanation claims and about 12% contain borderline-or-better cases; that explanatory practice varies by subfield, with coefficients of explanatory richness ranging from 0.77 for probability and statistics to 1.32 for logic and set theory, with combinatorics at 1.19; and that when asked to adjudicate philosophical theories, the model argues for a combination of unificationism and ontic structure-revealing accounts, supplemented by epistemic or pragmatic elements to handle reproofs, heuristic explanations, and demystifications. A further reported finding is the model's observation that obstructions, or reasons why a strategy fails, often function as explanatory concerns, which the author identifies as a promising new research direction.","pith_inferences":["If LLM annotation is validated against human gold standards, the prompting-plus-filtering pattern could be exported to other philosophically loaded concepts such as beauty, simplicity, depth, and naturalness, giving empirical philosophy a general tool rather than a one-off study.","The subfield variation suggests a testable sociological hypothesis beyond the paper's table: explanatory talk clusters where mathematical objects admit multiple representational perspectives, such as algebra, geometry, topology, and combinatorics, and thins where a single analytic or computational framework dominates.","The model's obstruction observation points to a concrete coding extension: future studies could explicitly tag explanations of failure, such as why a theorem is false, why a method fails, or why a construction is obstructed, to see whether that category forms a substantial share of mathematical explanation.","One could push the paper's quality-filtering idea further by using the model's chain-of-thought traces to build self-scored confidence labels for each example, reducing the need for human spot checks in larger follow-up studies."],"forward_implications":["The same pipeline can be scaled to much larger corpora, including the full repository of roughly 80,000 mathematics preprints, making corpus philosophy of mathematics a practical research program.","The resulting dataset gives philosophers hundreds of clear and borderline examples of mathematical explanation that were not selected for their convenience or theory-friendliness.","Explanatory practice is not uniform across mathematics: subfield-specific richness coefficients imply that combinatorics and logic-and-set-theory communities invoke explanation more than average, while probability and statistics invoke it less.","The rarity debate about mathematical explanation is sharpened: at least about 3–12% of recent preprint papers engage explanation talk, enough to count as a settled practice without being ubiquitous.","LLMs can act as interpretative assistants, not just annotators: querying the dataset yields targeted examples, such as tradeoffs between explanatory and other proof virtues, and can propose new philosophical projects like the study of obstructions as explanations."],"supporting_citations":[{"why":"Supplies the conceptual definition and examples of mathematical explanation embedded in the prompt that shape what the model looks for.","marker":"[Mancosu et al. 2023]"},{"why":"Introduces the explanatory versus non-explanatory proof distinction that the study operationalizes in the corpus.","marker":"[Steiner 1978]"},{"why":"Provides the skeptical claim that mathematicians rarely describe themselves as explaining, which the prevalence estimates are meant to address.","marker":"[Resnik & Kushner 1987]"},{"why":"Identifies the word-concept problem for keyword-count corpus methods that the LLM approach claims to overcome.","marker":"[Chartrand 2022]"},{"why":"Argues that LLM-based automated text analysis is a promising methodological path for philosophy of mathematical practice.","marker":"[D'Alessandro 2025]"},{"why":"Exemplifies the corpus-counting approach whose limitations motivate the shift to semantic analysis by language models.","marker":"[Mizrahi 2020]"}],"fun_headline_variants":["AI reads 5,000 math papers to test theories of explanation","Gemini analyzes 5,000 arXiv papers on mathematical explanation","LLM corpus study reveals how math explanations vary by field","Obstructions as explanations: AI scans 5,000 math papers","Testing philosophy of math with 5,000 papers and one LLM"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole empirical picture rests on trusting that Gemini 2.5 Pro reads mathematics papers accurately enough to label explanations, since the author checked its outputs only informally and even he estimates that a fifth of the final dataset is low quality and another sixty percent is borderline.","fun_headline_variants_meta":{"raw":{"variants":["AI reads 5,000 math papers to test theories of explanation","Gemini analyzes 5,000 arXiv papers on mathematical explanation","LLM corpus study reveals how math explanations vary by field","Obstructions as explanations: AI scans 5,000 math papers","Testing philosophy of math with 5,000 papers and one LLM"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000206,"raw_usage":{"total_tokens":1464,"prompt_tokens":1083,"completion_tokens":381,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":699,"completion_tokens_details":{"reasoning_tokens":291}},"tokens_in":699,"tokens_out":381,"duration_ms":4257,"temperature":1.0,"reasoning_tokens":291,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:24:21.240603+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of about 200 of the 5,000 corpus papers, have a panel of mathematicians and philosophers independently mark every passage that clearly discusses mathematical explanation, and compare their annotations with Gemini's flags on the same papers; if agreement is near chance or the model misses most passages experts agree on, the claim that the model can do accurate corpus work fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the conceptual definition and examples of mathematical explanation embedded in the prompt that shape what the model looks for."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the explanatory versus non-explanatory proof distinction that the study operationalizes in the corpus."},{"cited_title":"and David Kushner","cited_arxiv_id":null,"evidence_quote":"Provides the skeptical claim that mathematicians rarely describe themselves as explaining, which the prevalence estimates are meant to address."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Identifies the word-concept problem for keyword-count corpus methods that the LLM approach claims to overcome."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Argues that LLM-based automated text analysis is a promising methodological path for philosophy of mathematical practice."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Exemplifies the corpus-counting approach whose limitations motivate the shift to semantic analysis by language models."}],"review_version":1}