REVIEW 4 major objections 5 minor 42 references
A Matter of Interest: Understanding Interestingness of Math Problems in Humans and Language Models
T0 review · 4 major / 5 minor · reviewed 2026-08-03 · deepseek-v4-flash
Pith's one-line read Although language models broadly agree with human average ratings of math-problem interestingness, they largely fail to reproduce the distribution of human judgments and only weakly match the reasons humans give.
desk verdict Two solid new human datasets and a careful mean-vs-distribution comparison; the negative 'why' result is real but needs prompt robustness checks. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The analysis rests on two purpose-built datasets: 63 crowdsourced raters judging 18 AMC problems (originals and hand-written variants) and 48 IMO participants judging past IMO problems with rationale checklists. Three instruments carry the argument: per-problem mean squared correlation (R2) measures overall agreement; Wasserstein distance compares the full distribution of human versus model ratings; and rationale-importance ratings show whether models weight the same features (elegance, novelty, generality) as humans. The human split-half R2 (0.71) and split-half Wasserstein (9.5) serve as noise ceilings that calibrate how much agreement is even possible.
What would settle it
Run the same rating task with two or three differently worded prompts—including one that explicitly asks for a rating on a 0-100 scale and one that asks how a typical person would rate the problem—and check whether the Wasserstein distance to human ratings falls to the human split-half level; if it does, the distributional failure seen here is a prompt artifact rather than a stable model property.
Extended reading notes
Core claim
The paper's central finding is that current LLMs have a partial but systematically limited grasp of what makes a math problem interesting. Against two human datasets—AMC problems rated by crowdsourced participants and IMO problems rated by Olympiad competitors—models attain squared correlations with human mean ratings as high as 0.78, close to the human split-half ceiling of 0.71. Yet Wasserstein distances show most models place ratings in a narrower band than humans, and on IMO rationale questions most models assign only one or two importance levels where humans use the full scale. Only the smaller Mistral models came close on both distribution and rationales. Thus, judged by the paper's cr
Load-bearing premise
The paper assumes that fixing one prompt template, one sampling temperature, and 50 samples fairly captures each LLM's own interestingness distribution, so the narrow rating ranges reflect model nature rather than prompt or sampling artifacts; no prompt ablations are reported.
Editorial extensions
If this is right
- High R2 with human means is not sufficient evidence of alignment; LLM interestingness should be evaluated distributionally and by rationale alignment, not only by average correlation.
- In applications such as curriculum design or automated mathematical discovery, relying on a single LLM risks homogenized interestingness; multi-LLM and human-AI collaborative systems better reflect human diversity.
- Model family and size affect subjective judgment: the Mistral 7B and 24B models were closest to human distributions and rationales, so model selection matters for human-facing uses.
- LRM reasoning-token patterns reveal a 'flash judgment' effect for problems rated uninteresting, which could serve as a cheap signal in automated problem screening.
- After filtering for validity, LLMs can generate problems that people find engaging, suggesting a practical path for AI-assisted problem generation even while judgment alignment remains incomplete.
Reading between the lines
- Beyond the paper: the observed distributional collapse could be partly a prompt artifact rather than a stable model property; reworded prompts or explicit instructions to use the full scale might substantially reduce Wasserstein distances.
- Beyond the paper: the paper's rationale checklist is human-defined; LLMs might express different but internally coherent rationales, so an open-ended rationale comparison could change the conclusion about 'why' alignment.
- Beyond the paper: if the distributional gap persists across prompts, then LLM-driven discovery systems may systematically under-explore problems that humans find surprising or elegant—a testable prediction for automated conjecture generation.
- Beyond the paper: since the human samples are confined to competition-level mathematics and self-selected adults, calibrating models to beginners or research mathematicians could reveal different alignment patterns than the ones reported here.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies whether LLMs and large reasoning models (LRMs) align with human judgments of mathematical problem interestingness and difficulty. It contributes two datasets: a Prolific crowdsourced study with AMC-derived problems and a survey of IMO 2024 participants. Models are compared with humans on per-problem mean ratings (squared Pearson correlation), on rating distributions (Wasserstein distance with bootstrap confidence intervals), and on importance ratings for preset interestingness criteria. The main empirical claims are that (i) several LLMs correlate surprisingly well with human mean ratings (R2 up to 0.78 against a split-half human baseline of 0.71), but (ii) most models fail to match the distribution of human judgments, and (iii) most LLMs only weakly align with why humans find problems interesting, collapsing importance ratings to one or two response levels. The abstract also states that LLM-generated problems were evaluated for interestingness after validity filtering, but no such generation experiment appears in the manuscript.
Significance. If the results hold, the paper provides a useful empirical benchmark for a relatively underexplored capability: judging subjective properties of mathematical tasks rather than solving them. The distinction between matching average human ratings and matching the distribution of human judgments is important and generally well operationalized through Wasserstein distances with human split-half baselines. The inclusion of both crowd and expert populations, negative and positive controls, multiple model families, and two sampling temperatures are strengths. The rationale-alignment analysis is a novel and potentially valuable contribution, though its current support is incomplete. The manuscript is appropriate for a workshop or conference venue interested in AI evaluation, human-AI collaboration, and mathematical reasoning, provided the missing evidence and protocol details are addressed.
major comments (4)
- [Section 4, Appendix A1.2, Figures 19–30] The central negative claim that most LLMs do not reflect why humans find problems interesting rests on the IMO importance-rating replication, but the exact prompt template, response format, and scale anchors are not reported. Most models collapse to one or two importance levels despite 50 samples at temperature 1.0, so the result could be a prompt artifact rather than a model property. The paper's own Prolific data show large temperature sensitivity (Mistral 7B WD 12.4 at temperature 1.0, Table 1, vs 19.0 at temperature 0.3, Table 3), demonstrating that the elicitation protocol can materially change model behavior. Please report the full prompt and run at least one ablation (e.g., alternative wording, numeric vs verbal scale, or free-text rationales coded against the same criteria) to establish that the observed collapse is not an artifact.
- [Abstract and §5 (Discussion)] The abstract claims: 'Finally, we evaluate LLMs' ability to generate interesting problems and find that, after filtering for validity, LLMs are able to generate engaging problems.' No generation experiment, dataset, or analysis appears anywhere in the provided manuscript. This is a load-bearing claim in the abstract's summary of contributions. Either add the generation study with its methods and results, or remove the claim from the abstract and position the paper as covering only judgment and rationale alignment.
- [Section 4, first paragraph; Figure 2; Appendix A1.1.4] The model-human R2 values for per-problem mean interestingness are reported as point estimates over only 18 Prolific problems, with no confidence intervals or significance tests. The human split-half baseline is reported with a 95% CI [0.53, 0.87], but the model R2 values are not. Without uncertainty quantification, the statement that LLMs 'are able to approximate human perceptions of interestingness with surprising fidelity' and the ranking of model families (e.g., Mistral strongest) is not supported. Please provide bootstrap or permutation-based CIs for each model's human agreement and, where relevant, tests comparing models against the human split-half baseline and against each other.
- [Section 4, 'Most LLMs do not reflect why...'; Appendix A4.1] The rationale-alignment comparison uses 11 interestingness criteria authored by the research team. Human participants also provided free-text reasons, and the paper says LLM responses were sampled, but no analysis of free-text rationales is reported. Since the claim is about why humans find problems interesting, and the comparison is against a preset taxonomy, it would strengthen the result to show that the preset criteria capture the reasons humans actually give (e.g., by coding the free-text responses) or to compare free-text rationales directly. This is particularly important because the main 'why' finding is a negative result that could in part reflect the taxonomy rather than the models.
minor comments (5)
- [§3, Contributions] Typo: 'problem interestignness' should be 'problem interestingness'.
- [Figure 3 caption] The caption says 'Slow judgments occupy the bottom quartile and fast judgments occupy the top one.' This appears reversed relative to the text, which describes fast flash judgments for low-interest problems (short reasoning chains). Clarify which quartile corresponds to slow versus fast.
- [Section 4, LRM judgment lengths] The claim that all four LRMs make fast judgments of uninterestingness on the Prolific dataset is based on visual inspection of Figure 3 and Appendix Figure 6; no statistical test or effect size is reported for the difference in reasoning-token distributions between low- and high-interest problems. A simple nonparametric test would make the result more robust.
- [Section 3, IMO data collection] Participants were asked whether they wanted to see the solution before rating interestingness and difficulty, but the manuscript does not report whether seeing the solution affected ratings. This is a potentially important confound for the IMO rationale analysis.
- [Table 7, Appendix A3.2] The 'Positive control' row is a substantive math problem, while the 'Negative control' is a single arithmetic fact. The filtering rule in §3 uses only the negative control. Consider explaining why the positive control is not used in any filtering or calibration analysis.
Circularity Check
No significant circularity: the central comparisons are external empirical measurements against human judgments, with no fitted quantity renamed as a prediction and no load-bearing self-citation chain.
full rationale
The paper's central claims are empirical comparisons between human judgments and LLM judgments. Section 4 states: 'For each LLM, we compute the R2 between per-problem mean interestingness in humans and the model,' and 'We measure distributional similarity between human and LLM judgments using the Wasserstein Distance (WD).' These quantities are evaluated against external human data, are not fit to the models being evaluated, and are not derived from any assumption about what the models should output. Similarly, the rationale-alignment analysis compares LLM importance ratings to independently collected human importance ratings over the same survey items; although the 11 criteria were authored by the research team, this is a measurement-instrument choice, not a result that is true by construction. Self-citations such as [6] and [7] provide framing and related work but are not load-bearing: the conclusions do not rest on an unverified theorem or prior result from the same authors. The Discussion candidly notes limits ('the IMO study contained very few ratings for each problem...'), and the abstract's unsupported generation claim is a completeness/correctness concern, not circularity. No equation or fitted parameter reduces to an input; the derivation chain is self-contained empirical measurement.
Assumptions & free parameters
free parameters (3)
- Negative-control exclusion threshold =
90 (interestingness rating for '28+13')
- Median split for low/high-interest LRM analysis =
per-model median interestingness score
- Sampling budgets =
20 responses (Prolific, per model per temperature), 50 (IMO rationale), 5 (IMO LRMs)
assumptions (5)
- domain assumption A 0-100 interestingness rating after a 1-minute minimum reflection measures the intended construct.
- domain assumption The 11 interestingness and 12 uninterestingness criteria exhaust the main reasons people judge competition problems.
- domain assumption AMC and IMO problems are a usable sample for studying mathematical interestingness.
- ad hoc to paper A fixed prompt template and temperature sampling fairly represent each LLM's judgment distribution.
- ad hoc to paper The negative-control filter identifies unfaithful participants.
Cite this review
Pith. "Pith review of A Matter of Interest: Understanding Interestingness of Math Problems in Humans and Language Models." pith.science (2026). https://pith.science/paper/TNWFJVDW
@misc{pith2026251108548,
author = {Pith},
title = {Pith review of: A Matter of Interest: Understanding Interestingness of Math Problems in Humans and Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/TNWFJVDW}},
note = {Machine review of arXiv:2511.08548}
}
read the original abstract
The evolution of mathematics is shaped importantly by interestingness: researchers choose which problems to pursue, and students choose which problems to engage with, based on expectations of interest and challenge. As AI systems, particularly large language models (LLMs) that operate flexibly over natural language and formal mathematics, are increasingly used in mathematics research and education, it becomes crucial to characterize how closely their judgments align with people from different mathematical backgrounds. We study whether LLMs align with human interestingness judgments by comparing LLM ratings with those of two populations, crowdsourced participants with college math experience and International Math Olympiad competitors. Although many LLMs broadly agree with human notions of interestingness, they largely fail to match the distribution of human judgments. They also weakly align with why humans find problems interesting, with low correlation to human-selected rationales. Finally, we evaluate LLMs' ability to generate interesting problems and find that, after filtering for validity, LLMs are able to generate engaging problems. We conclude with takeaways, including the need for multi-LLM human-AI collaborative systems, that highlight both the promise and current limits of LLMs as partners in mathematical reasoning.
Figures
Figures from the paper (27 more)
Reference graph
Works this paper leans on
-
[1]
Aristotle: Imo-level automated theorem proving, 2025
Tudor Achim, Alex Best, Alberto Bietti, Kevin Der, Mathïs Fédérico, Sergei Gukov, Daniel Halpern-Leistner, Kirsten Henningsgard, Yury Kudryashov, Alexander Meiburg, Martin Michelsen, Riley Patterson, Eric Rodriguez, Laura Scharff, Vikram Shanker, Vladmir Sicca, Hari Sowrirajan, Aidan Swope, Matyas Tamas, Vlad Tenev, Jonathan Thomm, Harold Williams, and La...
2025
-
[2]
Daniel E. Berlyne. Complexity and incongruity variables as determinants of exploratory choice and evaluative ratings.Canadian Journal of Psychology / Revue canadienne de psychologie, 17(3):274–290, 1963
1963
-
[3]
Sparks of artificial general intelligence: Early experiments with gpt-4, 2023
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang. Sparks of artificial general intelligence: Early experiments with gpt-4, 2023
2023
-
[4]
What makes people think a puzzle is fun to solve? InProceedings of the Annual Meeting of the Cognitive Science Society, volume 47, 2025
Junyi Chu, Kristine Zheng, and Judith E Fan. What makes people think a puzzle is fun to solve? InProceedings of the Annual Meeting of the Cognitive Science Society, volume 47, 2025
2025
-
[5]
Training verifiers to solve math word problems, 2021
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems, 2021
2021
-
[6]
Building machines that learn and think with people.Nature human behaviour, 8(10):1851–1863, 2024
Katherine M Collins, Ilia Sucholutsky, Umang Bhatt, Kartik Chandra, Lionel Wong, Mina Lee, Cedegao E Zhang, Tan Zhi-Xuan, Mark Ho, Vikash Mansinghka, et al. Building machines that learn and think with people.Nature human behaviour, 8(10):1851–1863, 2024
2024
-
[7]
Evaluating language models’ evaluations of games.arXiv preprint arXiv:2510.10930, 2025
Katherine M Collins, Cedegao E Zhang, Graham Todd, Lance Ying, Mauricio Barba da Costa, Ryan Liu, Prafull Sharma, Adrian Weller, Ionatan Kuperwajs, Lionel Wong, et al. Evaluating language models’ evaluations of games.arXiv preprint arXiv:2510.10930, 2025
arXiv 2025
-
[8]
Katherine M Collins, Cedegao E Zhang, Lionel Wong, Mauricio Barba da Costa, Graham Todd, Adrian Weller, Samuel J Cheyette, Thomas L Griffiths, and Joshua B Tenenbaum. People use fast, flat goal-directed simulation to reason about novel problems.arXiv preprint arXiv:2510.11503, 2025
arXiv 2025
Show all 42 references
-
[9]
Automatic concept formation in pure mathematics
Simon Colton, Alan Bundy, and Toby Walsh. Automatic concept formation in pure mathematics. InInternational Joint Conference on Artificial Intelligence, 1999
1999
-
[10]
On the notion of interestingness in automated mathematical discovery.International Journal of Human-Computer Studies, 53(3):351–375, 2000
Simon Colton, Alan Bundy, and Toby Walsh. On the notion of interestingness in automated mathematical discovery.International Journal of Human-Computer Studies, 53(3):351–375, 2000
2000
-
[11]
Evaluations of subjective complexity, pleasingness and interestingness for a series of random polygons varying in complexity.Perception & Psychophysics, 2:281–286, 1967
Harris Day. Evaluations of subjective complexity, pleasingness and interestingness for a series of random polygons varying in complexity.Perception & Psychophysics, 2:281–286, 1967
1967
-
[12]
Role of specific curiosity in school achievement.Journal of Educational Psychology, 59:37–43, 02 1968
Hy Day. Role of specific curiosity in school achievement.Journal of Educational Psychology, 59:37–43, 02 1968
1968
-
[13]
Susan L. Epstein. On the discovery of mathematical theorems. InProceedings of the 10th International Joint Conference on Artificial Intelligence - Volume 1, IJCAI’87, page 194–197, San Francisco, CA, USA, 1987. Morgan Kaufmann Publishers Inc
1987
-
[14]
Alhussein Fawzi, Matej Balog, Aja Huang, Thomas Hubert, Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Francisco J. R. Ruiz, Julian Schrittwieser, Grzegorz Swirszcz, David Silver, Demis Hassabis, and Pushmeet Kohli. Discovering faster matrix multiplicat...
2022
-
[15]
Data for math- ematical copilots: Better ways of presenting proofs for machine learning.arXiv preprint arXiv:2412.15184, 2024
Simon Frieder, Jonas Bayer, Katherine M Collins, Julius Berner, Jacob Loader, András Juhász, Fabian Ruehle, Sean Welleck, Gabriel Poesia, Ryan-Rhys Griffiths, et al. Data for math- ematical copilots: Better ways of presenting proofs for machine learning.arXiv preprint arXiv:24...
2024
-
[16]
Smith, Kevin Buzzard, Timothy Gowers, Peter J
Simon Frieder, Sam Bealing, Arsenii Nikolaiev, Geoff C. Smith, Kevin Buzzard, Timothy Gowers, Peter J. Liu, Po-Shen Loh, Lester Mackey, Leonardo de Moura, Dan Roberts, D. Sculley, Terence Tao, David Balduzzi, Simon Coyle, Alex Gerko, Ryan Holbrook, Addison Howard, and XTX Mark...
2024
-
[17]
The problem of the problem.New directions for methodology of social and behavioral science: Question framing and response consistency, 11:37–49, 1982
Jacob W Getzels. The problem of the problem.New directions for methodology of social and behavioral science: Question framing and response consistency, 11:37–49, 1982
1982
-
[18]
Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai, 2025
Elliot Glazer, Ege Erdil, Tamay Besiroglu, Diego Chicharro, Evan Chen, Alex Gunning, Caroline Falkman Olsson, Jean-Stanislas Denain, Anson Ho, Emily de Oliveira Santos, Olli Järviniemi, Matthew Barnett, Robert Sandler, Matej Vrzala, Jaime Sevilla, Qiuyu Ren, Elizabeth Pratt, L...
2025
-
[19]
A survey on llm-as-a-judge, 2025
Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, Saizhuo Wang, Kun Zhang, Yuanzhuo Wang, Wen Gao, Lionel Ni, and Jian Guo. A survey on llm-as-a-judge, 2025
2025
-
[20]
Putnam-axiom: A functional and static benchmark for measuring higher level mathematical reasoning in llms, 2025
Aryan Gulati, Brando Miranda, Eric Chen, Emily Xia, Kai Fronsdal, Bruno Dumont, Elyas Obbad, and Sanmi Koyejo. Putnam-axiom: A functional and static benchmark for measuring higher level mathematical reasoning in llms, 2025
2025
-
[21]
Measuring mathematical problem solving with the math dataset, 2021
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the math dataset, 2021
2021
-
[22]
Beauty is not simplicity: An analysis of mathematicians’ proof appraisals†.Philosophia Mathematica, 23(1):87–109, 07 2014
Matthew Inglis and Andrew Aberdein. Beauty is not simplicity: An analysis of mathematicians’ proof appraisals†.Philosophia Mathematica, 23(1):87–109, 07 2014
2014
-
[23]
Ai mathematical olympiad - progress prize 1
XTX Investments. Ai mathematical olympiad - progress prize 1. https://kaggle.com/ competitions/ai-mathematical-olympiad-prize, 2024. Kaggle
2024
-
[24]
Hongchao Jiang, Yiming Chen, Yushi Cao, Hung yi Lee, and Robby T. Tan. Codejudgebench: Benchmarking llm-as-a-judge for coding tasks, 2025
2025
-
[25]
Douglas B. Lenat. Am, an artificial intelligence approach to discovery in mathematics as heuristic search. 1976
1976
-
[26]
The psychology of curiosity: A review and reinterpretation.Psychological Bulletin, 116(1):75–98, 1994
George Loewenstein. The psychology of curiosity: A review and reinterpretation.Psychological Bulletin, 116(1):75–98, 1994
1994
-
[27]
Advanced version of gemini with deep think officially achieves gold-medal standard at the international mathematical olympiad, July 2025
Thang Luong and Edward Lockhart. Advanced version of gemini with deep think officially achieves gold-medal standard at the international mathematical olympiad, July 2025. DeepMind blog post
2025
-
[28]
American mathematics competitions (AMC)
Mathematical Association of America. American mathematics competitions (AMC). https: //maa.org/math-competitions. Accessed: 2025-09-06
2025
-
[29]
Mathematical conjecture genera- tion using machine intelligence, 2023
Challenger Mishra, Subhayan Roy Moulik, and Rahul Sarkar. Mathematical conjecture genera- tion using machine intelligence, 2023
2023
-
[30]
Shubhra Mishra, Gabriel Poesia, and Noah D. Goodman. From next-token to mathematics: The learning dynamics of mathematical reasoning in language models, 2025
2025
-
[31]
What is a problem that we may solve it?Synthese, pages 85–118, 1981
Thomas Nickles. What is a problem that we may solve it?Synthese, pages 85–118, 1981
1981
-
[32]
Alexander Novikov, Ngân V˜u, Marvin Eisenberger, Emilien Dupont, Po-Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, Borislav Kozlovskii, Francisco J. R. Ruiz, Abbas Mehrabian, M. Pawan Kumar, Abigail See, Swarat Chaudhuri, George Holland, Alex Davies, Sebastian Nowozin, Pushm...
2025
-
[33]
Introducing OpenAI o3 and o4-mini
OpenAI. Introducing OpenAI o3 and o4-mini. https://openai.com/index/ introducing-o3-and-o4-mini/, April 2025. Accessed: 2025-08-29
2025
-
[34]
Prolific
Stefan Palan and Christian Schitter. Prolific. ac—a subject pool for online experiments.Journal of behavioral and experimental finance, 17:22–27, 2018
2018
-
[35]
Gabriel Poesia, David Broman, Nick Haber, and Noah D. Goodman. Learning formal mathe- matics from intrinsic motivation, 2024
2024
-
[36]
From calculation to adjudication: Examining llm judges on mathematical reasoning tasks, 2025
Andreas Stephan, Dawei Zhu, Matthias Aßenmacher, Xiaoyu Shen, and Benjamin Roth. From calculation to adjudication: Examining llm judges on mathematical reasoning tasks, 2025
2025
-
[37]
Putnambench: Evaluating neural theorem-provers on the putnam mathematical competition, 2024
George Tsoukalas, Jasper Lee, John Jennings, Jimmy Xin, Michelle Ding, Michael Jennings, Amitayush Thakur, and Swarat Chaudhuri. Putnambench: Evaluating neural theorem-provers on the putnam mathematical competition, 2024
2024
-
[38]
Amplifying human performance in combinatorial competitive programming, 2024
Petar Veliˇckovi´c, Alex Vitvitskyi, Larisa Markeeva, Borja Ibarz, Lars Buesing, Matej Ba- log, and Alexander Novikov. Amplifying human performance in combinatorial competitive programming, 2024
2024
-
[39]
Springer Berlin Heidelberg, Berlin, Heidelberg, 2009
Cédric Villani.Optimal Transport: Old and New, volume 338 ofGrundlehren der Mathematis- chen Wissenschaften. Springer Berlin Heidelberg, Berlin, Heidelberg, 2009
2009
-
[40]
Meta-reasoning: Deciding which game to play, which problem to solve, and when to quit
Lionel Wong, Tracey Mills, Ionatan Kuperwajs, Katherine M Collins, and Tom Griffiths. Meta-reasoning: Deciding which game to play, which problem to solve, and when to quit. In Proceedings of the Annual Meeting of the Cognitive Science Society, volume 47, 2025
2025
-
[41]
Xing, Hao Zhang, Joseph E
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena, 2023. 9 Appendix A1 Additional Results 11...
2023
-
[42]
Mean pairwise r
We sample at two temperatures and also plot the95%confidence interval. A1.1.2 Wasserstein Distance Tables In Tables 3, 4, and 5, we report the Wasserstein distances between the LLM and human distributions for interestingness and difficulty. A1.1.3 LRM Judgment Lengths We exami...
2024
Reviewed August 3, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.