REVIEW 2 major objections 5 minor 68 references
Adding an economically irrelevant passage to a graduate microeconomics problem lowers language models' chance of a correct final answer by 12.3 percentage points, even as the models keep producing fluent, coherent explanations and rate the
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 01:47 UTC pith:X3URG3VL
load-bearing objection A careful, unusually transparent LLM field experiment whose headline result (a within-task 12pp red-herring accuracy drop) is solid, but whose construct validity rests on the withheld test items; worth refereeing, not desk-rejecting. the 2 major comments →
Wrong and More Confident: A Field Experiment on Large Language Models Taking a Graduate Economics Exam
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that an economically irrelevant passage appended to a graduate-level microeconomics problem causes a causal, within-task drop in language models' probability of answering correctly, from a control mean of about 0.587 to roughly 0.49 on the treated versions, an estimated effect of β1 = −0.1234 (se 0.0267) in a linear probability model with model and task fixed effects. This drop is not accompanied by any detectable change in whether the model gives an explanation, whether the final answer follows from the shown reasoning, or whether the answer stays stable across repeated waves. The model does not notice the corruption: asked to rate difficulty, it rates the red-h
What carries the argument
The Graduate Economic Reasoning Benchmark (GERB): sixty graduate-level microeconomics problems, each with one verified answer and a step-by-step reference solution, each fielded in four versions—with or without a red herring, and with or without an explanation request—to thirty-eight language models, five times each. The within-task design is the load-bearing mechanism: task fixed effects absorb everything common to a problem's four versions, so the red-herring coefficient is identified from the contrast between the treated and control versions of the same task. Two-way clustering by task and model, plus robustness checks (alternative clustering, judge-graded vs. rule-graded cells, wild clus
Load-bearing premise
The load-bearing premise is that the red-herring passages are genuinely economically irrelevant and leave the operative mechanism and verified answer unchanged, and that the reference solutions are correct—but the 60 items and solutions are withheld, so an independent reader cannot verify this.
What would settle it
Have a panel of economists audit all 60 items and their red-herring versions without knowing which condition produced which estimate; if any passage changes the operative mechanism or any reference solution is incorrect, the treatment contrast would measure a difficulty change rather than distraction, and the 12.3-point estimate would not mean what the title claims. A simpler check: if re-inserting the same passage after the model has committed to an answer produces no accuracy drop, the mechanism is not corruption of reasoning.
If this is right
- A headline accuracy figure on hard, expert-level material can mix recall of familiar patterns with genuine construction of answers; a content-preserving perturbation separates the two.
- Asking a model to explain its reasoning raises accuracy by about 6 percentage points but does not narrow the red-herring gap, so explanation helps generally without conferring distraction-resistance.
- Reasoning ability raises overall accuracy but does not reduce the accuracy cost of the red herring; it changes only whether the model repeats the same wrong answer across attempts.
- The 12.3-point drop is broad: 37 of 38 models are less accurate under the red herring, and the effect survives alternative clustering, graders, and sample restrictions.
- Open-weight and closed-weight models reach similar accuracy, but open-weight models do so at substantially lower cost per correct answer.
Where Pith is reading between the lines
- Inference: If the effect generalizes beyond microeconomics, the same within-task red-herring design could be used to audit LLM reasoning in law, medicine, and software, where a fluent but wrong derivation is more dangerous than a refusal.
- Inference: The metacognitive inversion—rating the corrupted task as easier while answering it wrong more often—suggests that a model's self-reported difficulty is a poor alarm for reasoning failure; interfaces that rely on model confidence to flag uncertain outputs would miss this failure mode.
- Inference: Because the benchmark content is withheld, the published design cannot by itself be used to measure future models' absolute accuracy; a natural extension would be to release a held-out subset to a trusted auditor to confirm the construct validity of the red-herring passages.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper constructs a new benchmark, GERB, consisting of 60 graduate-level microeconomics problems with verified answers and reference solutions. Each problem is fielded in a 2×2 design: with or without an inserted 'red herring' passage, and with or without an explanation request. Thirty-eight LLMs answer all four versions of every problem in five independent temperature-zero waves, yielding 45,402 scored responses. The headline linear probability model (Eq. 1) with model and task fixed effects estimates a red-herring effect of β1 = −0.1234 (se 0.0267) on final-answer correctness, equivalent to a 12.3 percentage-point drop. The effect is robust across the specifications in Table 11, including a wild cluster bootstrap over 60 task clusters (p<0.001) and a deterministic rule-graded subset. Secondary results are that explanation requests improve accuracy and do not interact with the red herring; reasoning type does not moderate the accuracy loss; the red herring reduces reasoning correctness while leaving explanation, coherence, and consistency roughly unchanged; models rate red-herring versions as slightly easier; and open-weight models attain similar accuracy at lower cost per correct answer. The 60 problem statements, reference solutions, and red-herring passages are withheld from public release to prevent contamination.
Significance. The paper's core design is strong: within-task, within-model contrasts with fixed effects identify the treatment effect without relying on published benchmark scores; the five-wave panel permits cell-level aggregation; and the wild cluster bootstrap addresses the small number of model clusters. The deterministic rule-graded robustness row and the two-judge agreement analysis are also genuine strengths, as are the public analysis code, result tables, and API verification tests. If the treatment really is an irrelevant passage, the paper provides a clean causal estimate of distraction on expert-level economic reasoning — a meaningful contribution to the measurement literature. The main unresolved issue is construct validity: the claim that each herring leaves the operative mechanism and verified answer unchanged is asserted but not audited because the corpus is withheld. This is a fixable problem, so the appropriate decision is major revision rather than rejection.
major comments (2)
- [§2, footnote 1; Abstract] The treatment definition — 'It looks relevant and is not needed to solve the problem, and it leaves the verified final answer and the operative mechanism unchanged' — is the construct-validity assumption on which the entire headline claim rests. With task fixed effects in Eq. (1), β1 identifies the average effect of adding the author's passage on final-answer accuracy, but it cannot distinguish 'irrelevant text distracts reasoning' from 'the rewritten task is harder or different.' If even some herring passages alter the constraint set, the applicable equilibrium, or the operative mechanism, the −0.1234 estimate would mix the intended distraction effect with a difficulty/content effect. Footnote 1 and the abstract state that the 60 items and reference solutions are withheld, so the irrelevance property is currently an assertion rather than an audited property. This is not a demonstrated e
- [§8, Abstract, title] The 'More Confident' component of the headline is supported by a marginally significant estimate: the red-herring difficulty-rating change is −0.181 (se 0.100, p = 0.071). The text appropriately says the finding 'rests on the direction of the effect rather than its size,' but the abstract states without qualification that 'The red herring also leads a model to rate a problem as easier than its clean version,' and contribution (iii) describes a 'metacognitive inversion.' A two-sided p of 0.071 is not enough to carry the second half of the title unless the test was pre-specified or additional evidence is supplied. Please temper the abstract/title or add supporting robustness (e.g., ordinal rating models, a pre-registered one-sided test, or split-half stability).
minor comments (5)
- [§11, Table 11] The rule-graded versus judge-graded comparison (β1 = −0.210 vs −0.103) is presented as evidence against a lenient or confused judge, but the two subsamples are not randomly assigned: rule-graded cells are short exact-answer responses, while judge-graded cells are longer/open responses. The difference may reflect item composition. Please soften the causal interpretation of this comparison.
- [Eq. (2)] The centered rating term is written as γ δ^c_qm, using γ for both the model fixed effects and the coefficient on the rating. Use a distinct symbol (e.g., λ) for the rating coefficient.
- [Table 5, consistency column] The Type 3 interaction (+0.032, p<0.001) is estimated with only five model clusters. Table 10 already warns that Type-3 clustered standard errors are unreliable; the same caveat should appear with Table 5, where this result is used to claim differences across reasoning types.
- [§2] 'The tasks with a red herring form the treated group and the tasks without one the control group' reads as a between-task comparison; the design is within-task versions. Rephrase to avoid ambiguity.
- [§7, Table 7] The rating analysis is based on 43,028 responses, while the rest of the paper uses 45,402; clarify why 2,374 responses lack ratings (the text mentions 'rated pairs' but does not reconcile the N).
Circularity Check
No significant circularity: the headline estimate is identified from within-task, within-model contrasts, and no load-bearing step reduces to its own inputs.
full rationale
The central claim, β1 = −0.1234 (se 0.0267) from Eq. (1), is a linear-probability estimate of answer correctness on the red-herring indicator with model and task fixed effects. Identification comes from the within-task, within-model contrast between treated and control versions of the same problem, so the estimate does not reduce to a fitted parameter or to the definition of the treatment. Section 2 defines a red herring as a passage that 'leaves the verified final answer and the operative mechanism unchanged,' but this is a construct-validity assumption about the instrument, not an equation that mechanically forces the estimated effect; the paper measures the accuracy drop empirically rather than assuming it. No load-bearing step is justified by a self-citation: the within-subject design cites List (2025), related work cites external sources, and the author's own benchmark is the data-generating instrument rather than an imported theorem. The disclosed limitations — domain scope, router-based deployment, model drift, and the withheld problem content in footnote 1 — are auditability and external-validity caveats, not circular reductions. The robustness battery, including the rule-graded subsample and wild cluster bootstrap, further supports that the headline is an empirical contrast rather than a restatement of inputs. Therefore no circularity step is present.
Axiom & Free-Parameter Ledger
free parameters (3)
- Model fixed effects (38) =
38 intercepts, not individually reported
- Task fixed effects (60) =
60 intercepts, not individually reported
- Reasoning-type labels (Type 1/2/3) =
17 / 16 / 5 models
axioms (6)
- domain assumption Each red-herring passage is economically irrelevant and leaves the verified answer and operative mechanism unchanged
- domain assumption The verified final answers and reference solutions are correct and complete
- domain assumption API calls are stateless: each response is an independent draw with no carry-over
- domain assumption LLM-as-judge grading is a valid measurement of answer and reasoning correctness
- domain assumption Wave-level consistency is interpretable despite temperature being non-binding for some model families
- standard math Standard cluster-robust linear-probability-model inference with 60 task and 38 model clusters is valid
read the original abstract
A red herring, an irrelevant passage added to a problem, corrupts a language model's reasoning and, through it, its final answer, while the form of the response survives untouched. The benchmark, called the Graduate Economic Reasoning Benchmark (GERB), is sixty graduate-level microeconomics problems, each a detailed setup with a verified final answer and a step-by-step reference solution. Each problem has two versions, one with the red herring and one without, and each of those is asked in two ways, one requesting an explanation and one not. This is a within-subject $2\times2$ factorial experimental design. Thirty-eight language models answer all four versions of every problem. The clean problems (the control group) are already hard, with the models answering under sixty percent correctly on average. The red herring lowers the probability of a correct final answer by 12.3 percentage points, about a quarter of the models' mean accuracy of 0.525. The damage is largest on the problems the model rates as easy. Reasoning ability confers no protection, as the red herring's effect does not differ detectably across models with and without reasoning ability. It does change how the failure looks, since a model with no reasoning mode repeats one wrong answer across waves while a reasoning model wavers. The red herring also leads a model to rate a problem as easier than its clean version, while answering it wrong more often. Although open- and closed-weight models reach the same accuracy, the open-weight models reach it at a substantially lower cost per correct final answer. The form of the response is preserved even as its substance fails. The model still produces an explanation (explanation given), the final answer still follows from the reasoning shown (coherence), and, in the aggregate, it remains the same across waves (consistency).
Figures
Reference graph
Works this paper leans on
-
[1]
and Koller, Alexander , title =
Bender, Emily M. and Koller, Alexander , title =. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages =. 2020 , publisher =
2020
-
[2]
, title =
Brynjolfsson, Erik and Li, Danielle and Raymond, Lindsey R. , title =. Quarterly Journal of Economics , volume =. 2025 , doi =
2025
-
[3]
On the Measure of Intelligence , year =
Chollet, Fran. On the Measure of Intelligence , year =
-
[4]
arXiv preprint arXiv:2110.14168 , year =
Cobbe, Karl and Kosaraju, Vineet and Bavarian, Mohammad and Chen, Mark and Jun, Heewoo and Kaiser, Lukasz and Plappert, Matthias and Tworek, Jerry and Hilton, Jacob and Nakano, Reiichiro and Hesse, Christopher and Schulman, John , title =. arXiv preprint arXiv:2110.14168 , year =
-
[5]
Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics (
Deng, Chunyuan and Zhao, Yilun and Tang, Xiangru and Gerstein, Mark and Cohan, Arman , title =. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics (. 2024 , url =
2024
-
[6]
and Sanyal, Soumya and Welleck, Sean and Ren, Xiang and Ettinger, Allyson and Harchaoui, Zaid and Choi, Yejin , title =
Dziri, Nouha and Lu, Ximing and Sclar, Melanie and Li, Xiang Lorraine and Jiang, Liwei and Lin, Bill Yuchen and West, Peter and Bhagavatula, Chandra and Bras, Ronan Le and Hwang, Jena D. and Sanyal, Soumya and Welleck, Sean and Ren, Xiang and Ettinger, Allyson and Harchaoui, Zaid and Choi, Yejin , title =. Advances in Neural Information Processing Systems...
2023
-
[7]
Science , volume =
Eloundou, Tyna and Manning, Sam and Mishkin, Pamela and Rock, Daniel , title =. Science , volume =. 2024 , doi =
2024
-
[8]
Journal of Economic Perspectives , volume =
Frederick, Shane , title =. Journal of Economic Perspectives , volume =. 2005 , doi =
2005
-
[9]
2026 , url =
Grady, Pat and Huang, Sonya , title =. 2026 , url =
2026
-
[10]
Nature Communications , volume =
Griot, Maxime and Hemptinne, Coralie and Vanderdonckt, Jean and Yuksel, Demet , title =. Nature Communications , volume =. 2025 , doi =
2025
-
[11]
Nature , volume =
Guo, Daya and Yang, Dejian and Zhang, Haowei and Song, Junxiao and Zhang, Ruoyu and Xu, Runxin and Zhu, Qihao and Ma, Shirong and Wang, Peiyi and Bi, Xiao and others , title =. Nature , volume =. 2025 , note =
2025
-
[12]
2026 , url =
Hassabis, Demis , title =. 2026 , url =
2026
-
[13]
Proceedings of the International Conference on Learning Representations (
Hendrycks, Dan and Burns, Collin and Basart, Steven and Zou, Andy and Mazeika, Mantas and Song, Dawn and Steinhardt, Jacob , title =. Proceedings of the International Conference on Learning Representations (. 2021 , url =
2021
-
[14]
arXiv preprint arXiv:2207.05221 , year =
Kadavath, Saurav and Conerly, Tom and Askell, Amanda and Henighan, Tom and Drain, Dawn and Perez, Ethan and Schiefer, Nicholas and Hatfield-Dodds, Zac and DasSarma, Nova and Tran-Johnson, Eli and others , title =. arXiv preprint arXiv:2207.05221 , year =
-
[15]
Econometrica , volume =
Kaji, Tetsuya and Manresa, Elena and Pouliot, Guillaume , title =. Econometrica , volume =. 2023 , doi =
2023
-
[16]
2022 , howpublished =
LeCun, Yann , title =. 2022 , howpublished =
2022
-
[17]
2026 , url =
LeCun, Yann , title =. 2026 , url =
2026
-
[18]
and Mattar, Marcelo G
Li, Ji-An and Xiong, Hua-Dong and Wilson, Robert C. and Mattar, Marcelo G. and Benna, Marcus K. , title =. Advances in Neural Information Processing Systems (. 2025 , url =
2025
-
[19]
arXiv preprint arXiv:2205.14334 , year =
Lin, Stephanie and Hilton, Jacob and Evans, Owain , title =. arXiv preprint arXiv:2205.14334 , year =
-
[20]
Signs of Introspection in Large Language Models , howpublished =
Lindsey, Jack and. Signs of Introspection in Large Language Models , howpublished =. 2025 , url =
2025
-
[21]
Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Short Papers) , pages =
Magar, Inbal and Schwartz, Roy , title =. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Short Papers) , pages =. 2022 , publisher =
2022
-
[22]
2026 , url =
Marcus, Gary , title =. 2026 , url =
2026
-
[23]
arXiv preprint arXiv:2410.05229 , year =
Mirzadeh, Iman and Alizadeh, Keivan and Shahrokhi, Hooman and Tuzel, Oncel and Bengio, Samy and Farajtabar, Mehrdad , title =. arXiv preprint arXiv:2410.05229 , year =
-
[24]
2024 , howpublished =
Learning to Reason with. 2024 , howpublished =
2024
-
[25]
arXiv preprint arXiv:2501.14249 , year =
Phan, Long and Gatti, Alice and Han, Ziwen and Li, Nathaniel and others , title =. arXiv preprint arXiv:2501.14249 , year =
-
[26]
Advances in Neural Information Processing Systems (
Qu, Yuxiao and Zhang, Tianjun and Garg, Naman and Kumar, Aviral , title =. Advances in Neural Information Processing Systems (. 2024 , url =
2024
-
[27]
Rein, David and Hou, Betty Li and Stickland, Asa Cooper and Petty, Jackson and Pang, Richard Yuanzhe and Dirani, Julien and Michael, Julian and Bowman, Samuel R. , title =. arXiv preprint arXiv:2311.12022 , year =
-
[28]
ACM Computing Surveys , volume =
Shi, Haizhou and Xu, Zihao and Wang, Hengyi and Qin, Weiyi and Wang, Wenyuan and Wang, Yibin and Wang, Zifeng and Ebrahimi, Sayna and Wang, Hao , title =. ACM Computing Surveys , volume =. 2025 , doi =
2025
-
[29]
Steyvers, Mark and Peters, Megan A. K. , title =. Current Directions in Psychological Science , year =
-
[30]
Nature Machine Intelligence , volume =
Steyvers, Mark and Tejeda, Heliodoro and Kumar, Aakriti and Belem, Catarina and Karny, Sheer and Hu, Xinyue and Mayer, Lukas and Smyth, Padhraic , title =. Nature Machine Intelligence , volume =. 2025 , doi =
2025
-
[31]
, title =
Tian, Katherine and Mitchell, Eric and Zhou, Allan and Sharma, Archit and Rafailov, Rafael and Yao, Huaxiu and Finn, Chelsea and Manning, Christopher D. , title =. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (. 2023 , doi =
2023
-
[32]
and Chi, Ed H
Wang, Xuezhi and Wei, Jason and Schuurmans, Dale and Le, Quoc V. and Chi, Ed H. and Narang, Sharan and Chowdhery, Aakanksha and Zhou, Denny , title =. Proceedings of the International Conference on Learning Representations (. 2023 , url =
2023
-
[33]
and Le, Quoc V
Wei, Jason and Wang, Xuezhi and Schuurmans, Dale and Bosma, Maarten and Ichter, Brian and Xia, Fei and Chi, Ed H. and Le, Quoc V. and Zhou, Denny , title =. Advances in Neural Information Processing Systems (. 2022 , url =
2022
-
[34]
and Zhang, Hao and Gonzalez, Joseph E
Zheng, Lianmin and Chiang, Wei-Lin and Sheng, Ying and Zhuang, Siyuan and Wu, Zhanghao and Zhuang, Yonghao and Lin, Zi and Li, Zhuohan and Li, Dacheng and Xing, Eric P. and Zhang, Hao and Gonzalez, Joseph E. and Stoica, Ion , title =. Advances in Neural Information Processing Systems (. 2023 , url =
2023
-
[35]
, title =
List, John A. , title =. 2025 , month = jan, note =
2025
-
[36]
2026 , howpublished =
2026
-
[37]
Provider Routing: Intelligent Multi-Provider Request Routing , year =
-
[38]
2026 , howpublished =
How. 2026 , howpublished =
2026
-
[39]
2026 , howpublished =
Reasoning Tokens: Enhanced. 2026 , howpublished =
2026
-
[40]
2026 , howpublished =
Chat: Inference. 2026 , howpublished =
2026
-
[41]
2026 , howpublished =
Azure. 2026 , howpublished =
2026
-
[42]
2026 , howpublished =
Reasoning Model (. 2026 , howpublished =
2026
-
[43]
2026 , howpublished =
Grounding with. 2026 , howpublished =
2026
-
[44]
2025 , howpublished =
From Precision to Quantization: A Practical Guide to Faster, Cheaper. 2025 , howpublished =
2025
-
[45]
arXiv preprint arXiv:2407.04069 , year =
Laskar, Md Tahmid Rahman and Alqahtani, Sawsan and Bari, M Saiful and Rahman, Mizanur and Khan, Mohammad Abdullah Matin and Khan, Haidar and Jahan, Israt and Bhuiyan, Amran and Tan, Chee Wei and Parvez, Md Rizwan and Hoque, Enamul and Joty, Shafiq and Huang, Jimmy , title =. arXiv preprint arXiv:2407.04069 , year =
-
[46]
Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source
Balloccu, Simone and Schmidtov. Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (. 2024 , url =
2024
-
[47]
Under review at the Conference on Neural Information Processing Systems (
Truong, Sang and Wang, Serena and Haber, Nick and Koyejo, Sanmi , title =. Under review at the Conference on Neural Information Processing Systems (. 2026 , note =
2026
-
[48]
Under review at the Conference on Neural Information Processing Systems (
Truong, Sang and Wang, Serena and Wang, Angelina and Truong, Nhi and Reuel, Anka and Haber, Nick and Koyejo, Sanmi , title =. Under review at the Conference on Neural Information Processing Systems (. 2026 , note =
2026
-
[49]
2026 , howpublished =
Benchmarks Capability Index (. 2026 , howpublished =
2026
-
[50]
, title =
Reuel, Anka and Hardy, Amelia and Smith, Chandler and Lamparth, Max and Hardy, Malcolm and Kochenderfer, Mykel J. , title =. Advances in Neural Information Processing Systems (. 2024 , url =
2024
-
[51]
2025 , howpublished =
Liubimov, Nikolai , title =. 2025 , howpublished =
2025
-
[52]
Demystifying Evals for
Grace, Will and Hadfield, Joe and Olivares, Diego and. Demystifying Evals for. 2026 , howpublished =
2026
-
[53]
Atil, Berk and Aykent, Sarp and Chittams, Alexa and Fu, Lisheng and Passonneau, Rebecca J. and Radcliffe, Evan and Rajagopal, Guru Rajan and Sloan, Adam and Tudrej, Tomasz and Ture, Ferhan and Wu, Zhe and Xu, Lixinyu and Baldwin, Breck , title =. arXiv preprint arXiv:2408.04667 , year =
-
[54]
arXiv preprint arXiv:2405.14782 , year =
Biderman, Stella and Schoelkopf, Hailey and Sutawika, Lintang and Gao, Leo and Tow, Jonathan and others , title =. arXiv preprint arXiv:2405.14782 , year =
-
[55]
Harvard Data Science Review , volume =
Chen, Lingjiao and Zaharia, Matei and Zou, James , title =. Harvard Data Science Review , volume =. 2024 , doi =
2024
-
[56]
and Shadbolt, Nigel and Wooldridge, Michael , title =
La Malfa, Emanuele and Petrov, Aleksandar and Frieder, Simon and Weinhuber, Christoph and Burnell, Ryan and Nazar, Raza and Cohn, Anthony G. and Shadbolt, Nigel and Wooldridge, Michael , title =. Journal of Artificial Intelligence Research , volume =. 2024 , doi =
2024
-
[57]
Shi, Freda and Chen, Xinyun and Misra, Kanishka and Scales, Nathan and Dohan, David and Chi, Ed H. and Sch. Large Language Models Can Be Easily Distracted by Irrelevant Context , booktitle =. 2023 , url =
2023
-
[58]
Advances in Neural Information Processing Systems (NeurIPS) , year =
Naeini, Saeid and Saqur, Raeid and Saeidi, Mozhgan and Giorgi, John and Taati, Babak , title =. Advances in Neural Information Processing Systems (NeurIPS) , year =
-
[59]
arXiv preprint arXiv:2406.11020 , year =
Wang, Yuqing and Zhao, Yun , title =. arXiv preprint arXiv:2406.11020 , year =
-
[60]
arXiv preprint arXiv:2504.01167 , year =
Chen, Yaoyu and Hu, Yuheng and Lu, Yingda , title =. arXiv preprint arXiv:2504.01167 , year =
-
[61]
and Filippas, Apostolos and Manning, Benjamin S
Horton, John J. and Filippas, Apostolos and Manning, Benjamin S. , title =. 2023 , doi =
2023
-
[62]
arXiv preprint arXiv:2501.09686 , year =
Xu, Fengli and Hao, Qianyue and Zong, Zefang and Wang, Jingwei and Zhang, Yunke and Zhang, Jingyi and Lan, Xiaochong and Gong, Jiahui and Ouyang, Tianjian and Meng, Fanjin and others , title =. arXiv preprint arXiv:2501.09686 , year =
-
[63]
Reasoning Tokens , year =
-
[64]
Evans, Jonathan St. B. T. and Stanovich, Keith E. , title =. Perspectives on Psychological Science , volume =. 2013 , doi =
2013
-
[65]
Kahneman, Daniel , title =
-
[66]
2024 , note =
Guo, Yue and Yang, Yi , booktitle =. 2024 , note =
2024
-
[67]
2024 , note =
Quan, Yinzhu and Liu, Zefang , booktitle =. 2024 , note =
2024
-
[68]
Fish, Sara and Shephard, Julia and Li, Minkai and Shorrer, Ran I. and Gonczarowski, Yannai A. , year =. 2503.18825 , archivePrefix =
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.