REVIEW 3 minor 87 references
This paper claims that for economics meta-analyses, authors find a single-pass AI feedback report more useful than reports from two multi-agent debate tools, even though the debate tools spend far more computation.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 01:14 UTC pith:B6UGAFRJ
load-bearing objection A transparent, pre-registered null that single-pass AI feedback beats two debate tools the authors built and expected to win; the length-normalization caveat is real but disclosed and doesn't sink the paper.
Does Multi-Agent Debate Improve AI Feedback on Research Papers?
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central discovery is a null result with a clear direction: when the same paper gets three fixed-length, identity-masked AI reports - a single model pass, a two-model adversarial audit, and a multi-agent workshop - the paper's authors rank the single pass first, on average, by 0.66 rank points over one debate tool and 0.57 over the other. The result is stable to clustering, worst-case nonresponse, and tie recoding. In a separate exercise, authors who recalled their real journal referee report usually placed it first and never last, while three AI judges almost always placed the human report last, and the AI judge external to all report-writing model families preferred the most expensive t
What carries the argument
The comparison rests on a within-paper ranking design: each paper's author sees three reports of matched length and template, in random order and without tool labels, and ranks them by usefulness for improving the paper. The arms are bundles that vary model family, internal prompting, and number of calls - about 1, 6, and 10, with token costs in a roughly 1:9:30 ratio. Inference uses exact permutation tests on rank differences, a Friedman test, and percentile bootstrap confidence intervals, with the first reply per paper as baseline. The key is that the author, the person who would act on the feedback, is the judge rather than a model.
Load-bearing premise
The load-bearing premise is that normalizing all reports to a common length and template, and running the workshop tool in its deliberately light configuration, preserves the relative usefulness of the configurations; if the longer, full-depth workshop output contained its best insights, the length budget could have suppressed the debate tools' advantage.
What would settle it
Re-run the same three-way comparison without imposing a common word count and with the workshop tool at full depth; if authors then rank the workshop report above the single pass, the original null is an artifact of the length budget rather than a property of debate. Alternatively, show that the normalization pass systematically removed criticisms authors judged most useful, which would break the link between the ranked reports and the tools themselves.
If this is right
- For fixed-length feedback on economics meta-analyses, the extra computation in these two multi-agent configurations did not improve authors' usefulness rankings.
- A researcher wanting a second read on such a draft has a sensible default in the single-pass configuration; the design does not identify the occasions where the more elaborate tools would pay.
- AI judges disagree with authors: the fully external AI judge would have ranked the most expensive tool first, reversing the single-pass preference, so substituting a model for the intended user can flip conclusions about which AI feedback tool is best.
- Authors valued real journal referee feedback over the AI reports, while AI judges ranked human referee feedback last, showing that usefulness depends on who is doing the judging.
- The result does not show that multi-agent debate fails generally; cost caps and length normalization may account for part of it.
Where Pith is reading between the lines
- Inference: If the common-length normalization erased the debate tools' advantage, then a test that lets each configuration speak at its natural length - or at equal dollar cost rather than equal word count - could recover a debate benefit; the paper's design cannot distinguish 'debate doesn't help' from 'debate's help is hard to compress into 1,000 words.'
- Inference: The author-versus-AI-judge divergence suggests usefulness is a private signal: authors weight feasibility, effort, and tacit knowledge of their own paper, which models do not share. Evaluations of AI research aids that rely on model judges should be validated against the intended human user before replacing them.
- Inference: The weak agreement among co-authors ranking the same paper hints that usefulness rankings are noisy even among humans; a larger sample or forced pairwise comparisons could yield sharper estimates of the true ordering.
- Inference: A natural next experiment is the weaker-model comparison proposed in the paper: if debate's value comes from catching a single model's errors, then with a stronger base model the single pass should do relatively better, and running the same design with an older or smaller model would test that substitution.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a pre-registered, within-paper experiment on 55 economics meta-analyses (44 with author rankings). For each paper, three AI feedback reports were generated: a single pass by Claude Opus 4.8, a cross-model adversarial audit (mad-research), and a light multi-agent workshop (paper-workshop Act I). Reports were identity-masked, randomized, and normalized to a common ~1,000-word template. Authors ranked the single pass as most useful, with mean ranks 1.59 versus 2.25 (mad-research) and 2.16 (paper-workshop); the pairwise contrasts survive Holm correction, and an exact permutation test rejects equality across the three arms. The paper also reports that recalled human referee reports were usually ranked first by authors but last by three AI judges, and that an external Gemini judge would have reversed the main ordering of the three arms. The authors interpret the result as evidence against a perceived-usefulness return to extra computation in this fixed-length setting, not as a general verdict on multi-agent debate.
Significance. If the result stands, it is a useful contribution to the empirical literature on test-time compute and LLM-as-a-judge. The design has notable strengths: the study was pre-registered before report generation; the outcome is measured from the papers' authors rather than the tool builders; the analysis uses exact permutation tests, pre-specified robustness checks, a worst-case non-response bound, and a complete replication archive. The fact that the authors' own tools were ranked below the single pass runs against the authors' stated prior and conflicts of interest, which increases credibility. The paper is appropriately cautious: it repeatedly scopes the finding to fixed-length, light configurations and to economics meta-analyses, and it labels the human-referee and AI-judge comparisons as descriptive/exploratory. The result is unlikely to settle the general debate question, but it provides a well-measured data point and a transferable evaluation protocol.
minor comments (3)
- [Section 2.1] The common-length normalization is the most important residual threat to the title's broad phrasing. The manuscript explicitly says it cannot rule out changes in emphasis or in which criticisms survived normalization. Because the single pass starts near the length budget while paper-workshop condenses roughly 800,000 tokens, differential content loss is not implausible. This does not undermine the stated fixed-length claim, which is already carefully scoped, but the abstract's 'Probably not' is slightly broader than the evidence. Consider adding a sentence to the abstract making the fixed-length qualification more prominent and, if feasible, a supplemental content-retention audit on the own-paper subsample.
- [Section 4.4 / Table 5] The human-referee comparison is clearly labeled as descriptive, and the table notes explain why Panels A and B score different objects. This handling is appropriate. A minor suggestion: state explicitly in the text that the recalled-placement analysis rests on 21 self-selected recollections rather than a pre-specified random sample; this is implied but could be made more prominent.
- [Section 4.2 / Table 4] The cost table is clear and helpful. One small clarification would help: the 'API-equivalent dollars' are based on July 2026 list rates that may change; adding a sentence that the qualitative conclusion is insensitive to plausible price movements would preempt a reader concern. This is a presentation point, not a substantive one.
Circularity Check
No significant circularity: the central ranking outcome comes from independent author judgments against pre-registered hypotheses, not from fitted inputs or self-cited derivations.
full rationale
The paper's central claim is an empirical comparison of three AI report configurations. The load-bearing data are usefulness rankings from the authors of 44 economics meta-analyses, elicited under identity masking, pre-registered before any report was generated, and analyzed with permutation tests. There is no equation chain in which an output reduces to an input by definition: the pairwise contrasts (single minus mad-research = -0.66, single minus workshop = -0.57) are computed from observed ranks, not fitted. The authors built two of the three tools and cite their own tool archives for provenance, but the outcome went against their pre-registered expectation, so the self-citations are not being used to force the result. The paper also explicitly discloses that the normalizing rewrite may change emphasis and that the light workshop configuration and fixed length may partly reflect constraints rather than debate itself (Sections 2.1, 2.5 deviation 7, and Conclusion); these are honest limitations, not circular steps. The AI-judge analyses, including the Gemini reversal, are labeled exploratory and do not enter the primary inference. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors' prior work, and no known result is repackaged under new coordinates. The central derivation is therefore self-contained with respect to the outcome data, and the disclosed self-citations are not load-bearing justifications for the finding.
Axiom & Free-Parameter Ledger
free parameters (1)
- Report length budget =
≈1,000 words (single pass 1,096; paper-workshop 1,070)
axioms (4)
- domain assumption Authors' usefulness rankings are a valid measure of feedback quality for improving their own papers.
- domain assumption The normalization and blinding pass preserved the substantive set of criticisms in each report.
- standard math The per-paper ranking permutation tests assume exchangeability of ranks under the null.
- domain assumption The two strata (own and external) can be pooled.
read the original abstract
Probably not, at least for meta-analyses in economics. In a pre-registered, identity-masked, within-paper experiment, the authors of 44 meta-analyses ranked three AI reports on their own paper by usefulness for improving it: a single pass by a frontier model against two multi-agent debate tools we built and expected to win. All reports were held to a common length and template. The authors preferred the single pass, by 0.66 rank points over mad-research (95% CI 0.32 to 1.00) and 0.57 over paper-workshop (0.16 to 0.95), though paper-workshop spent roughly thirty times the tokens. Authors who recalled their journal referee report usually placed it first and never last; in a separate exercise, three AI judges almost always placed the real journal referee report last. Among the three AI reports, Gemini (the judge whose model family wrote none of the reports) would have ranked paper-workshop first in the authors' place, reversing the single-pass preference. The reversal warns against substituting an AI judge for the author. We measure perceived usefulness for finished papers; whether AI should referee papers is a separate question.
Figures
Reference graph
Works this paper leans on
-
[1]
A. Anwar, C. F. Mang, and S. Plaza. Remittances and the labor supply choices of recipient households: Insights from meta-regression analysis. Journal of Economic Surveys, 2026. doi:10.1111/joes.70011
-
[2]
A. Astakhov, T. Havranek, and J. Novak. Firm size and stock returns: A quantitative survey. Journal of Economic Surveys, 33 0 (5): 0 1463--1492, 2019. doi:10.1111/joes.12335
-
[3]
S. Awaworyi Churchill, H. M. Luong, and M. Ugur. Does intellectual property protection deliver economic benefits? a multi-outcome meta-regression analysis of the evidence. Journal of Economic Surveys, 2022. doi:10.1111/joes.12489
- [4]
-
[5]
J. Bajzik, T. Havranek, Z. Irsova, and J. Novak. Does shareholder activism create value? a meta-analysis. Corporate Governance: An International Review, 33 0 (5): 0 1039--1061, 2025. doi:10.1111/corg.12637
-
[6]
G. Bonanno, L. Errico, N. Fiorino, and R. Ricciuti. The impact of government size on corruption: A meta-regression analysis. Journal of Economic Surveys, 2025. doi:10.1111/joes.12672
-
[7]
P. Cala, T. Havranek, Z. Irsova, M. Luskova, J. Matousek, and J. Novak. Financial incentives and performance: A meta-analysis of experiments in economics. Journal of Political Economy Microeconomics, 2026. URL https://meta-analysis.cz/incentives. Forthcoming
2026
-
[8]
A. Cazachevici, T. Havranek, and R. Horvath. Remittances and economic growth: A meta-analysis. World Development, 134: 0 105021, 2020. doi:10.1016/j.worlddev.2020.105021
arXiv 2020
-
[9]
M. Chletsos and A. Sintos. Financial development and income inequality: A meta-analysis. Journal of Economic Surveys, 2023. doi:10.1111/joes.12528
-
[10]
G. Christensen and E. Miguel. Transparency, reproducibility, and the credibility of economics research. Journal of Economic Literature, 56 0 (3): 0 920--980, 2018. doi:10.1257/jel.20171350
-
[11]
N. Cook, F. Bartoš, P. R. D. Bom, et al. Guidance for the use of AI in the meta-analysis of economics research. Journal of Economic Surveys, 2026 a . doi:10.1111/joes.70105
-
[12]
N. Cook, F. Bartoš, P. R. D. Bom, et al. Reporting guidelines for meta-analysis in economics: Updated for AI . Journal of Economic Surveys, 2026 b . doi:10.1111/joes.70116
-
[13]
Q. Dammerer, L. List, M. Rehm, and M. Schnetzer. Macroeconomic effects of a declining wage share: A meta-analysis of the functional income distribution and aggregate demand. Journal of Economic Surveys, 2025. doi:10.1111/joes.12614
-
[14]
M. D'Arcy, T. Hope, L. Birnbaum, and D. Downey. MARG : Multi-agent review generation for scientific papers. arXiv preprint arXiv:2401.04259, 2024
Pith/arXiv arXiv 2024
-
[15]
L. de Batz and E. Kocenda. Financial crime and punishment: A meta-analysis. Journal of Economic Surveys, 2024. doi:10.1111/joes.12580
-
[16]
G. Di Pietro. Studying abroad and earnings: A meta-analysis. Journal of Economic Surveys, 2022. doi:10.1111/joes.12472
-
[17]
S. Donovan, T. de Graaff, H. L. F. de Groot, and C. C. Koopmans. Unraveling urban advantages: A meta-analysis of agglomeration economies. Journal of Economic Surveys, 2024. doi:10.1111/joes.12543
-
[18]
Y. Du, S. Li, A. Torralba, J. B. Tenenbaum, and I. Mordatch. Improving factuality and reasoning in language models through multiagent debate. In Proceedings of the 41st International Conference on Machine Learning (ICML), volume 235 of PMLR, pages 11733--11763, 2024
2024
-
[19]
D. Ehrenbergerova, J. Bajzik, and T. Havranek. When does monetary policy sway house prices? a meta-analysis. IMF Economic Review, 71 0 (2): 0 538--573, 2023. doi:10.1057/s41308-022-00185-5
-
[20]
A. Elminejad, T. Havranek, R. Horvath, and Z. Irsova. Intertemporal substitution in labor supply: A meta-analysis. Review of Economic Dynamics, 51: 0 1095--1113, 2023. doi:10.1016/j.red.2023.10.001
-
[21]
A. Elminejad, T. Havranek, and Z. Irsova. Relative risk aversion: A meta-analysis. Journal of Economic Surveys, 39 0 (5): 0 2315--2333, 2025. doi:10.1111/joes.12689
-
[22]
M. D. Ernst. Permutation methods: A basis for exact inference. Statistical Science, 19 0 (4): 0 676--685, 2004. doi:10.1214/088342304000000396
-
[23]
A. Ferreira-Lopes, P. Linhares, L. F. Martins, and T. N. Sequeira. Quantitative easing and economic growth in Japan : A meta-analysis. Journal of Economic Surveys, 2022. doi:10.1111/joes.12449
-
[24]
M. Filomena and M. Picchio. Retirement and health outcomes in a meta-analytical framework. Journal of Economic Surveys, 2023. doi:10.1111/joes.12527
- [25]
-
[26]
F. Gachi. Assessing offshore wind employment: A systematic meta-analysis of investment and policy impacts in China , Denmark , and the US (2010--2023). Journal of Economic Surveys, 2026. doi:10.1111/joes.70028
-
[27]
J. S. Gans. Can author manipulation of AI referees be welfare improving? NBER Working Paper 34082, National Bureau of Economic Research, 2025
2025
-
[28]
S. Gechert, T. Havranek, Z. Irsova, and D. Kolcunova. Measuring capital-labor substitution: The importance of method choices and publication bias. Review of Economic Dynamics, 45: 0 55--82, 2022. doi:10.1016/j.red.2021.05.003
-
[29]
D. Guarascio, G. Piccirillo, and J. Reljic. Robots vs. workers: Evidence from a meta-analysis. Journal of Economic Surveys, 2025. doi:10.1111/joes.12699
- [30]
-
[31]
T. Havranek and Z. Irsova. research-audit-duel-protocol, 2026 a . URL https://github.com/tjhavranek/research-audit-duel-protocol. doi:10.5281/zenodo.19105954
-
[32]
T. Havranek and Z. Irsova. erc-ai-feedback, 2026 b . URL https://github.com/tjhavranek/erc-ai-feedback. doi:10.5281/zenodo.20829165
-
[33]
T. Havranek and Z. Irsova. mad-research, 2026 c . URL https://github.com/tjhavranek/mad-research. doi:10.5281/zenodo.20829175
-
[34]
T. Havranek and Z. Irsova. paper-workshop, 2026 d . URL https://github.com/tjhavranek/paper-workshop. doi:10.5281/zenodo.20828996
-
[35]
T. Havranek and O. Kokes. Income elasticity of gasoline demand: A meta-analysis. Energy Economics, 47: 0 77--86, 2015. doi:10.1016/j.eneco.2014.11.004
-
[36]
T. Havranek and A. Sokolova. Do consumers really follow a rule of thumb? three thousand estimates from 144 studies say ‘probably not’. Review of Economic Dynamics, 35: 0 97--122, 2020. doi:10.1016/j.red.2019.05.004
-
[37]
T. Havranek, R. Horvath, Z. Irsova, and M. Rusnak. Cross-country heterogeneity in intertemporal substitution. Journal of International Economics, 96 0 (1): 0 100--118, 2015 a . doi:10.1016/j.jinteco.2015.01.012
-
[38]
T. Havranek, Z. Irsova, K. Janda, and D. Zilberman. Selective reporting and the social cost of carbon. Energy Economics, 51: 0 394--406, 2015 b . doi:10.1016/j.eneco.2015.08.009
-
[39]
T. Havranek, R. Horvath, and A. Zeynalov. Natural resources and economic growth: A meta-analysis. World Development, 88: 0 134--151, 2016. doi:10.1016/j.worlddev.2016.07.016
-
[40]
T. Havranek, M. Rusnak, and A. Sokolova. Habit formation in consumption: A meta-analysis. European Economic Review, 95: 0 142--167, 2017. doi:10.1016/j.euroecorev.2017.03.009
-
[41]
T. Havranek, D. Herman, and Z. Irsova. Does daylight saving save electricity? a meta-analysis. The Energy Journal, 39 0 (2): 0 35--61, 2018 a . doi:10.5547/01956574.39.2.thav
-
[42]
T. Havranek, Z. Irsova, and T. Vlach. Measuring the income elasticity of water demand: The importance of publication and endogeneity biases. Land Economics, 94 0 (2): 0 259--283, 2018 b . doi:10.3368/le.94.2.259
-
[43]
T. Havranek, Z. Irsova, and O. Zeynalova. Tuition fees and university enrolment: A meta-regression analysis. Oxford Bulletin of Economics and Statistics, 80 0 (6): 0 1145--1184, 2018 c . doi:10.1111/obes.12240
-
[44]
T. Havranek, T. D. Stanley, H. Doucouliagos, P. Bom, J. Geyer-Klingeberg, I. Iwasaki, W. R. Reed, K. Rost, and R. C. M. van Aert. Reporting guidelines for meta-analysis in economics. Journal of Economic Surveys, 34 0 (3): 0 469--475, 2020. doi:10.1111/joes.12363
-
[45]
T. Havranek, Z. Irsova, L. Laslopova, and O. Zeynalova. Publication and attenuation biases in measuring skill substitution. Review of Economics and Statistics, 106 0 (5): 0 1187--1200, 2024. doi:10.1162/rest_a_01227
-
[46]
P. Heimberger. Do higher public debt levels reduce economic growth? Journal of Economic Surveys, 2023. doi:10.1111/joes.12536
-
[47]
S. Hirsch, T. Petersen, M. Koppenberg, and M. Hartmann. CSR and firm profitability: Evidence from a meta-regression analysis. Journal of Economic Surveys, 2023. doi:10.1111/joes.12523
-
[48]
S. Holm. A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, 6 0 (2): 0 65--70, 1979. URL https://www.jstor.org/stable/4615733
arXiv 1979
-
[49]
N. Horie, I. Iwasaki, O. Kupets, X. Ma, S. Mizobata, and M. Satogami. Wage-experience profiles in China and Eastern Europe : A large meta-analysis. Journal of Economic Surveys, 2025. doi:10.1111/joes.12605
-
[50]
N. Hussain, M. Khan, D. K. Nguyen, A. Stocchetti, and S. Corbet. Board-level governance and corporate social responsibility: A meta-analytic review. Journal of Economic Surveys, 2025. doi:10.1111/joes.12603
-
[51]
J. P. A. Ioannidis. What meta-research has taught us about research and changes to research practices. Journal of Economic Surveys, 39 0 (4): 0 1823--1834, 2025. doi:10.1111/joes.12666
-
[52]
T. Iorngurum. The exchange rate pass-through to domestic prices: A meta-analysis. Journal of Economic Surveys, 2025. doi:10.1111/joes.12647
-
[53]
Z. Irsova, H. Doucouliagos, T. Havranek, and T. D. Stanley. Meta-analysis of social science research: A practitioner's guide. Journal of Economic Surveys, 38 0 (5): 0 1547--1566, 2024. doi:10.1111/joes.12595
-
[54]
S. Jiang, J. Wan, Y. Wang, and G. Xiao. Social support and the adoption of climate-smart agriculture: A meta-analysis. Journal of Economic Surveys, 2026. doi:10.1111/joes.70098. In press
-
[55]
A. Khan, J. Hughes, D. Valentine, L. Ruis, K. Sachan, A. Radhakrishnan, E. Grefenstette, S. R. Bowman, T. Rocktäschel, and E. Perez. Debating with more persuasive LLM s leads to more truthful answers. arXiv preprint arXiv:2402.06782, 2024. ICML 2024
Pith/arXiv arXiv 2024
-
[56]
J. Knaisch and C. Pöschel. Wage response to corporate income taxes: A meta-regression analysis. Journal of Economic Surveys, 2024. doi:10.1111/joes.12557
-
[57]
E. Kocenda and I. Iwasaki. Bank survival around the world: A meta-analytic review. Journal of Economic Surveys, 2022. doi:10.1111/joes.12451
-
[58]
A. Korinek. Generative AI for economic research: Use cases and implications for economists. Journal of Economic Literature, 61 0 (4): 0 1281--1317, 2023. doi:10.1257/jel.20231736
-
[59]
A. Korinek. AI agents for economic research. NBER Working Paper 34202, National Bureau of Economic Research, 2025
2025
-
[60]
K. Kroupova, T. Havranek, and Z. Irsova. Student employment and education: A meta-analysis. Economics of Education Review, 100: 0 102539, 2024. doi:10.1016/j.econedurev.2024.102539
arXiv 2024
-
[61]
T. Liang, Z. He, W. Jiao, X. Wang, Y. Wang, R. Wang, Y. Yang, S. Shi, and Z. Tu. Encouraging divergent thinking in large language models through multi-agent debate. arXiv preprint arXiv:2305.19118, 2023. EMNLP 2024
Pith/arXiv arXiv 2023
-
[62]
W. Liang, Y. Zhang, H. Cao, B. Wang, D. Ding, X. Yang, K. Vodrahalli, S. He, D. Smith, Y. Yin, D. McFarland, and J. Zou. Can large language models provide useful feedback on research papers? a large-scale empirical analysis. NEJM AI, 1 0 (8), 2024. doi:10.1056/AIoa2400196
-
[63]
S. Malovana, M. Hodula, J. Bajzik, and Z. Gric. Bank capital, lending, and regulation: A meta-analysis. Journal of Economic Surveys, 2024. doi:10.1111/joes.12560
-
[64]
S. Malovana, M. Hodula, Z. Gric, and J. Bajzik. Borrower-based macroprudential measures and credit growth: How biased is the existing literature? Journal of Economic Surveys, 2025. doi:10.1111/joes.12608
-
[65]
J. Matousek, T. Havranek, and Z. Irsova. Individual discount rates: A meta-analysis of experimental evidence. Experimental Economics, 25 0 (1): 0 318--358, 2022. doi:10.1007/s10683-021-09716-9
-
[66]
J. Mun, C. Jung, X. Zhou, H. Kim, and M. Sap. GoodPoint : Learning constructive scientific paper feedback from author responses. arXiv preprint arXiv:2604.11924, 2026
Pith/arXiv arXiv 2026
-
[67]
J. Núñez, D. Martín-Barroso, J. A. Núñez-Serrano, and F. J. Velázquez. How much are we willing to pay for quality wine? a meta-analysis and meta-regression analysis. Journal of Economic Surveys, 2025. doi:10.1111/joes.12668
-
[68]
Opatrny, T
M. Opatrny, T. Havranek, Z. Irsova, and M. Scasny. Publication bias and model uncertainty in measuring the effect of class size on achievement. Journal of Labor Economics, 2026. URL https://meta-analysis.cz/class. Forthcoming
2026
-
[69]
A. Panickssery, S. R. Bowman, and S. Feng. LLM evaluators recognize and favor their own generations. arXiv preprint arXiv:2404.13076, 2024. NeurIPS 2024
Pith/arXiv arXiv 2024
-
[70]
P. Pataranutaporn, N. Powdthavee, C. Achiwaranguprok, and P. Maes. Can AI solve the peer review crisis? a large-scale cross-model experiment of LLM s' performance and biases in evaluating over 1,000 economics papers. arXiv preprint arXiv:2502.00070, 2025
Pith/arXiv arXiv 2025
-
[71]
M. Picchio and M. Ubaldi. Unemployment and health: A meta-analysis. Journal of Economic Surveys, 2024. doi:10.1111/joes.12588
-
[72]
K. Saito, A. Wachi, K. Wataoka, and Y. Akimoto. Verbosity bias in preference labeling by large language models. arXiv preprint arXiv:2310.10076, 2023. NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following
Pith/arXiv arXiv 2023
-
[73]
S. Schneider. Do robots boost productivity? a quantitative meta-study. Journal of Economic Surveys, 2026. doi:10.1111/joes.70042
-
[74]
A. Sintos. Population diversity and economic growth: A meta-regression analysis. Journal of Economic Surveys, 2025. doi:10.1111/joes.12681
-
[75]
A. Sintos, M. Chletsos, and A. Xydea. Revisiting the health spending--growth nexus. Journal of Economic Surveys, 2026. doi:10.1111/joes.70095. In press
-
[76]
A. Smit, N. Grinsztajn, P. Duckworth, T. D. Barrett, and A. Pretorius. Should we be going MAD ? a look at multi-agent debate strategies for LLM s. In Proceedings of the 41st International Conference on Machine Learning (ICML), volume 235 of PMLR, pages 45883--45905, 2024
2024
-
[77]
B. Su, N. Collina, G. Wen, D. Li, K. Cho, J. Fan, B. Zhao, and W. Su. How to find fantastic AI papers: Self-rankings as a powerful predictor of scientific impact beyond peer review. arXiv preprint arXiv:2510.02143, 2025
arXiv 2025
-
[78]
P. Valickova, T. Havranek, and R. Horvath. Financial development and economic growth: A meta-analysis. Journal of Economic Surveys, 29 0 (3): 0 506--526, 2015. doi:10.1111/joes.12068
-
[79]
P. Wang, L. Li, L. Chen, Z. Cai, D. Zhu, B. Lin, Y. Cao, Q. Liu, T. Liu, and Z. Sui. Large language models are not fair evaluators. arXiv preprint arXiv:2305.17926, 2023. ACL 2024
Pith/arXiv arXiv 2023
-
[80]
Q. Wang, Z. Wang, Y. Su, H. Tong, and Y. Song. Rethinking the bounds of LLM reasoning: Are multi-agent discussions the key? arXiv preprint arXiv:2402.18272, 2024. ACL 2024
Pith/arXiv arXiv 2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.