Pith. sign in

REVIEW 3 minor 87 references

This paper claims that for economics meta-analyses, authors find a single-pass AI feedback report more useful than reports from two multi-agent debate tools, even though the debate tools spend far more computation.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 01:14 UTC pith:B6UGAFRJ

load-bearing objection A transparent, pre-registered null that single-pass AI feedback beats two debate tools the authors built and expected to win; the length-normalization caveat is real but disclosed and doesn't sink the paper.

arxiv 2607.14713 v1 pith:B6UGAFRJ submitted 2026-07-16 econ.GN cs.CLcs.MAq-fin.EC

Does Multi-Agent Debate Improve AI Feedback on Research Papers?

classification econ.GN cs.CLcs.MAq-fin.EC
keywords multi-agent debateAI feedbackresearch paper reviewmeta-analysisLLM-as-a-judgetest-time computeauthor rankingseconomics
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper asks whether spending more compute on AI feedback - multiple models critiquing each other in debate - buys feedback authors actually find more useful. Across 44 economics meta-analyses, the authors who wrote the papers ranked the cheapest configuration, a single pass by one frontier model, as most useful for improving their paper, ahead of two multi-agent debate tools the study's authors built and expected to win. The gap is about two-thirds of a rank point and survives correction for multiple testing. Extra computation did not translate into perceived usefulness, and an AI judge standing in for the author would have reversed the ranking. The study measures perceived usefulness on already-published papers, not whether AI should referee.

Core claim

The central discovery is a null result with a clear direction: when the same paper gets three fixed-length, identity-masked AI reports - a single model pass, a two-model adversarial audit, and a multi-agent workshop - the paper's authors rank the single pass first, on average, by 0.66 rank points over one debate tool and 0.57 over the other. The result is stable to clustering, worst-case nonresponse, and tie recoding. In a separate exercise, authors who recalled their real journal referee report usually placed it first and never last, while three AI judges almost always placed the human report last, and the AI judge external to all report-writing model families preferred the most expensive t

What carries the argument

The comparison rests on a within-paper ranking design: each paper's author sees three reports of matched length and template, in random order and without tool labels, and ranks them by usefulness for improving the paper. The arms are bundles that vary model family, internal prompting, and number of calls - about 1, 6, and 10, with token costs in a roughly 1:9:30 ratio. Inference uses exact permutation tests on rank differences, a Friedman test, and percentile bootstrap confidence intervals, with the first reply per paper as baseline. The key is that the author, the person who would act on the feedback, is the judge rather than a model.

Load-bearing premise

The load-bearing premise is that normalizing all reports to a common length and template, and running the workshop tool in its deliberately light configuration, preserves the relative usefulness of the configurations; if the longer, full-depth workshop output contained its best insights, the length budget could have suppressed the debate tools' advantage.

What would settle it

Re-run the same three-way comparison without imposing a common word count and with the workshop tool at full depth; if authors then rank the workshop report above the single pass, the original null is an artifact of the length budget rather than a property of debate. Alternatively, show that the normalization pass systematically removed criticisms authors judged most useful, which would break the link between the ranked reports and the tools themselves.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • For fixed-length feedback on economics meta-analyses, the extra computation in these two multi-agent configurations did not improve authors' usefulness rankings.
  • A researcher wanting a second read on such a draft has a sensible default in the single-pass configuration; the design does not identify the occasions where the more elaborate tools would pay.
  • AI judges disagree with authors: the fully external AI judge would have ranked the most expensive tool first, reversing the single-pass preference, so substituting a model for the intended user can flip conclusions about which AI feedback tool is best.
  • Authors valued real journal referee feedback over the AI reports, while AI judges ranked human referee feedback last, showing that usefulness depends on who is doing the judging.
  • The result does not show that multi-agent debate fails generally; cost caps and length normalization may account for part of it.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Inference: If the common-length normalization erased the debate tools' advantage, then a test that lets each configuration speak at its natural length - or at equal dollar cost rather than equal word count - could recover a debate benefit; the paper's design cannot distinguish 'debate doesn't help' from 'debate's help is hard to compress into 1,000 words.'
  • Inference: The author-versus-AI-judge divergence suggests usefulness is a private signal: authors weight feasibility, effort, and tacit knowledge of their own paper, which models do not share. Evaluations of AI research aids that rely on model judges should be validated against the intended human user before replacing them.
  • Inference: The weak agreement among co-authors ranking the same paper hints that usefulness rankings are noisy even among humans; a larger sample or forced pairwise comparisons could yield sharper estimates of the true ordering.
  • Inference: A natural next experiment is the weaker-model comparison proposed in the paper: if debate's value comes from catching a single model's errors, then with a stronger base model the single pass should do relatively better, and running the same design with an older or smaller model would test that substitution.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 3 minor

Summary. The paper reports a pre-registered, within-paper experiment on 55 economics meta-analyses (44 with author rankings). For each paper, three AI feedback reports were generated: a single pass by Claude Opus 4.8, a cross-model adversarial audit (mad-research), and a light multi-agent workshop (paper-workshop Act I). Reports were identity-masked, randomized, and normalized to a common ~1,000-word template. Authors ranked the single pass as most useful, with mean ranks 1.59 versus 2.25 (mad-research) and 2.16 (paper-workshop); the pairwise contrasts survive Holm correction, and an exact permutation test rejects equality across the three arms. The paper also reports that recalled human referee reports were usually ranked first by authors but last by three AI judges, and that an external Gemini judge would have reversed the main ordering of the three arms. The authors interpret the result as evidence against a perceived-usefulness return to extra computation in this fixed-length setting, not as a general verdict on multi-agent debate.

Significance. If the result stands, it is a useful contribution to the empirical literature on test-time compute and LLM-as-a-judge. The design has notable strengths: the study was pre-registered before report generation; the outcome is measured from the papers' authors rather than the tool builders; the analysis uses exact permutation tests, pre-specified robustness checks, a worst-case non-response bound, and a complete replication archive. The fact that the authors' own tools were ranked below the single pass runs against the authors' stated prior and conflicts of interest, which increases credibility. The paper is appropriately cautious: it repeatedly scopes the finding to fixed-length, light configurations and to economics meta-analyses, and it labels the human-referee and AI-judge comparisons as descriptive/exploratory. The result is unlikely to settle the general debate question, but it provides a well-measured data point and a transferable evaluation protocol.

minor comments (3)
  1. [Section 2.1] The common-length normalization is the most important residual threat to the title's broad phrasing. The manuscript explicitly says it cannot rule out changes in emphasis or in which criticisms survived normalization. Because the single pass starts near the length budget while paper-workshop condenses roughly 800,000 tokens, differential content loss is not implausible. This does not undermine the stated fixed-length claim, which is already carefully scoped, but the abstract's 'Probably not' is slightly broader than the evidence. Consider adding a sentence to the abstract making the fixed-length qualification more prominent and, if feasible, a supplemental content-retention audit on the own-paper subsample.
  2. [Section 4.4 / Table 5] The human-referee comparison is clearly labeled as descriptive, and the table notes explain why Panels A and B score different objects. This handling is appropriate. A minor suggestion: state explicitly in the text that the recalled-placement analysis rests on 21 self-selected recollections rather than a pre-specified random sample; this is implied but could be made more prominent.
  3. [Section 4.2 / Table 4] The cost table is clear and helpful. One small clarification would help: the 'API-equivalent dollars' are based on July 2026 list rates that may change; adding a sentence that the qualitative conclusion is insensitive to plausible price movements would preempt a reader concern. This is a presentation point, not a substantive one.

Circularity Check

0 steps flagged

No significant circularity: the central ranking outcome comes from independent author judgments against pre-registered hypotheses, not from fitted inputs or self-cited derivations.

full rationale

The paper's central claim is an empirical comparison of three AI report configurations. The load-bearing data are usefulness rankings from the authors of 44 economics meta-analyses, elicited under identity masking, pre-registered before any report was generated, and analyzed with permutation tests. There is no equation chain in which an output reduces to an input by definition: the pairwise contrasts (single minus mad-research = -0.66, single minus workshop = -0.57) are computed from observed ranks, not fitted. The authors built two of the three tools and cite their own tool archives for provenance, but the outcome went against their pre-registered expectation, so the self-citations are not being used to force the result. The paper also explicitly discloses that the normalizing rewrite may change emphasis and that the light workshop configuration and fixed length may partly reflect constraints rather than debate itself (Sections 2.1, 2.5 deviation 7, and Conclusion); these are honest limitations, not circular steps. The AI-judge analyses, including the Gemini reversal, are labeled exploratory and do not enter the primary inference. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors' prior work, and no known result is repackaged under new coordinates. The central derivation is therefore self-contained with respect to the outcome data, and the disclosed self-citations are not load-bearing justifications for the finding.

Axiom & Free-Parameter Ledger

1 free parameters · 4 axioms · 0 invented entities

The central claim depends on design choices rather than fitted parameters. The only hand-chosen magnitude is the common report length budget; the analysis itself estimates effect sizes from data. Assumptions about the validity of author self-assessment and the fairness of normalization are load-bearing.

free parameters (1)
  • Report length budget = ≈1,000 words (single pass 1,096; paper-workshop 1,070)
    All three arms were normalized to a common template and length budget to prevent format/length from identifying the arm; the budget is a hand-chosen design value that could constrain the multi-agent tools, which ordinarily produce longer output.
axioms (4)
  • domain assumption Authors' usefulness rankings are a valid measure of feedback quality for improving their own papers.
    The outcome variable is perceived usefulness, and the paper explicitly targets the author as the intended user; there is no objective improvement measure.
  • domain assumption The normalization and blinding pass preserved the substantive set of criticisms in each report.
    If shortening systematically removed the multi-agent reports' best points, the null could be an artifact of the fixed-length constraint (Section 2.1).
  • standard math The per-paper ranking permutation tests assume exchangeability of ranks under the null.
    Exact permutation inference over within-paper rank differences is valid under the null of random ranking within each paper.
  • domain assumption The two strata (own and external) can be pooled.
    They pool after a stratum-by-arm interaction test (p=0.97), but the samples come from related networks (the authors serve as associate editors at the Journal of Economic Surveys).

pith-pipeline@v1.3.0-alltime-deepseek · 20928 in / 11136 out tokens · 95592 ms · 2026-08-02T01:14:27.850916+00:00 · methodology

0 comments
read the original abstract

Probably not, at least for meta-analyses in economics. In a pre-registered, identity-masked, within-paper experiment, the authors of 44 meta-analyses ranked three AI reports on their own paper by usefulness for improving it: a single pass by a frontier model against two multi-agent debate tools we built and expected to win. All reports were held to a common length and template. The authors preferred the single pass, by 0.66 rank points over mad-research (95% CI 0.32 to 1.00) and 0.57 over paper-workshop (0.16 to 0.95), though paper-workshop spent roughly thirty times the tokens. Authors who recalled their journal referee report usually placed it first and never last; in a separate exercise, three AI judges almost always placed the real journal referee report last. Among the three AI reports, Gemini (the judge whose model family wrote none of the reports) would have ranked paper-workshop first in the authors' place, reversing the single-pass preference. The reversal warns against substituting an AI judge for the author. We measure perceived usefulness for finished papers; whether AI should referee papers is a separate question.

Figures

Figures reproduced from arXiv: 2607.14713 by Tomas Havranek, Zuzana Irsova.

Figure 1
Figure 1. Figure 1: Cost and author rankings across the three configurations. Author mean rank (lower is more useful) against token cost per paper, on a log scale. 4.3 Robustness The single-pass preference does not depend on how we handle the rankings. Averaging all of each paper’s rankings instead of only the first reply, with the resampling clustered by paper so every paper stays one equal-weight unit, leaves it in place: s… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

87 extracted references · 58 canonical work pages

  1. [1]

    Anwar, C

    A. Anwar, C. F. Mang, and S. Plaza. Remittances and the labor supply choices of recipient households: Insights from meta-regression analysis. Journal of Economic Surveys, 2026. doi:10.1111/joes.70011

  2. [2]

    Astakhov, T

    A. Astakhov, T. Havranek, and J. Novak. Firm size and stock returns: A quantitative survey. Journal of Economic Surveys, 33 0 (5): 0 1463--1492, 2019. doi:10.1111/joes.12335

  3. [3]

    Awaworyi Churchill, H

    S. Awaworyi Churchill, H. M. Luong, and M. Ugur. Does intellectual property protection deliver economic benefits? a multi-outcome meta-regression analysis of the evidence. Journal of Economic Surveys, 2022. doi:10.1111/joes.12489

  4. [4]

    Bajzik, T

    J. Bajzik, T. Havranek, Z. Irsova, and J. Schwarz. Estimating the armington elasticity: The importance of study design and publication bias. Journal of International Economics, 127: 0 103383, 2020. doi:10.1016/j.jinteco.2020.103383

  5. [5]

    Bajzik, T

    J. Bajzik, T. Havranek, Z. Irsova, and J. Novak. Does shareholder activism create value? a meta-analysis. Corporate Governance: An International Review, 33 0 (5): 0 1039--1061, 2025. doi:10.1111/corg.12637

  6. [6]

    Bonanno, L

    G. Bonanno, L. Errico, N. Fiorino, and R. Ricciuti. The impact of government size on corruption: A meta-regression analysis. Journal of Economic Surveys, 2025. doi:10.1111/joes.12672

  7. [7]

    P. Cala, T. Havranek, Z. Irsova, M. Luskova, J. Matousek, and J. Novak. Financial incentives and performance: A meta-analysis of experiments in economics. Journal of Political Economy Microeconomics, 2026. URL https://meta-analysis.cz/incentives. Forthcoming

  8. [8]

    Cazachevici, T

    A. Cazachevici, T. Havranek, and R. Horvath. Remittances and economic growth: A meta-analysis. World Development, 134: 0 105021, 2020. doi:10.1016/j.worlddev.2020.105021

  9. [9]

    Chletsos and A

    M. Chletsos and A. Sintos. Financial development and income inequality: A meta-analysis. Journal of Economic Surveys, 2023. doi:10.1111/joes.12528

  10. [10]

    Christensen and E

    G. Christensen and E. Miguel. Transparency, reproducibility, and the credibility of economics research. Journal of Economic Literature, 56 0 (3): 0 920--980, 2018. doi:10.1257/jel.20171350

  11. [11]

    N. Cook, F. Bartoš, P. R. D. Bom, et al. Guidance for the use of AI in the meta-analysis of economics research. Journal of Economic Surveys, 2026 a . doi:10.1111/joes.70105

  12. [12]

    N. Cook, F. Bartoš, P. R. D. Bom, et al. Reporting guidelines for meta-analysis in economics: Updated for AI . Journal of Economic Surveys, 2026 b . doi:10.1111/joes.70116

  13. [13]

    Dammerer, L

    Q. Dammerer, L. List, M. Rehm, and M. Schnetzer. Macroeconomic effects of a declining wage share: A meta-analysis of the functional income distribution and aggregate demand. Journal of Economic Surveys, 2025. doi:10.1111/joes.12614

  14. [14]

    D'Arcy, T

    M. D'Arcy, T. Hope, L. Birnbaum, and D. Downey. MARG : Multi-agent review generation for scientific papers. arXiv preprint arXiv:2401.04259, 2024

  15. [15]

    de Batz and E

    L. de Batz and E. Kocenda. Financial crime and punishment: A meta-analysis. Journal of Economic Surveys, 2024. doi:10.1111/joes.12580

  16. [16]

    Di Pietro

    G. Di Pietro. Studying abroad and earnings: A meta-analysis. Journal of Economic Surveys, 2022. doi:10.1111/joes.12472

  17. [17]

    Donovan, T

    S. Donovan, T. de Graaff, H. L. F. de Groot, and C. C. Koopmans. Unraveling urban advantages: A meta-analysis of agglomeration economies. Journal of Economic Surveys, 2024. doi:10.1111/joes.12543

  18. [18]

    Y. Du, S. Li, A. Torralba, J. B. Tenenbaum, and I. Mordatch. Improving factuality and reasoning in language models through multiagent debate. In Proceedings of the 41st International Conference on Machine Learning (ICML), volume 235 of PMLR, pages 11733--11763, 2024

  19. [19]

    Ehrenbergerova, J

    D. Ehrenbergerova, J. Bajzik, and T. Havranek. When does monetary policy sway house prices? a meta-analysis. IMF Economic Review, 71 0 (2): 0 538--573, 2023. doi:10.1057/s41308-022-00185-5

  20. [20]

    Elminejad, T

    A. Elminejad, T. Havranek, R. Horvath, and Z. Irsova. Intertemporal substitution in labor supply: A meta-analysis. Review of Economic Dynamics, 51: 0 1095--1113, 2023. doi:10.1016/j.red.2023.10.001

  21. [21]

    Elminejad, T

    A. Elminejad, T. Havranek, and Z. Irsova. Relative risk aversion: A meta-analysis. Journal of Economic Surveys, 39 0 (5): 0 2315--2333, 2025. doi:10.1111/joes.12689

  22. [22]

    M. D. Ernst. Permutation methods: A basis for exact inference. Statistical Science, 19 0 (4): 0 676--685, 2004. doi:10.1214/088342304000000396

  23. [23]

    Ferreira-Lopes, P

    A. Ferreira-Lopes, P. Linhares, L. F. Martins, and T. N. Sequeira. Quantitative easing and economic growth in Japan : A meta-analysis. Journal of Economic Surveys, 2022. doi:10.1111/joes.12449

  24. [24]

    Filomena and M

    M. Filomena and M. Picchio. Retirement and health outcomes in a meta-analytical framework. Journal of Economic Surveys, 2023. doi:10.1111/joes.12527

  25. [25]

    Friedman

    M. Friedman. The use of ranks to avoid the assumption of normality implicit in the analysis of variance. Journal of the American Statistical Association, 32 0 (200): 0 675--701, 1937. doi:10.1080/01621459.1937.10503522

  26. [26]

    F. Gachi. Assessing offshore wind employment: A systematic meta-analysis of investment and policy impacts in China , Denmark , and the US (2010--2023). Journal of Economic Surveys, 2026. doi:10.1111/joes.70028

  27. [27]

    J. S. Gans. Can author manipulation of AI referees be welfare improving? NBER Working Paper 34082, National Bureau of Economic Research, 2025

  28. [28]

    Gechert, T

    S. Gechert, T. Havranek, Z. Irsova, and D. Kolcunova. Measuring capital-labor substitution: The importance of method choices and publication bias. Review of Economic Dynamics, 45: 0 55--82, 2022. doi:10.1016/j.red.2021.05.003

  29. [29]

    Guarascio, G

    D. Guarascio, G. Piccirillo, and J. Reljic. Robots vs. workers: Evidence from a meta-analysis. Journal of Economic Surveys, 2025. doi:10.1111/joes.12699

  30. [30]

    Hampl, T

    M. Hampl, T. Havranek, and Z. Irsova. Foreign capital and domestic productivity in the Czech Republic : A meta-regression analysis. Applied Economics, 52 0 (18): 0 1949--1958, 2020. doi:10.1080/00036846.2020.1726864

  31. [31]

    Havranek and Z

    T. Havranek and Z. Irsova. research-audit-duel-protocol, 2026 a . URL https://github.com/tjhavranek/research-audit-duel-protocol. doi:10.5281/zenodo.19105954

  32. [32]

    Havranek and Z

    T. Havranek and Z. Irsova. erc-ai-feedback, 2026 b . URL https://github.com/tjhavranek/erc-ai-feedback. doi:10.5281/zenodo.20829165

  33. [33]

    Havranek and Z

    T. Havranek and Z. Irsova. mad-research, 2026 c . URL https://github.com/tjhavranek/mad-research. doi:10.5281/zenodo.20829175

  34. [34]

    Havranek and Z

    T. Havranek and Z. Irsova. paper-workshop, 2026 d . URL https://github.com/tjhavranek/paper-workshop. doi:10.5281/zenodo.20828996

  35. [35]

    Havranek and O

    T. Havranek and O. Kokes. Income elasticity of gasoline demand: A meta-analysis. Energy Economics, 47: 0 77--86, 2015. doi:10.1016/j.eneco.2014.11.004

  36. [36]

    Havranek and A

    T. Havranek and A. Sokolova. Do consumers really follow a rule of thumb? three thousand estimates from 144 studies say ‘probably not’. Review of Economic Dynamics, 35: 0 97--122, 2020. doi:10.1016/j.red.2019.05.004

  37. [37]

    Havranek, R

    T. Havranek, R. Horvath, Z. Irsova, and M. Rusnak. Cross-country heterogeneity in intertemporal substitution. Journal of International Economics, 96 0 (1): 0 100--118, 2015 a . doi:10.1016/j.jinteco.2015.01.012

  38. [38]

    Havranek, Z

    T. Havranek, Z. Irsova, K. Janda, and D. Zilberman. Selective reporting and the social cost of carbon. Energy Economics, 51: 0 394--406, 2015 b . doi:10.1016/j.eneco.2015.08.009

  39. [39]

    Havranek, R

    T. Havranek, R. Horvath, and A. Zeynalov. Natural resources and economic growth: A meta-analysis. World Development, 88: 0 134--151, 2016. doi:10.1016/j.worlddev.2016.07.016

  40. [40]

    Havranek, M

    T. Havranek, M. Rusnak, and A. Sokolova. Habit formation in consumption: A meta-analysis. European Economic Review, 95: 0 142--167, 2017. doi:10.1016/j.euroecorev.2017.03.009

  41. [41]

    Havranek, D

    T. Havranek, D. Herman, and Z. Irsova. Does daylight saving save electricity? a meta-analysis. The Energy Journal, 39 0 (2): 0 35--61, 2018 a . doi:10.5547/01956574.39.2.thav

  42. [42]

    Havranek, Z

    T. Havranek, Z. Irsova, and T. Vlach. Measuring the income elasticity of water demand: The importance of publication and endogeneity biases. Land Economics, 94 0 (2): 0 259--283, 2018 b . doi:10.3368/le.94.2.259

  43. [43]

    Havranek, Z

    T. Havranek, Z. Irsova, and O. Zeynalova. Tuition fees and university enrolment: A meta-regression analysis. Oxford Bulletin of Economics and Statistics, 80 0 (6): 0 1145--1184, 2018 c . doi:10.1111/obes.12240

  44. [44]

    Havranek, T

    T. Havranek, T. D. Stanley, H. Doucouliagos, P. Bom, J. Geyer-Klingeberg, I. Iwasaki, W. R. Reed, K. Rost, and R. C. M. van Aert. Reporting guidelines for meta-analysis in economics. Journal of Economic Surveys, 34 0 (3): 0 469--475, 2020. doi:10.1111/joes.12363

  45. [45]

    Havranek, Z

    T. Havranek, Z. Irsova, L. Laslopova, and O. Zeynalova. Publication and attenuation biases in measuring skill substitution. Review of Economics and Statistics, 106 0 (5): 0 1187--1200, 2024. doi:10.1162/rest_a_01227

  46. [46]

    Heimberger

    P. Heimberger. Do higher public debt levels reduce economic growth? Journal of Economic Surveys, 2023. doi:10.1111/joes.12536

  47. [47]

    Hirsch, T

    S. Hirsch, T. Petersen, M. Koppenberg, and M. Hartmann. CSR and firm profitability: Evidence from a meta-regression analysis. Journal of Economic Surveys, 2023. doi:10.1111/joes.12523

  48. [48]

    S. Holm. A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, 6 0 (2): 0 65--70, 1979. URL https://www.jstor.org/stable/4615733

  49. [49]

    Horie, I

    N. Horie, I. Iwasaki, O. Kupets, X. Ma, S. Mizobata, and M. Satogami. Wage-experience profiles in China and Eastern Europe : A large meta-analysis. Journal of Economic Surveys, 2025. doi:10.1111/joes.12605

  50. [50]

    Hussain, M

    N. Hussain, M. Khan, D. K. Nguyen, A. Stocchetti, and S. Corbet. Board-level governance and corporate social responsibility: A meta-analytic review. Journal of Economic Surveys, 2025. doi:10.1111/joes.12603

  51. [51]

    J. P. A. Ioannidis. What meta-research has taught us about research and changes to research practices. Journal of Economic Surveys, 39 0 (4): 0 1823--1834, 2025. doi:10.1111/joes.12666

  52. [52]

    Iorngurum

    T. Iorngurum. The exchange rate pass-through to domestic prices: A meta-analysis. Journal of Economic Surveys, 2025. doi:10.1111/joes.12647

  53. [53]

    Irsova, H

    Z. Irsova, H. Doucouliagos, T. Havranek, and T. D. Stanley. Meta-analysis of social science research: A practitioner's guide. Journal of Economic Surveys, 38 0 (5): 0 1547--1566, 2024. doi:10.1111/joes.12595

  54. [54]

    Jiang, J

    S. Jiang, J. Wan, Y. Wang, and G. Xiao. Social support and the adoption of climate-smart agriculture: A meta-analysis. Journal of Economic Surveys, 2026. doi:10.1111/joes.70098. In press

  55. [55]

    A. Khan, J. Hughes, D. Valentine, L. Ruis, K. Sachan, A. Radhakrishnan, E. Grefenstette, S. R. Bowman, T. Rocktäschel, and E. Perez. Debating with more persuasive LLM s leads to more truthful answers. arXiv preprint arXiv:2402.06782, 2024. ICML 2024

  56. [56]

    Knaisch and C

    J. Knaisch and C. Pöschel. Wage response to corporate income taxes: A meta-regression analysis. Journal of Economic Surveys, 2024. doi:10.1111/joes.12557

  57. [57]

    Kocenda and I

    E. Kocenda and I. Iwasaki. Bank survival around the world: A meta-analytic review. Journal of Economic Surveys, 2022. doi:10.1111/joes.12451

  58. [58]

    A. Korinek. Generative AI for economic research: Use cases and implications for economists. Journal of Economic Literature, 61 0 (4): 0 1281--1317, 2023. doi:10.1257/jel.20231736

  59. [59]

    A. Korinek. AI agents for economic research. NBER Working Paper 34202, National Bureau of Economic Research, 2025

  60. [60]

    Kroupova, T

    K. Kroupova, T. Havranek, and Z. Irsova. Student employment and education: A meta-analysis. Economics of Education Review, 100: 0 102539, 2024. doi:10.1016/j.econedurev.2024.102539

  61. [61]

    Liang, Z

    T. Liang, Z. He, W. Jiao, X. Wang, Y. Wang, R. Wang, Y. Yang, S. Shi, and Z. Tu. Encouraging divergent thinking in large language models through multi-agent debate. arXiv preprint arXiv:2305.19118, 2023. EMNLP 2024

  62. [62]

    Liang, Y

    W. Liang, Y. Zhang, H. Cao, B. Wang, D. Ding, X. Yang, K. Vodrahalli, S. He, D. Smith, Y. Yin, D. McFarland, and J. Zou. Can large language models provide useful feedback on research papers? a large-scale empirical analysis. NEJM AI, 1 0 (8), 2024. doi:10.1056/AIoa2400196

  63. [63]

    Malovana, M

    S. Malovana, M. Hodula, J. Bajzik, and Z. Gric. Bank capital, lending, and regulation: A meta-analysis. Journal of Economic Surveys, 2024. doi:10.1111/joes.12560

  64. [64]

    Malovana, M

    S. Malovana, M. Hodula, Z. Gric, and J. Bajzik. Borrower-based macroprudential measures and credit growth: How biased is the existing literature? Journal of Economic Surveys, 2025. doi:10.1111/joes.12608

  65. [65]

    Matousek, T

    J. Matousek, T. Havranek, and Z. Irsova. Individual discount rates: A meta-analysis of experimental evidence. Experimental Economics, 25 0 (1): 0 318--358, 2022. doi:10.1007/s10683-021-09716-9

  66. [66]

    J. Mun, C. Jung, X. Zhou, H. Kim, and M. Sap. GoodPoint : Learning constructive scientific paper feedback from author responses. arXiv preprint arXiv:2604.11924, 2026

  67. [67]

    Núñez, D

    J. Núñez, D. Martín-Barroso, J. A. Núñez-Serrano, and F. J. Velázquez. How much are we willing to pay for quality wine? a meta-analysis and meta-regression analysis. Journal of Economic Surveys, 2025. doi:10.1111/joes.12668

  68. [68]

    Opatrny, T

    M. Opatrny, T. Havranek, Z. Irsova, and M. Scasny. Publication bias and model uncertainty in measuring the effect of class size on achievement. Journal of Labor Economics, 2026. URL https://meta-analysis.cz/class. Forthcoming

  69. [69]

    Panickssery, S

    A. Panickssery, S. R. Bowman, and S. Feng. LLM evaluators recognize and favor their own generations. arXiv preprint arXiv:2404.13076, 2024. NeurIPS 2024

  70. [70]

    Pataranutaporn, N

    P. Pataranutaporn, N. Powdthavee, C. Achiwaranguprok, and P. Maes. Can AI solve the peer review crisis? a large-scale cross-model experiment of LLM s' performance and biases in evaluating over 1,000 economics papers. arXiv preprint arXiv:2502.00070, 2025

  71. [71]

    Picchio and M

    M. Picchio and M. Ubaldi. Unemployment and health: A meta-analysis. Journal of Economic Surveys, 2024. doi:10.1111/joes.12588

  72. [72]

    Saito, A

    K. Saito, A. Wachi, K. Wataoka, and Y. Akimoto. Verbosity bias in preference labeling by large language models. arXiv preprint arXiv:2310.10076, 2023. NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following

  73. [73]

    Schneider

    S. Schneider. Do robots boost productivity? a quantitative meta-study. Journal of Economic Surveys, 2026. doi:10.1111/joes.70042

  74. [74]

    A. Sintos. Population diversity and economic growth: A meta-regression analysis. Journal of Economic Surveys, 2025. doi:10.1111/joes.12681

  75. [75]

    Sintos, M

    A. Sintos, M. Chletsos, and A. Xydea. Revisiting the health spending--growth nexus. Journal of Economic Surveys, 2026. doi:10.1111/joes.70095. In press

  76. [76]

    A. Smit, N. Grinsztajn, P. Duckworth, T. D. Barrett, and A. Pretorius. Should we be going MAD ? a look at multi-agent debate strategies for LLM s. In Proceedings of the 41st International Conference on Machine Learning (ICML), volume 235 of PMLR, pages 45883--45905, 2024

  77. [77]

    B. Su, N. Collina, G. Wen, D. Li, K. Cho, J. Fan, B. Zhao, and W. Su. How to find fantastic AI papers: Self-rankings as a powerful predictor of scientific impact beyond peer review. arXiv preprint arXiv:2510.02143, 2025

  78. [78]

    Valickova, T

    P. Valickova, T. Havranek, and R. Horvath. Financial development and economic growth: A meta-analysis. Journal of Economic Surveys, 29 0 (3): 0 506--526, 2015. doi:10.1111/joes.12068

  79. [79]

    P. Wang, L. Li, L. Chen, Z. Cai, D. Zhu, B. Lin, Y. Cao, Q. Liu, T. Liu, and Z. Sui. Large language models are not fair evaluators. arXiv preprint arXiv:2305.17926, 2023. ACL 2024

  80. [80]

    Q. Wang, Z. Wang, Y. Su, H. Tong, and Y. Song. Rethinking the bounds of LLM reasoning: Are multi-agent discussions the key? arXiv preprint arXiv:2402.18272, 2024. ACL 2024

Showing first 80 references.