{"id":"8d0b70eb-b65c-4aa4-ba7e-bda08d3bd536","arxiv_id":"2507.08350","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A controlled study finds that more agents, deeper critique-revision loops, and diverse personas increase the diversity of LLM-generated research ideas, with critic-side diversity best improving feasibility.","lead":"This paper tests how to arrange multiple AI agents in a dialogue to brainstorm research ideas, comparing different numbers of agents, reviewer roles, and rounds of critique. It finds that more agents, more rounds, and diverse expert personas make ideas more varied, and that varied critics make final proposals more feasible, though quality is judged only by another AI.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The feasibility claim rests on a single LLM judge with no human calibration; a self-preference confound and unvalidated scale make the central quality conclusion unsupported as stated.","rationale":"The reader's weakest assumption identified exactly the load-bearing concern: GPT-4 as the sole quality judge with acknowledged bias and circularity. My analysis agrees. The paper's main contribution is empirical guidance for multi-agent ideation design; its diversity claims are grounded in a mechanical deduplication measure and are comparatively secure. The feasibility claim, however, is entirely downstream of an unvalidated LLM preference tournament, and the link is even more fragile than the reader stated because the critique prompt explicitly conditions the generator on the same feasibility criterion the judge later rewards. The abstract's word 'feasibility' overclaims what the evidence shows: a GPT-4 preference signal, not demonstrated practicality. The paper's own limitations section concedes this, which is why the appropriate verdict is CONDITIONAL rather than REJECT: the concern is concrete, testable, and would require human validation or at least a second independent judge to resolve, but the study otherwise is a reproducible, controlled comparison with code and prompts. No other concern seems more load-bearing: the small effect sizes and missing statistical tests are real but secondary, and the abstract/body configuration count mismatch (seven vs. ten) is a reporting inconsistency rather than a threat to the central claim. The concrete test I propose directly targets the weakest link: human rating of a modest sample of pairs, or a swap of the judge model, would settle whether the feasibility result is an artifact of GPT-4 self-preference.","tokens_in":11468,"tokens_out":1576,"duration_ms":16409,"concrete_test":"Run a human evaluation on a stratified sample of, say, 100 proposal pairs (Baseline vs. Diverse Critic; Baseline vs. L=3) with 2–3 expert raters using the same novelty and feasibility rubric, and compare human preferences with the GPT-4 tournament outcome. Alternatively, re-run the tournament with an independent judge (e.g., Claude or Llama) and check whether Diverse Critic still beats Baseline; if the ranking flips or the preference drops to chance, the feasibility claim is judge-specific.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that 'increasing critic-side diversity... further boosts the feasibility of the final proposals' is supported only by GPT-4 preference tournaments (Section 4.2, Section 7). The critique prompt used to generate the ideas explicitly asks reviewers to flag parts 'not feasible for the student to complete the project within two months' (Appendix A), so GPT-4 is effectively judging proposals it was itself prompted to make feasible. This creates a circularity: the evaluator and the generator share the same model family and the same feasibility criterion, so the measured 'feasibility boost' may reflect prompt-aligned self-preference rather than any property a human expert would endorse. The paper acknowledges this in Section 7: 'relying on a single automatic judge introduces model bias and potential circularity.' No human annotation or external benchmark validates the judge. The effect sizes are also small (e.g., Precision@10 rises from 0.47 at Baseline to 0.52 at L=3, and the Diverse Critic win rate is 0.55 against Baseline), with no confidence intervals or significance tests, so the headlining feasibility result could plausibly vanish under human evaluation or even under a different LLM judge. The diversity findings are more robust because they are measured by an embedding-based deduplication ratio, not by the LLM judge, but the novelty/feasibility claims inherit the judge's bias.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a controlled empirical study of multi-agent LLM dialogue design for research ideation. It varies three design axes—agent parallelism (number of critics), interaction depth (number of critique-revision turns), and persona diversity—within an ideation-critique-revision framework. Using GPT-4o-mini to generate ideas across seven AI/NLP topics with 20 seeds per condition, it evaluates output diversity via an embedding-based Non-Duplicate Ratio and output quality via GPT-4 preference tournaments. The paper concludes that larger cohorts, deeper interaction, and broader persona heterogeneity increase diversity, and that adding a domain-specialized critic further increases feasibility. Code and full prompts are released.","tokens_in":11732,"tokens_out":6764,"duration_ms":72438,"significance":"The paper's controlled factorial design and complete prompt/code release are clear strengths, and the diversity results, if confirmed with proper uncertainty quantification, would provide practical guidance for building multi-agent ideation systems. The central feasibility claim, however, depends entirely on a single LLM judge with no human calibration, and Section 5 contains a concrete misreading of Table 1. The contribution is empirical rather than theoretical, and the current evidence does not yet support the abstract's unqualified feasibility conclusion.","major_comments":[{"comment":"The text states that Precision@20 rises from 0.47 at N=2 to 0.50 at N=3 before dropping at N=4, but Table 1 reports Precision@20 values of 0.47, 0.47, and 0.49 for N=2, N=3, and N=4, respectively. The only 0.50 in the N=3 row is Precision@40, not Precision@20. The claim that the precision trend mirrors the Non-Duplicate Ratio trend is therefore not supported by the printed data. In addition, the phrase 'marginally worse than Baseline' assumes an unreported Baseline Precision@20 value of 0.50; the paper should either report this reference value or reframe the comparison. Please correct the text or the table and re-derive the affected conclusions.","section":"§5, Table 1"},{"comment":"The central feasibility claim rests solely on the GPT-4 preference tournament, with no human annotation and no comparison of judge scores to human expert ratings. The critique prompt in Appendix A explicitly instructs the critic to flag parts 'not feasible for the student to complete the project within two months,' and the generator (GPT-4o-mini) and judge (GPT-4) are models from the same family. The measured 'feasibility boost' may therefore reflect the judge's alignment with the prompt-injected feasibility criterion rather than a property that human experts would endorse. Because the abstract and conclusion make an unqualified causal claim ('increasing critic-side diversity ... further boosts the feasibility'), this load-bearing point needs direct support: at minimum a small human evaluation or a calibration of the LLM judge against expert ratings, plus an analysis of whether the judge's preferences are driven by the specific feasibility categories listed in the critique prompt.","section":"§4.2, §7, Appendix A"},{"comment":"No confidence intervals, error bars, or significance tests are reported for any comparison, despite 20 seeds per topic-condition. Differences such as Non-Duplicate Ratio 0.77 vs. 0.80 and Precision@10 0.52 vs. 0.48 are small relative to the seed variance one would expect, so the prose claims of 'clear and largely orthogonal effects' and 'consistently raises diversity' are not statistically grounded. Please provide bootstrap or permutation intervals and test the specific comparisons that support the headline claims.","section":"§5"},{"comment":"The text reports that the specialized critic achieves 'a win rate of 0.55 against Baseline at N=10,' but the value 0.55 in Table 3 is Precision@10, i.e., the fraction of the top-10 proposals that come from the non-Baseline configuration, not the pairwise win rate in the tournament as defined in §4.2. This conflates two different evaluation quantities and should be corrected; the corresponding quality conclusion should be rephrased in terms of the metric actually computed.","section":"§5, Table 3"}],"minor_comments":[{"comment":"The phrase 'largely orthogonal effects' is not supported by any interaction analysis or factorial test; please soften the claim or add an explicit interaction analysis.","section":"§5"},{"comment":"The Non-Duplicate Ratio is computed with a fixed cosine threshold (0.8) and MiniLM embeddings; a brief discussion of the sensitivity of the diversity conclusions to this threshold and embedding choice would strengthen the claims.","section":"§4.1"},{"comment":"The dashes for Single and Baseline in the Precision columns are never explained; please state whether these configurations are undefined by design or serve as a 0.50 reference point.","section":"Tables 1 and 2"},{"comment":"Several qualitative examples described as generated for the bias topic instead illustrate code-generation ideas; please verify that the examples correspond to the intended topic and add a note about their representativeness.","section":"Appendix C"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe one thing to know: the diversity results are probably real, the feasibility results are not supported as stated. The study cleanly isolates three design axes—number of critics, number of refinement turns, and where personas are injected—and measures diversity with an embedding deduplication ratio. That part is reproducible and robust. The quality/feasibility story is carried entirely by GPT-4 preference tournaments against Baseline, with no human annotation, no confidence intervals, and effect sizes like a 0.55 win rate. The paper admits the circularity risk in Section 7, but it may be worse than “potential”: the critique prompt itself instructs reviewers to flag anything not feasible within two months, so the generator is explicitly steered toward the same criterion the judge rewards.\n\nThat is a self-confounded design, and it makes the abstract’s “further boosts feasibility” overreach. What the data show is that a specific LLM judge prefers proposals from a pipeline prompted to make them more feasible. That is an interesting observation, but not evidence of human-judged feasibility.\n\nWhere the paper genuinely helps: the factorial comparison is careful, the configurations are clearly defined, and the code and prompts are public. The finding that critic-side personas produce judged quality gains while proposer-side personas produce diversity is a useful empirical lead. The limitations paragraph is refreshingly direct—many papers would not have stated “single automatic judge introduces model bias and potential circularity.”\n\nSoft spots, in order: (1) no human validation, (2) no significance testing despite 20 seeds—0.77 vs 0.80 could be noise, (3) a misreading of Table 1 (the prose says Precision@20 rises to 0.50 at three critics, but the table shows 0.47), (4) abstract says seven configurations while the body lists ten. These are fixable, but as written the central quality claim should not be taken at face value.\n\nBottom line: worth a serious referee because the empirical setup is transparent and the question is practical. For a practitioner, use the diversity guidance; treat the feasibility guidance as a hypothesis.\n\nRecommendation: send to review with a request for human evaluation, or at least a second judge and bootstrap confidence intervals. I would cite it for the diversity findings and the factorial design.","headline":"Honest factorial study with robust diversity findings and an unsupported feasibility claim that rests on a self-preferring LLM judge.","tokens_in":12302,"tokens_out":2615,"would_cite":true,"duration_ms":28436,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"More agents and deeper debates make LLM research ideas more diverse.","keywords":["multi-agent LLM dialogue","research ideation","ideation-critique-revision","agent parallelism","interaction depth","persona diversity","LLM-as-a-judge","idea diversity"],"falsifier":"Take the same generated proposals and have human NLP researchers rank paired outputs from the baseline self-critique configuration versus the three-parallel-critics configuration; if human rankings show no feasibility advantage for the multi-agent proposals, the paper's central feasibility claim collapses.","tokens_in":11248,"feed_emoji":"💡","tokens_out":5361,"duration_ms":55495,"temperature":0.7,"pith_summary":"This paper tries to establish a practical answer to how multi-agent LLM dialogues should be arranged when the goal is generating new research ideas. It compares a series of configurations built on an ideation–critique–revision loop, varying the number of agents, the number of critique–revision turns, and whether agents carry specialized personas. The central finding is that each of these three levers increases the diversity of the ideas produced, and that placing specialized personas on the critic side specifically makes the final proposals more feasible. If correct, the paper gives system builders a concrete recipe: use about three parallel critics, two to three refinement rounds, and a domain-specialized critic. The authors present the work as empirical guidance rather than a new generation framework.","feed_headline":"More agents and deeper debates make LLM research ideas more diverse","feed_subtitle":"A controlled study of 7,000 GPT-generated proposals shows dialogue design changes novelty and feasibility.","key_machinery":"The ideation–critique–revision loop is the object that carries the argument: one or more LLM proposers generate ideas, one or more critics give constructive feedback, and a reviser updates the proposal, with the cycle repeated a set number of times. The paper manipulates three axes of this loop—agent diversity (domain personas such as Physics-AI or Psychology-AI placed on the critic or on the proposer/reviser), agent parallelism (two, three, or four independent critics whose feedback is aggregated before revision), and interaction depth (two, three, or four sequential critique–revision turns). Each configuration is measured by non-duplicate ratio for diversity and by a GPT-4 preference tournament for quality, which produces Precision@N against a self-critique baseline. The controlled comparisons along these axes are what let the paper attribute changes in output quality to design choices rather than to prompt content.","core_discovery":"Using the research-ideation setup of Si et al. (2025) as a fixed pipeline—paper retrieval, idea generation, embedding-based deduplication, and LLM-as-judge evaluation—the paper replaces single-shot generation with a multi-agent dialogue and measures what changes. It reports three separable effects averaged over seven AI/NLP topics. Growing the number of parallel critics from one to three raises the non-duplicate ratio from 0.77 to 0.80 and keeps precision roughly at parity; a fourth critic adds little. Deepening the critique–revision loop from one to three turns raises the non-duplicate ratio from 0.77 to 0.85 and gives the best precision at rank 10 (0.52), while a fourth turn yields diminishing returns. Injecting a specialized persona into the critic role improves the win rate against baseline to 0.55, while putting the persona in the proposer/reviser role raises diversity to 0.81. The paper concludes that these axes are largely orthogonal and combine best at three critics with two to three refinement turns and a specialized critic.","pith_inferences":["The same factorial result likely transfers to other open-ended creative generation tasks where novelty and feasibility matter, such as product concept generation or experiment design, though this paper only tests research ideation.","Replacing the GPT-4 judge with multiple human expert raters could shift the precision numbers; the diversity findings rest on an embedding-based metric and would probably survive human evaluation more robustly than the feasibility claims.","A testable extension is to randomize which persona is assigned to which critic rather than fixing a single specialized persona, to separate the effect of persona content from the effect of critic heterogeneity itself."],"forward_implications":["System builders get a concrete default: three parallel critics and two to three refinement turns outperform both single-shot generation and self-critique on diversity, with no loss in quality as measured by the paper's judge.","Specializing the critic role is the cheapest route to feasibility gains: a domain persona on the critic side beats the baseline without sacrificing diversity, whereas a persona on the proposer side buys diversity but not a precision advantage.","Returns saturate: a fourth critic or a fourth refinement turn offers negligible gains, so adding agents or dialogue rounds blindly is not supported by the evidence.","Because the three axes show largely independent effects, they can be tuned separately and combined, giving a factorial recipe for designing multi-agent ideation systems beyond the seven tested topics."],"supporting_citations":[{"why":"Supplies the base pipeline (retrieval, generation, deduplication, LLM-judge evaluation), the seven topics, and the single-agent baseline this study modifies.","marker":"Si et al. (2025)"},{"why":"Establishes the multi-agent idea-generation setting and the high redundancy problem that motivates replacing single-shot generation.","marker":"Su et al. (2025)"},{"why":"Provides the self-refine critique–revision mechanism that the dialogue loop is built on.","marker":"Madaan et al. (2023)"},{"why":"Grounds multi-agent debate as prior evidence that interaction improves LLM outputs, a baseline the paper extends to ideation.","marker":"Du et al. (2023)"},{"why":"Supports role-play and persona assignment among LLM agents, the mechanism behind the agent-diversity manipulation.","marker":"Li et al. (2023)"},{"why":"Represents the single-agent scientific idea-generation approach whose quality the multi-agent configurations are compared against.","marker":"Wang et al. (2024b)"}],"fun_headline_variants":["More LLM critics boost research idea diversity","Deeper AI dialogues yield more feasible research ideas","Critic diversity boosts feasibility in AI ideation","Study of 7,000 ideas: more agents, more diversity","Three AI critics beat one for research brainstorming"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes GPT-4 is a valid and impartial judge of idea novelty and feasibility; every quality conclusion in Sections 5 and the abstract comes from GPT-4 preference tournaments with no human annotation, a limitation the authors explicitly acknowledge in Section 7.","fun_headline_variants_meta":{"raw":{"variants":["More LLM critics boost research idea diversity","Deeper AI dialogues yield more feasible research ideas","Critic diversity boosts feasibility in AI ideation","Study of 7,000 ideas: more agents, more diversity","Three AI critics beat one for research brainstorming"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000138,"raw_usage":{"total_tokens":1155,"prompt_tokens":946,"completion_tokens":209,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":562,"completion_tokens_details":{"reasoning_tokens":136}},"tokens_in":562,"tokens_out":209,"duration_ms":3180,"temperature":1.0,"reasoning_tokens":136,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:20:35.836360+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same generated proposals and have human NLP researchers rank paired outputs from the baseline self-critique configuration versus the three-parallel-critics configuration; if human rankings show no feasibility advantage for the multi-agent proposals, the paper's central feasibility claim collapses.","supporting_citations":[],"review_version":1}