{"id":"c9df0c95-379f-4078-af3f-3ce4600eefb7","arxiv_id":"2607.22375","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":7,"one_line_summary":"IDEAgent, a multi-agent LLM pipeline with lineage-based diversity archives and repair/refinement, is reported to achieve up to 3.89x higher Yield and 8x more successful topics than baselines across 32 CS topics.","lead":"This paper introduces IDEAgent, a multi-agent LLM system that generates research ideas by explicitly optimizing both quality and diversity through repair, refinement, and lineage archives. It also proposes a joint metric, Yield, and reports that IDEAgent outperforms baselines across 32 computer-science topics.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract's headline ratios are not reproduced by the paper's own tables: '8x more topics' only holds for Yield≥2 (8/1 vs 27/8 at ≥1), and '3.89x' arises from averaging two low-agreement judges; per-judge ratios are ~2.3x.","rationale":"The reader's weakest assumption was the validity of LLM judges as proxies for expert scientific judgment. That is a real and important limitation, and it is partially responsible for the fragility I identify. However, the most load-bearing concern about the central claim is more specific and internal to the paper: the Abstract's headline numbers do not match the paper's own tables. The '8x more topics' claim is only true at Yield≥2, not at non-zero Yield≥1, and the '3.89x' Yield advantage only emerges after averaging two low-agreement judges; neither judge individually produces that ratio. Because these numbers are the paper's primary advertised results, their instability is a direct threat to the central claim. The qualitative conclusion—IDEAgent yields more qualifying diverse ideas than baselines—survives per-judge inspection, so I would not reject the paper. The conditional verdict stands, but the authors should correct the Abstract and prominently report per-judge ratios and the Φ=2 qualifier. I mark agreement as 'partial' because I share the reader's underlying concern about LLM-judge reliability but focus on a narrower, verifiable inconsistency in the reported results.","tokens_in":24913,"tokens_out":10297,"duration_ms":118239,"concrete_test":"Recompute Table 1's headline rows directly from the per-judge Tables 12 and 13, without averaging scores across judges. Specifically: (1) count successful topics at Yield≥1 and Yield≥2 for NOVA and IDEAgent under gate D≥7,S≥7,C≥6,NB≥7, and check whether the 8x ratio appears only at Yield≥2; (2) compute IDEAgent/NOVA Yield ratios separately for each judge. If the Abstract's 'non-zero Yield on 8x more topics' and '3.89x' are not reproduced in either per-judge analysis, the central quantitative claims must be revised to report the per-judge range and the Φ=2 qualifier.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The Abstract's central quantitative claims—'outperforms the best baseline by 3.89x on Yield, while achieving non-zero Yield on 8x more topics'—are not supported by the paper's own reported numbers.\n\nFirst, '8x more topics' is only true for Yield≥2, not for non-zero Yield. In Table 1 at the headline gate (D≥7, S≥7, C≥6, NB≥7), the successful-topic counts for Yield≥1 are NOVA 8/32 vs IDEAgent 27/32, a ratio of 3.4x. The 8x figure appears only in the Yield≥2 row: NOVA 1/32 vs IDEAgent 8/32. Since the paper defines a 'successful topic' via a parameter Φ and reports Φ=1,2,3, the natural reading of 'non-zero Yield' is Φ=1, where the claimed 8x is false.\n\nSecond, the '3.89x' Yield ratio is an artifact of averaging the two external judges' per-idea scores before applying hard thresholds. The per-judge tables (Tables 12 and 13) give IDEAgent/NOVA Yield ratios of 3.531/1.562 = 2.26x (Claude Opus) and 1.219/0.500 = 2.44x (Claude Sonnet). The averaged Table 1 ratio becomes 1.094/0.281 = 3.89x because averaging suppresses the baseline's marginal ideas more than IDEAgent's: NOVA's mean Yield falls from 1.562/0.500 to 0.281, while IDEAgent falls from 3.531/1.219 to 1.094. This is especially fragile given the paper's own inter-judge agreement is weak on soundness (κ=0.268, Table 2).\n\nThese are not cosmetic discrepancies: the Abstract's exact quantitative claims would need to be reworded or qualified. The qualitative direction (IDEAgent > NOVA) survives per-judge inspection, so the work is not invalidated, but the headline numbers as stated are not reproducible from the reported tables.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes IDEAgent, an LLM-based multi-agent system that frames research ideation as a Quality-Diversity (QD) search. IDEAgent generates seed ideas, evaluates them along non-obviousness, soundness, clarity, and diversity using LLM judges, and conditionally applies repair/refinement while maintaining lineage archives and compact memories. The authors also introduce Yield, a set-level metric that counts the largest subset of ideas passing quality thresholds and pairwise diversity constraints. They evaluate on 32 CS topics against four baselines (Stateless, One-Shot, Sequential-Memory, NOVA), reporting that IDEAgent outperforms the best baseline by 3.89x on Yield and achieves non-zero Yield on 8x more topics, with quality improvements attributed to repair/refinement. The code, rubrics, and per-judge tables are included in the paper and appendix.","tokens_in":25479,"tokens_out":10837,"duration_ms":114442,"significance":"If validated, the paper makes a useful contribution: formulating ideation as QD search with lineage-based archives, a joint Yield metric, and an open-source implementation are all practical assets. The appendix's per-judge results, cost analysis, and scaling experiments are welcome. However, the evaluation is entirely LLM-judged and the headline quantitative claims overstate what the paper's own tables show. The qualitative direction (IDEAgent > baselines) survives per-judge inspection, but the reported magnitudes are artifacts of a particular pooling choice and are not robust across judges. With corrected reporting and at least minimal human calibration, the contribution would be publishable; in its current form the evidence does not support the abstract's precise claims.","major_comments":[{"comment":"The claimed 3.89x and 8x improvements are not supported by the per-judge tables. At the headline gate (D≥7, S≥7, C≥6, NB≥7), Table 1 gives Yield 1.094 vs 0.281, but this is a 'jury' result obtained by averaging the two judges' per-idea scores before thresholding. Tables 12 and 13 show per-judge ratios of 3.531/1.562 = 2.26x (Claude Opus) and 1.219/0.500 = 2.44x (Claude Sonnet). Similarly, '8x more topics' is not the non-zero-Yield ratio: Table 1 at Yield≥1 shows 27/32 vs 8/32 = 3.4x; 8x appears only at Yield≥2 (8/32 vs 1/32), and even there per-judge ratios are 1.8x and 5x. The abstract and §7 should report per-judge ranges and unambiguously define the jury averaging procedure.","section":"Abstract, §7, Tables 1, 12, 13"},{"comment":"The central evaluation is entirely LLM-judged. Both the internal qualification/repair decisions and the external Yield metric rely on model scores, with no human expert ratings. Soundness has linear-weighted κ=0.268 and Spearman ρ=0.409 between the two external judges; hard Yield thresholds can amplify judge disagreement into large differences. The jury Yield is much lower than either judge's individual Yield (e.g., NOVA: 0.281 vs 1.562 and 0.500), showing the method's advantage is partly a product of averaging near-threshold scores. I request either (a) a human-evaluated subset of 40-60 ideas with expert scores and Yield recomputed on those scores, or (b) an explicit re-scoping of all conclusions to 'according to the chosen LLM judges' with disagreement-aware intervals. The paper is transparent about this limitation, but it remains load-bearing for the central quantitative claim.","section":"§5, §9 (Limitations), Table 2"},{"comment":"The claim that repair/refinement are 'crucial' is based on small external-judge deltas. For example, repair raises external soundness from 6.39 to 6.88 on a 0-9 scale using only 30 lineages, and refinement raises Yield by +0.78 to +0.88. These differences are of the same order as the inter-judge soundness disagreement (κ=0.268), and no confidence intervals or effect sizes are reported. The internal scores in Fig. 3a are on a 0-100 scale and cannot be compared directly. The paper should report uncertainty (e.g., bootstrap CIs across topics) and paired tests for the external repair/refinement deltas, or temper the claim.","section":"§7, Fig. 3b"}],"minor_comments":[{"comment":"The header spells the system as 'IDEAAgent' while the abstract and body use 'IDEAgent'; please standardize.","section":"Title/header"},{"comment":"'Scaling LLM Reasoning via Reinforcement Learning' is listed twice under cs.AI. If these are two distinct topic instances they should be distinguished; otherwise the dataset contains fewer than 32 unique topics.","section":"Table 3"},{"comment":"The relationship between Table 1's jury Yield and the per-judge Yields should be defined explicitly (per-idea score averaging before thresholding) so readers are not misled by the seemingly inconsistent means across tables.","section":"Tables 1, 12, 13"},{"comment":"The color scale and axis labels are small; also the caption 'Claude-only two-judge Yield surface' is confusing since two Claude models are used, not a single judge. Clarify.","section":"Figures 2 and 4"}],"recommendation":"major_revision","confidential_remarks":"The paper's direction is promising and the open-source release is a strength. The main blockers are the overstated abstract and the absence of any non-LLM validation. I would not reject outright, but major revision should include corrected headline numbers and at least a minimal human-evaluated subset, or an explicit re-scoping of all performance claims to the specific LLM judges used."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the useful stuff: IDEAgent is the first system I've seen that treats research ideation as a Quality-Diversity search with explicit lineage tracking. The design is coherent — compact signatures, separate archives for active, historical, and rejected ideas, and a repair/refine loop that gives targeted feedback. Yield is a sensible joint metric: it forces you to count only mutually distinct ideas that clear quality thresholds. The paper ships code, runs on 32 topics across 8 domains, and the scaling experiments show that simply sampling more fresh ideas does not match IDEAgent's performance. That is real evidence that the search structure matters. The analysis of repair/refinement is thorough, and the authors are transparent about their limitations. Now the soft spots, in order of seriousness. First, the abstract's headline numbers do not reproduce from the paper's own tables. The '3.89x' Yield advantage comes from averaging the two external judges before applying thresholds; per-judge the advantage is about 2.3x. And '8x more topics' only holds for Yield ≥ 2 (1 vs 8 topics), not for non-zero Yield (8 vs 27, which is ~3.4x). That is a real mismatch. The qualitative direction is consistent, but the specific claims need rephrasing or a corrected table. Second, the evaluation rests entirely on LLM judges. The authors acknowledge this, but it is load-bearing: both internal gating and external Yield use model scores, and the inter-annotator agreement on soundness is κ = 0.268. Without some human-anchored validation, 'pursuable research directions' is an unsupported interpretation. A small human sample on a few topics would make a big difference. Third, the closest QD baseline, QDAIF, is discussed in Related Work but never compared. A direct run would strengthen the contribution. There are also many free thresholds, and the headline numbers are picked at one operating point. None of this kills the paper. The framework is clearly specified, the code is out, and the direction of results holds across two independent judges. The citation pattern looks fine; the most relevant prior work is covered. This deserves a serious referee, but the revision needs to fix the abstract, add or explicitly scope human validation, and ideally add QDAIF. I would cite it for the QD formulation and Yield.","headline":"Genuinely new QD-for-ideation formulation with a solid engineering story, but the abstract's headline ratios overstate the results; the qualitative direction holds.","tokens_in":678,"tokens_out":944,"would_cite":true,"duration_ms":41098,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Treating research ideation as a quality-diversity search, with lineage archives and targeted repair, yields four times more pursuable, mutually distinct ideas than the best baseline.","keywords":["Quality-Diversity search","research idea generation","LLM agents","scientific discovery","repair and refinement","lineage archives","yield metric","novelty-diversity tradeoff"],"falsifier":"Have domain experts independently score a sample of the IDEAgent-generated ideas on the same rubrics and recompute Yield with human scores; if the human-based Yield no longer shows IDEAgent ahead of the baselines — or if repair or refinement lowers human-rated soundness — the central claim collapses. A lighter check: run the pipeline on a benchmark where idea soundness can be verified computationally (e.g., known results), and compare judge scores to ground truth.","tokens_in":24764,"feed_emoji":"💡","tokens_out":7468,"duration_ms":67300,"temperature":0.7,"pith_summary":"LLM-based research ideation has been treated as an isolated optimization problem for either quality or diversity, which produces near-duplicate ideas or trivial proposals. This paper argues that ideation should be a Quality-Diversity (QD) search: the goal is a set of ideas that are each non-obvious, sound, and clear, while remaining mutually distinct. The authors build IDEAgent, a multi-agent framework that evolves ideas through lineages, keeps compact summaries of active, historical, and rejected ideas, and repairs or refines drafts that miss thresholds. They also introduce Yield, a joint metric that counts the largest set of mutually diverse ideas passing quality thresholds. Across 32 computer-science topics, IDEAgent achieves 3.89x the Yield of the best baseline at the strictest gate and produces at least one qualifying idea on 27 of 32 topics, versus 8 for the baseline.","feed_headline":"Quality-diversity search yields 4x more AI research ideas","feed_subtitle":"Lineage tracking and targeted repair lift distinct high-quality ideas from 0.28 to 1.09 per topic.","key_machinery":"The load-bearing machinery is the lineage-archive search loop paired with the Yield metric. Each idea is condensed into a five-field signature (problem, mechanism, value-add, assumptions, expected effect); a controller routes candidates through four cases — pure historical match, single-lineage match, below-diversity-floor, or eligible — before one repair or two refinements are allowed. The qualification gate requires non-obviousness, soundness, and clarity scores of at least 60 on a 0–100 scale, no invalid or disputed soundness judgment, and a diversity floor of 60 relative to all archives. Yield operationalizes the QD conjunction: it filters ideas by quality thresholds, then extracts the m","core_discovery":"The central discovery is that the conjunction of quality and diversity — enforced by a lineage-archive search with targeted repair and refinement — is what converts a fixed generation budget into a dense portfolio of distinct, viable research directions. IDEAgent's controller maintains active, historical, and rejected archives, comparing each new idea's compact signature against all prior generations before deciding whether to accept, repair, refine, or reject it. The paper's quantitative claim is that this pipeline, evaluated with the new Yield metric, achieves a mean Yield of 1.094 at the gate NB≥7, S≥7, C≥6, D≥7, compared with 0.281 for the NOVA-inspired baseline — a 3.89x improvement — a","pith_inferences":["Editorial: Because both internal routing and external Yield depend on LLM judges, the 3.89x advantage could shrink or vanish under human expert review if the judges reward stylistic polish rather than genuine logical rigor. A human-scored replication on a sample of the 320 ideas would settle this.","Editorial: The lineage-archive mechanism — logging active, historical, and rejected ideas as compact signatures — mirrors how human researchers record dead ends and superseded hypotheses; the same pattern could generalize to other open-ended search loops such as experiment design or automated molecule discovery.","Editorial: The maximum-clique Yield metric could be adopted as a general-purpose evaluation for generation tasks that value both quality and coverage, potentially replacing averaged novelty scores that near-duplicate outputs can game.","Editorial: The finding that one-shot generation achieved zero Yield suggests that fully shared context during generation contaminates ideas; future systems may need a controlled information flow between generations, a direction the paper leaves implicit."],"forward_implications":["A fixed ideation budget — ten fresh seeds plus at most two auxiliary drafts — can be converted into roughly four times as many distinct, high-quality ideas as independent generation, and the gain grows at stricter non-obviousness thresholds.","Removing the repair/refinement loop (the Sequential-Memory baseline) drops Yield at the strictest gate from 1.094 to 0.281, so the quality-improvement subroutine is the main driver of the advantage.","Yield is model-agnostic and can be recomputed with any judge, including human evaluators, so it offers a game-resistant way to measure the quality-diversity tradeoff in any open-ended generation system.","Because the framework relies on compact signatures rather than full texts, it scales sub-linearly in memory and cost, and the authors argue the same principles should transfer to future models.","Cost per successful diverse portfolio is lower for IDEAgent than for the strongest baseline when at least two diverse qualifying ideas are required, and the gap widens with three."],"fun_headline_variants":["IDEAgent quadruples viable research ideas","Quality-diversity search yields 4x research ideas","Lineage archive lifts AI idea yield 4x","Repair and refine: 4x more distinct ideas","Joint quality-diversity search yields 4x ideas"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The evaluation rests entirely on LLM judges' scores for non-obviousness, soundness, clarity, and pairwise diversity — both the internal gating that decides repair, refinement, and rejection, and the external Yield that reports success; the paper's own inter-judge agreement on soundness is only κ=0.268, and no human validation is provided.","fun_headline_variants_meta":{"raw":{"variants":["IDEAgent quadruples viable research ideas","Quality-diversity search yields 4x research ideas","Lineage archive lifts AI idea yield 4x","Repair and refine: 4x more distinct ideas","Joint quality-diversity search yields 4x ideas"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000197,"raw_usage":{"total_tokens":1235,"prompt_tokens":813,"completion_tokens":422,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":557,"completion_tokens_details":{"reasoning_tokens":348}},"tokens_in":557,"tokens_out":422,"duration_ms":4939,"temperature":1.0,"reasoning_tokens":348,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T04:56:01.797535+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have domain experts independently score a sample of the IDEAgent-generated ideas on the same rubrics and recompute Yield with human scores; if the human-based Yield no longer shows IDEAgent ahead of the baselines — or if repair or refinement lowers human-rated soundness — the central claim collapses. A lighter check: run the pipeline on a benchmark where idea soundness can be verified computationally (e.g., known results), and compare judge scores to ground truth.","supporting_citations":[],"review_version":1}