{"id":"067e4af3-b2b1-43e4-8a31-9913658e95f2","arxiv_id":"2508.19200","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"An LLM pipeline that recombines mined themes, domains, and methods generates diverse research ideas, with 99.5% of surveyed papers decomposable into these three axes but only 16.4% reconstructible from them.","lead":"This paper builds a modern version of Ramon Llull's medieval combinatorial machine: it mines research papers for themes, domains, and methods, then prompts LLMs to combine them into research ideas. It compares the resulting ideas with prior LLM ideation methods and asks how much of research ideation is combinatorial.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 99.5% decomposability result is not probative: the decomposition prompt supplies the mined element lists and only requires one selection per axis, so near-universal decomposability may reflect list coverage rather than an inherent A+B+C structure.","rationale":"The reader's weakest assumption already identifies the decomposition protocol as the soft spot, and my reading agrees: the 99.5% decomposability is the most load-bearing quantitative claim in the paper, yet the evaluation design makes it nearly tautological. The paper deserves credit for a transparent, clearly written pipeline, a human-pilot element set, and explicit limitations acknowledging that quantitative metrics do not establish scientific merit. However, the abstract's opening claim that the machine produces \"diverse, relevant, and grounded\" ideas and that most ML research is A+B+C-decomposable rests on metrics that are at best partially validated. Since the authors explicitly plan open-sourcing and the fix is a simple control experiment, the appropriate outcome is the same conditional acceptance the reader recommended, with the control as an explicit condition. I would not move the verdict to reject because the pipeline itself is a reasonable baseline and the overclaim is correctable; I would not move to accept because the central empirical support is currently an artifact risk.","tokens_in":17053,"tokens_out":4799,"duration_ms":45035,"concrete_test":"Hold the Section 4.3 / A.8 decomposition protocol fixed and rerun it on a random 200-title sample with control element lists: (a) the mined elements randomly permuted across Theme/Domain/Method, and (b) frequency-matched generic keywords not mined from these conferences. If decomposability remains near 99% under either control, the reported 99.5% is uninformative and the abstract's coverage claim must be weakened. As a secondary check, recompute the reconstruction rate at Jaccard thresholds 0.2 and 0.4 to quantify how much the 16.4% figure depends on the hand-chosen 30% cut-off.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.3 reports 99.5% decomposability as evidence that the three-disk design captures fundamental aspects of machine-learning research ideation. The protocol in Appendix A.8 cannot support that interpretation. The decomposition prompt gives Gemini the full mined lists (682 themes, 633 domains, 866 methods; Table 2) and asks it to \"find the MOST SPECIFIC and ESSENTIAL concepts from these lists\". The success criterion is at least one selected element per disk. With lists this broad and populated by generic terms like \"efficiency\", \"reasoning\", and \"LLMs\", a compliant model can map essentially any title onto the lists; the 99.5% figure mostly measures vocabulary coverage, not whether papers are inherently A+B+C combinations. The reconstruction rate of 16.4% is also threshold-dependent: a title pair is called reconstructible at >=30% token-level Jaccard, a low bar for generic titles, and the 5 candidate titles are generated from the same broad lists. The abstract's companion claim that ideas are \"grounded in current literature\" is not directly measured by any metric reported in Section 4; relevance is lexical BLEU against ACL 2025 titles and is lower than both baselines. The manuscript is transparent about these limitations, but as published the central coverage claim is an artifact risk rather than a demonstrated property.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a combinatorial ideation pipeline inspired by Ramon Llull's Ars combinatoria. It defines three compositional axes — Theme, Domain, and Method — whose elements are mined from top-tier conference papers using Gemini, then combined with templates and rewritten by LLMs into research titles and abstracts. The authors report conference-level statistics of the mined elements, compare their generated ideas against prior ideation systems (Si et al. 2024; Yu et al. 2024) on diversity, similarity, and relevance, and perform a two-stage coverage analysis claiming 99.5% decomposability and 16.4% reconstructibility of research papers into the three-axis framework.","tokens_in":17305,"tokens_out":3359,"duration_ms":31286,"significance":"If the coverage claims were supported, this would be a lightweight, interpretable, and falsifiable baseline for LLM-driven ideation and a useful quantitative perspective on the combinatorial structure of machine-learning research. The paper is transparent about its limitations, plans to open-source code and data, and introduces no fitted parameters, which are notable strengths. However, the empirical support for the headline claims is currently partial: relevance scores in Table 4 are below those of prior systems, no groundedness metric is reported, and the decomposition protocol in Section 4.3 is vulnerable to a list-coverage artifact.","major_comments":[{"comment":"The decomposition protocol supplies the full mined element lists (682 themes, 633 domains, 866 methods; Table 2) and asks Gemini to select \"the MOST SPECIFIC and ESSENTIAL concepts from these lists,\" with success defined as at least one selection per disk. With lists of this breadth and granularity, near-universal decomposability (99.5%) mostly measures vocabulary coverage rather than an inherent A+B+C structure. This undermines the claim that the three-disk design captures fundamental aspects of research ideation. I recommend a control condition: ask the model to propose elements without being shown the lists and then test membership, or compare against randomly sampled element lists, and report the decomposability rate under those conditions.","section":"Section 4.3 and Appendix A.8"},{"comment":"The abstract claims the generated ideas are \"diverse, relevant, and grounded in current literature,\" but Table 4 shows relevance of 0.11 (Top) and 0.05 (Random), well below the 0.28 and 0.18 of the prior systems, and no groundedness metric is reported anywhere. The relevance measure is lexical BLEU against ACL 2025 titles, which is a weak proxy for actual relevance, and the authors' own footnote in Section 4.2 states that the comparison does not suggest superior quality. The abstract should be softened or, preferably, a direct groundedness evaluation should be added — for example, checking whether generated ideas cite or are traceable to a specific relevant paper in the mined corpus.","section":"Abstract and Table 4"},{"comment":"The reconstruction rate of 16.4% is threshold-dependent and the protocol is self-confirming: the five candidate titles are generated from the same elements by the same model that selected them, and the 30% Jaccard threshold is hand-chosen with no sensitivity analysis. Without a control — for example, reconstructing titles from random elements or from elements selected by a different model — the 16.4% figure is difficult to interpret. Please report how the rate varies with the Jaccard threshold and include a random-element baseline.","section":"Section 4.3, reconstruction criterion"},{"comment":"The comparison in Table 4 is not matched on important dimensions: Si et al. (2024) uses human-filtered ideas restricted to seven NLP topics, and Yu et al. (2024) grounds ideas in seed papers, whereas the Ramón Llull variants use unfiltered ACL 2024 elements with no seed grounding. The paper acknowledges this in a footnote, but the table presentation invites direct numerical comparison. A matched condition — e.g., human-filtered Ramón Llull ideas, or topic-constrained sampling — would strengthen the diversity/relevance trade-off analysis.","section":"Section 4.2, Table 4 comparison"}],"minor_comments":[{"comment":"The term \"bijective coverage\" is used for a one-way decomposition followed by a one-way reconstruction; this is not a bijection in the mathematical sense. Please clarify the terminology, e.g., \"bidirectional coverage.\"","section":"Section 4.3"},{"comment":"The table caption does not define \"Similarity.\" The text explains it as average top-K Jaccard similarity with K equal to the number of generated ideas, but this should be stated in the caption for clarity.","section":"Table 4"},{"comment":"The reconstruction prompt instructs the model to generate five diverse titles but does not ask it to reconstruct the original title; reporting the maximum Jaccard over five unconstrained candidates may underestimate reconstructibility. Consider a prompt that explicitly asks for a title close in content to the original paper.","section":"Appendix A.8"},{"comment":"The table contains placeholder entries \"Value A3,\" \"Value B3,\" and \"Value C3\" in the RL Theory row; these appear to be unfinished and should be filled or removed.","section":"Table 8"},{"comment":"There are several typographical inconsistencies, including \"Ram´on\" with inconsistent spacing and the use of non-ASCII apostrophes. A careful proofreading pass is recommended.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest about its limitations and the historical framing is engaging, but the empirical core is not yet at journal strength. The decomposition result as designed is close to a tautology, and the relevance/groundedness claims outrun the evidence. The requested controls and a direct groundedness check are well within the scope of a revision, so I see this as major revision rather than reject."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: this is a clearly written, modest workshop paper that builds a lightweight LLM ideation pipeline around three axes (Theme, Domain, Method) and mines those elements from ICLR/ACL/COLT/COLM papers. The genuinely new parts are the cross-conference element statistics, the resource of mined themes/domains/methods, and the decomposition/reconstruction analysis. The recombination idea itself is not new—Scideator and CHIMERA already do facet recombination—but the three-disk framing and the conference-level comparisons are a reasonable incremental step.\n\nThe paper does several things well. The pipeline is simple and reproducible in spirit, the authors are unusually honest in the limitations appendix (they explicitly say their quantitative metrics are not sufficient to judge scientific merit, that the comparison does not establish superior quality, and that the tool is not meant to attack peer review). That transparency is real and should be credited. The cross-conference statistics, e.g., similar domain counts but more method elements in ICLR than ACL, are interesting and could be useful to the community.\n\nThe soft spots are real but not fatal. The abstract claims ideas are diverse, relevant, and grounded, yet the reported relevance numbers (BLEU 0.11 for Top, 0.05 for Random) are well below the prior systems cited (0.28 and 0.18), and no groundedness metric is reported anywhere. That is an overclaim relative to the evidence. The bigger issue is the 99.5% decomposability result. The stress-test note is correct: the decomposition prompt hands Gemini the full mined lists (682 themes, 633 domains, 866 methods) and asks it to select at least one concept per axis. With lists that broad and populated by generic terms like \"efficiency\" and \"reasoning,\" near-universal decomposability mostly measures vocabulary coverage, not an inherent A+B+C structure. The 16.4% reconstruction rate is also threshold-dependent (30% token Jaccard, a low bar, with candidate titles generated from the same broad lists). The authors do not hide the prompts, so the analysis is inspectable, but the headline claim should be reworded or the experiment redesigned.\n\nOther issues are minor: no error bars, the baselines are sampled for different purposes so the comparison is only suggestive (the authors acknowledge this), and there is no quantitative comparison with Scideator or CHIMERA despite citing them. No code or data is released yet, though a public repo is promised.\n\nWho this is for: people working on LLM-based scientific ideation and AI-for-science tools. A serious referee should engage with it as a workshop-quality contribution that deserves conditional acceptance with revisions. The central idea is defensible and the weaknesses are addressable. I would bring it to a reading group, but I would not cite it yet in my own work until the coverage analysis is reworked and the resource is actually released.","headline":"A modest, transparent workshop paper whose recombination baseline is worth having, but whose headline decomposability number is largely an artifact of the evaluation prompt and should be downgraded.","tokens_in":17887,"tokens_out":1171,"would_cite":false,"duration_ms":12427,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Most research ideas in machine learning decompose into theme, domain, and method.","keywords":["Ramon Llull","Ars combinatoria","automated ideation","large language models","research idea generation","theme-domain-method decomposition","scientific creativity","combinatorial generation"],"falsifier":"Take a random sample of papers from the coverage analysis, erase the mined element lists, and ask human annotators or a fresh LLM to propose Theme/Domain/Method triples for each title without seeing any predefined elements; if the share of titles that decompose drops far below 99.5%, the headline decomposability is an artifact of list-guided selection rather than an intrinsic property of research ideas. A second check is to have experts rate a sample of generated ideas for relevance and groundedness and correlate those ratings with the paper's BLEU and Jaccard scores; a near-zero correlation would show the metric-based evaluation does not measure the qualities claimed.","tokens_in":16802,"feed_emoji":"⚙️","tokens_out":9769,"duration_ms":80430,"temperature":0.7,"pith_summary":"This paper attempts to show that a large part of machine-learning research ideation can be mechanized as combinatorial recombination across three axes: Theme (the motivation, such as efficiency), Domain (the problem setting, such as question answering), and Method (the technique, such as linear attention). It builds a modern version of Ramon Llull's thirteenth-century thinking machine, using three rotating disks of elements plus templates that combine them, and an LLM to rewrite raw combinations into research titles and abstracts. The authors mine elements from expert-written lists and from 7,483 accepted papers at major conferences, then evaluate the generated ideas with lexical diversity and relevance metrics. Their headline result is that 99.5% of paper titles can be decomposed into the three axes, while only 16.4% can be reconstructed from the elements alone, which they interpret as evidence that the axes capture near-universal structural building blocks but not the specific creative spark. If the claim holds, the pipeline offers a lightweight, interpretable baseline for automated ideation and a map of how different research communities distribute their attention across themes, domains, and methods.","feed_headline":"99.5% of machine-learning research ideas fit three axes","feed_subtitle":"Recombining theme, domain, and method with an LLM gives a lightweight baseline for automated ideation.","key_machinery":"The central object is the 'thinking machine' itself: a triple (A, B, C) of element lists—Theme, Domain, Method—together with a set of templates T that specify how elements combine, and an LLM that rewrites the raw combination into a polished research idea. The machinery does three jobs: element mining (extracting A/B/C elements and templates from paper titles and abstracts via an LLM, then merging synonyms), combinatorial generation (sampling or enumerating combinations, optionally filtered by visit counts), and LLM rewriting (converting the raw idea into a title and abstract). The load-bearing evaluation device is the bijective coverage test: decomposition measures whether a title maps onto existing A, B, C elements; reconstruction measures whether those elements, fed to an LLM, can regenerate a title with at least 30% token Jaccard similarity to the original.","core_discovery":"On the paper's own terms, the central discovery is that the combinatorial structure of Ramon Llull's Ars combinatoria can be revived as a competitive, lightweight method for LLM-based research ideation. The authors define three disks of elements—Theme, Domain, Method—and a small set of templates, such as 'we did a in b with c' or 'compare c1 and c2 in b1 under a1'. Elements are harvested either from human experts or mined automatically from conference papers using an LLM, then merged by semantic similarity. Prompting a separate LLM to rewrite raw combinations produces titles and abstracts that, by lexical measures, are more diverse than prior single-pass or community-simulation baselines when elements are sampled randomly, and more similar to accepted ACL 2025 titles when top elements are enumerated. The bijective coverage analysis shows near-universal decomposability (99.5% across 7,483 papers) but limited reconstructibility (16.4% at a 30% Jaccard threshold), which the authors read as evidence that the three axes form a nearly complete descriptive vocabulary for ML research while the specific instantiation of an idea still demands human or model creativity beyond recombination.","pith_inferences":["A list-guided decomposition test may overstate decomposability; an unguided variant that asks annotators to propose A/B/C elements without seeing the mined lists would separate the claim that papers are inherently three-axis from the claim that an LLM can map any title onto a broad menu.","If the compositional model is right, mixing elements across conferences (an ICLR method with an ACL domain under a COLM theme) should produce ideas that expert judges rate as more novel on average than within-conference combinations; this is a testable prediction the paper does not run.","The 16.4% reconstruction ceiling suggests a quantitative definition of the 'non-combinatorial remainder' of an idea, and formalizing that remainder as a fourth axis or as perturbation and negation operators is a natural extension.","The paper's relevance metric is average BLEU against ACL 2025 titles; a semantic embedding-based measure would likely reorder the methods, so the reported diversity–relevance trade-off should be read as specific to lexical overlap."],"forward_implications":["If the 99.5% decomposability figure is accepted, Theme–Domain–Method is a nearly complete descriptive ontology for the surface structure of ML research papers across NLP, vision, and theory venues.","The pipeline provides a reproducible baseline: 'Llull (Top)' enumerates combinations of the most visited elements and achieves the highest similarity to accepted titles, while 'Llull (Random)' samples elements and achieves the highest diversity with lower relevance.","The released element lists and templates let researchers inspect community differences (ICLR yields more method elements than ACL; ACL domain elements are more stable across years than themes or methods) and track how field interests drift.","The decomposition–reconstruction gap delimits what pure recombination can accomplish: the axes supply structural raw material, while the paper's '4th axis', perturbation, and negation categories identify what still escapes the A+B+C frame."],"supporting_citations":[{"why":"Supplies the historical description of Llull's thinking machine that motivates the combinatorial ideation pipeline.","marker":"Borges (1937)"},{"why":"Provides the baseline ideation method and the LLM title-and-abstract rewriting setup that the paper adopts for its comparison.","marker":"Si et al. (2024)"},{"why":"Provides the community-simulation baseline whose 100 generated ideas are compared against Llull variants.","marker":"Yu et al. (2024)"},{"why":"Gemini 2.0 Flash carries out element mining, merging, and the decomposition/reconstruction evaluation, while Gemini 1.5 Pro rewrites raw combinations into ideas.","marker":"Gemini Team et al. (2024)"},{"why":"Defines BLEU, which the paper uses to measure how relevant generated titles are to ACL 2025 paper titles.","marker":"Papineni et al. (2002)"},{"why":"Defines the distinct-1 metric used to quantify diversity of generated idea titles.","marker":"Li et al. (2015)"}],"fun_headline_variants":["Three axes + LLM = diverse research ideas","Llull's combinatorics: automated ideation for AI research","Recombine theme, domain, method to spark ideas","Medieval combinatorics meets LLMs for fresh ideas"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that lexical measures—BLEU relevance to accepted titles, token Jaccard similarity, and distinct-1 diversity—together with a prompt that hands the LLM the element lists are valid proxies for whether research ideas are relevant, novel, and decomposable into Theme–Domain–Method; if those proxies are flawed, the coverage percentages and the comparative claims do not support the abstract's conclusion.","fun_headline_variants_meta":{"raw":{"variants":["Three axes + LLM = diverse research ideas","Llull's combinatorics: automated ideation for AI research","Recombine theme, domain, method to spark ideas","Medieval combinatorics meets LLMs for fresh ideas"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001158,"raw_usage":{"total_tokens":4789,"prompt_tokens":929,"completion_tokens":3860,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":545,"completion_tokens_details":{"reasoning_tokens":3795}},"tokens_in":545,"tokens_out":3860,"duration_ms":29054,"temperature":1.0,"reasoning_tokens":3795,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:53:27.928654+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of papers from the coverage analysis, erase the mined element lists, and ask human annotators or a fresh LLM to propose Theme/Domain/Method triples for each title without seeing any predefined elements; if the share of titles that decompose drops far below 99.5%, the headline decomposability is an artifact of list-guided selection rather than an intrinsic property of research ideas. A second check is to have experts rate a sample of generated ideas for relevance and groundedness and correlate those ratings with the paper's BLEU and Jaccard scores; a near-zero correlation would show the metric-based evaluation does not measure the qualities claimed.","supporting_citations":[],"review_version":1}