{"id":"bfc30d8b-895f-43c8-8348-06b2aaa0579b","arxiv_id":"2509.05072","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"A scalable pipeline constructs Functional Concept Graphs from patents, and the MUSE algorithm samples analogical inspirations that appear to improve creative ideation in a user study.","lead":"The paper builds a large graph of 'purposes' and 'solutions' from 500K patents, then samples related ideas from the graph to inspire people solving everyday problems. A user study suggests these inspirations raise the ratio of creative ideas, though the effect is only clearly significant under the more lenient creativity definition.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Graph-value claim is under-controlled: no random/keyword inspiration baseline, and the headline effect rests on one uncorrected liberal-threshold comparison.","rationale":"The reader's verdict is CONDITIONAL with moderate confidence. My concern is not primarily with graph construction (though NLI edges are noisy) but with the causal attribution in the user study. The headline statistic is a between-condition ratio difference; the minimal condition for attributing it to MUSE is a control that supplies an equal amount of external material without graph semantics. The paper's own introduction motivates MUSE by contrast with keyword search, yet no keyword condition is run. The absence is particularly damaging because all inspiration conditions required reading and processing; the time-course data show empty-condition participants produce more early ideas, consistent with an attention/effort mechanism. Additionally, the significance claim is weaker than stated: only one of four pairwise comparisons is tested, and the strict novelty threshold is marginal. The reader's weakest_assumption (NLI edge quality) would matter if the graph were the only input, but it is prior to the user-study confound; even a perfect graph would not support the conclusion without a random baseline. Therefore the paper should be accepted conditionally, with the baseline study as the explicit condition. This does not change the reader's verdict, but it names the precise experiment that would settle it.","tokens_in":13532,"tokens_out":5053,"duration_ms":59430,"concrete_test":"Conduct a follow-up user study with the same materials, judges, and 15-minute protocol, adding two between-subject conditions: (1) random inspirations: select purpose nodes uniformly at random from the FCG (or random patent titles), and (2) keyword inspirations: top Sentence-BERT matches to the problem text. Pre-register analyses comparing creative ratios at k=2 and k=3 against empty and against the original MUSE conditions, with multiple-comparison correction. If random/keyword conditions match MUSE's 0.75 ratio, the graph-based sampling is not the active ingredient; if they are significantly below, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 7.1 and Table 1 report that all inspiration-based conditions yielded significantly higher creative ratios than empty, with the sentence condition at 75% vs. 49%. However, the experiment has no control in which participants receive inspirations from a non-MUSE source (random patent nodes, keyword search, or unrelated concepts). The three non-empty arms differ only in display format; all draw from the FCG. Without such a baseline, the 26-point ratio gap could reflect the general effect of any external stimulus, demand characteristics, or the extra time spent reading before generating—not the abstraction paths MUSE samples. The only reported inferential test is a t-test between sentence and empty at novelty threshold k=2 (p=0.004); the k=3 result is not significant (p=0.07), and no significance test is reported for the purpose or purpose+mechanism conditions despite the claim that 'all inspiration-based conditions' were significant. Since RQ1 is framed as whether FCG inspirations enhance creativity, not merely whether some stimulus helps, this missing control is the load-bearing gap.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a scalable pipeline for building Functional Concept Graphs (FCGs) from patent data, where nodes are purpose/mechanism clusters and edges encode problem–solution and abstraction relations. The authors also introduce MUSE, an algorithm that samples inspiration nodes from the FCG via \"up–down\" paths, and report a user study with 61 participants on two everyday problems (\"Seal a leak\" and \"Cool a room\"). The central empirical claim is that inspirations from the FCG increase the ratio of creative solutions, with the purpose+mechanism-sentence condition yielding 75% creative solutions versus 49% in the empty condition. The paper releases the graph, code, and data.","tokens_in":13876,"tokens_out":2001,"duration_ms":23956,"significance":"If the central claim holds, the paper would make a useful contribution to computational creativity and design-by-analogy: it scales functional concept graphs to 500K patents, explicitly encodes abstraction relations, and provides a concrete inspiration-sampling mechanism. The release of the graph and pipeline is a valuable resource. The user study design also goes beyond many prior systems by testing the inspirations with human participants and by analyzing trajectory types (NLI, LLM, verb) in Section 7.2. However, the strength of the empirical conclusion is currently limited by the absence of a non-MUSE inspiration baseline and by the incomplete statistical reporting.","major_comments":[{"comment":"The headline RQ1 claim is that FCG inspirations enhance creativity, but the experiment has no control condition in which participants receive inspirations from a non-MUSE source (e.g., random patent nodes, keyword-search results, or unrelated concepts). All three non-empty arms draw from the FCG and differ only in display format. Thus the observed 26-point gap in creative ratio (75% vs. 49%) is consistent with the possibility that any external stimulus, demand characteristics, or the extra time spent reading before generating ideas improves the ratio. A non-MUSE baseline is load-bearing for the paper's central claim and should be added or, if impossible, the claims must be correspondingly restricted.","section":"§7.1, Table 1"},{"comment":"The text states that \"participants in all inspiration-based conditions produced a significantly higher ratio of creative ideas compared to participants in the empty condition,\" but the only reported inferential test is a t-test between the sentence and empty conditions at novelty threshold k=2 (p=0.004). The k=3 test is reported as not significant (p=0.07), and no significance tests are reported for the purpose or purpose+mechanism conditions. To substantiate the \"all conditions\" claim, the authors should report pairwise comparisons (or an overall ANOVA/permutation test) for each condition and threshold, with appropriate multiple-comparison correction. Alternatively, the claim should be softened to what the data actually support.","section":"§7.1"},{"comment":"The abstraction edges—the \"up\" steps in MUSE—are induced by NLI entailment on a single randomly selected representative purpose tag per cluster, using the prefix \"I want\" and a threshold t=0.5 chosen with recall 0.65 and precision 0.9. If the representative tag is noisy, the abstraction step can lead to semantically arbitrary nodes, and the creativity gain could come from any surprising external stimulus rather than meaningful functional abstraction. The paper currently validates the graph only indirectly through the user study outcomes. A direct validation of abstraction-edge quality, or a sensitivity analysis varying the representative tag, prefix, and threshold, would substantially strengthen the claim that MUSE's specific traversal mechanism, not merely graph connectivity, drives the improvement.","section":"§4.3, Appendix C.4"}],"minor_comments":[{"comment":"Typo: \"we rub agglomerative clustering\" should be \"we run agglomerative clustering.\"","section":"Appendix C.3"},{"comment":"The subsection header \"V erb-based connections\" contains an extra space; should be \"Verb-based connections.\"","section":"§4.3"},{"comment":"The text in §7.1 says the sentence condition provided the highest absolute number of feasible solutions, but Table 1 shows the sentence and empty conditions both have 4.8 feasible solutions on average. The wording should acknowledge the tie.","section":"Table 1"},{"comment":"The caption and text describe line styles, but the figure's legend should be checked to ensure the colors/line types match the text (e.g., \"purpose solid blue\" and \"purpose+mechanism sentence dashed green\"). If the figure is rendered in grayscale, line styles alone should be distinguishable.","section":"Figure 3"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nBottom line: this is a substantive systems paper with a suggestive but not yet convincing evaluation. The authors build an FCG from 500K patents, using NLI entailment for abstraction edges plus LLM- and verb-based virtual nodes to boost connectivity, and they release code and data. That is real work and a useful resource for anyone working on analogy mining or ideation support. The MUSE sampler is a reasonable design choice, and the user study as a pilot is worth seeing.\n\nWhere it falls short is the causal claim. RQ1 asks whether FCG inspirations enhance creativity, but every non-empty arm draws from the same FCG. There is no control where participants get random or keyword-based inspirations, so the 75%-vs-49% gap could reflect any external stimulus, extra reading time, or demand characteristics. The within-condition comparison in Table 2, inspired versus non-inspired ideas, is suggestive, but it still does not isolate the graph's specific contribution. And the only significance test reported is for sentence versus empty at k=2 (p=0.004); k=3 is borderline (p=0.07), and the other conditions get no inferential statistics despite being called significant. With 61 participants and 2 problems, that is respectable pilot evidence, not a demonstrated effect.\n\nThe graph itself is evaluated only indirectly, through cluster purity and NLI precision/recall on a small hand-built set. The abstraction edges rest on a representative purpose tag per node and a fixed \"I want\" prefix, which is a real fragility. Still, the authors are transparent about their choices and thresholds, and they ship code and data, so these weaknesses are testable.\n\nMy take: the paper is worth a serious referee and probably worth publishing after revision, but the headline creativity claim needs a randomized non-MUSE baseline and honest multiple-comparison handling. If you care about large-scale functional representations, this is a paper to know.\n\nRecommendation: send it to peer review, expecting a major revision. I would bring it to a reading group to discuss what counts as a control in creativity experiments.","headline":"Substantial graph-building resource with a suggestive but under-controlled user study; worth reviewing, but the creativity claim needs a real baseline.","tokens_in":14305,"tokens_out":2595,"would_cite":true,"duration_ms":29958,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MUSE claims that graph-sampled functional analogies from 500K patents raised users' creative idea ratio from 49% to 75% in a sentence-condition user study.","keywords":["functional concept graph","creative ideation","analogical inspiration","design fixation","patent mining","abstraction","LLM-based annotation","user study"],"falsifier":"Compare the sentence condition against a control that receives the same number of purpose+mechanism sentences sampled at random from the graph, or from non-abstraction paths. If random sentences produce a similar ratio of creative solutions, MUSE's abstraction structure is not the cause of the 75% versus 49% gain; if they do not, the effect is attributable to the FCG paths. A cheaper check is to manually audit 100 sampled NLI abstraction edges and measure precision against human abstraction judgments.","tokens_in":13485,"feed_emoji":"💡","tokens_out":7843,"duration_ms":64371,"temperature":0.7,"pith_summary":"MUSE is a method for building a large graph of functional concepts—purposes and mechanisms—from 500K patents, then using that graph to offer an inventor analogical inspirations for a problem. The paper's central claim is that these inspirations measurably reduce cognitive fixation: in a user study, participants shown purpose-plus-mechanism inspirations phrased as sentences produced creative ideas 75% of the time, versus 49% for participants given no inspirations. The authors argue that explicit abstraction edges, not surface keyword similarity, are what make the sampled analogies useful. If right, the work offers a scalable way to navigate the design space and generate creative options, with the graph itself released for further study.","feed_headline":"Patent-inspired analogies lift creative ideas from 49% to 75%","feed_subtitle":"A 500K-patent functional graph powers MUSE; users with sentence inspirations beat empty-condition users on creative ratio.","key_machinery":"The load-bearing object is the Functional Concept Graph (FCG): purpose nodes from clustered purpose tags, solution nodes from mechanism tags clustered by co-occurrence, and directed edges encoding 'mechanism achieves purpose' and 'purpose is more abstract than purpose'. The abstraction edges come from a natural-language-inference (NLI) entailment model—a classifier that checks whether one sentence is implied by another—run on representative purpose tags with an 'I want' prefix, then cleaned by cycle removal and transitive-edge removal; virtual nodes proposed by an LLM and verb-synonym groups add far links. MUSE samples the classic analogy v-structure—one or two 'up' steps to an abstraction a","core_discovery":"Patents become purpose and mechanism tags; similar purposes form problem nodes, mechanisms cluster by co-occurrence, and problem nodes link through abstraction edges scored by a natural-language entailment model plus LLM and verb virtual nodes. MUSE embeds a user problem, finds its nearest node, and samples 'up-up-down' analogy paths with diversity reranking. In a 61-participant study, the purpose+mechanism-sentence condition yielded 75% creative ideas (k=2) versus 49% with no inspirations, plus the most creative ideas in absolute terms. The paper reads this as evidence that structured functional analogy reduces fixation; the 500K-patent graph is released.","pith_inferences":["If the active ingredient is simply receiving any structured external prompt, a random-purpose control could produce a similar lift; the paper did not include such a condition, so this is an open alternative explanation.","The NLI abstraction edges are computed from one representative tag per cluster; aggregating multiple tags could strengthen edge precision and possibly make the strict novelty threshold significant.","The graph's patents are US, English, and from three CPC sections; a multilingual or cross-domain FCG could reveal far analogies across cultures, directly testing the paper's geographic-bias limitation.","MUSE-inspired ideation could be measured by downstream solution quality rather than human novelty ratings, which would connect the ratio increase to real-world innovation."],"forward_implications":["The released FCG over 500K patents gives other researchers a reusable map of functional analogies, not just a demo corpus.","Design tools can present 'purpose + mechanism sentences' rather than bare keywords; the sentence condition showed the strongest creativity gain.","MUSE-style sampling can be added to LLM prompting pipelines, possibly reducing the online-rehashed solutions the paper observes from a state-of-the-art LLM.","The time-course result implies inspiration tools need a warm-up: users start slower but overtake uninspired users, so evaluation should not be cut short.","The same graph structure can reframe a problem before solution search: going up an abstraction edge changes the problem statement itself, not just candidate answers."],"supporting_citations":[{"why":"Introduces the FCG idea the paper rebuilds; its co-occurrence edges and crowdworker annotation are the baseline being improved.","marker":"(Hope et al., 2022)"},{"why":"Documents design fixation, the cognitive obstacle MUSE targets.","marker":"(Jansson and Smith, 1991)"},{"why":"Supplies GPT-3 few-shot purpose-tag generation used to annotate patent descriptions.","marker":"(Brown et al., 2020)"},{"why":"Provides Sentence-BERT embeddings used to cluster purpose tags and to find the anchor node for a problem.","marker":"(Reimers and Gurevych, 2019)"},{"why":"Supplies the pretrained NLI entailment model that scores abstraction edges between problem nodes.","marker":"(Laurer et al., 2024)"},{"why":"Provides the method for removing cycles from noisy hierarchies while preserving the abstraction structure.","marker":"(Sun et al., 2017)"},{"why":"Supplies MMR, the diversity-based reranking MUSE uses to select up to five inspirations per source.","marker":"(Carbonell and Goldstein, 1998)"},{"why":"Provides the Faiss index used to map a problem text to the nearest node in the graph.","marker":"(Douze et al., 2024)"},{"why":"Defines ideation quality dimensions that the paper's creativity scoring builds on.","marker":"(Reinig et al., 2007)"},{"why":"Frames analogy mining for creativity, the evaluation approach MUSE extends.","marker":"(Hope et al., 2017)"}],"fun_headline_variants":["MUSE: Patent graphs lift creative ideas to 75%","Functional patent analogies boost creative ratio by 26%","MUSE mines patents to spark 75% creative ideas","How a 500K-patent graph breaks fixation with MUSE","Patent analogies: MUSE lifts creativity from 49% to 75%"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The claim collapses if the entailment-based abstraction edges do not faithfully represent real 'is a more general problem' relations; noisy edges would make MUSE's 'up' steps land on arbitrary nodes, and any creativity gain could then come from generic distraction rather than functional analogy.","fun_headline_variants_meta":{"raw":{"variants":["MUSE: Patent graphs lift creative ideas to 75%","Functional patent analogies boost creative ratio by 26%","MUSE mines patents to spark 75% creative ideas","How a 500K-patent graph breaks fixation with MUSE","Patent analogies: MUSE lifts creativity from 49% to 75%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000142,"raw_usage":{"total_tokens":942,"prompt_tokens":617,"completion_tokens":325,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":361,"completion_tokens_details":{"reasoning_tokens":246}},"tokens_in":361,"tokens_out":325,"duration_ms":4204,"temperature":1.0,"reasoning_tokens":246,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T05:36:19.712533+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the sentence condition against a control that receives the same number of purpose+mechanism sentences sampled at random from the graph, or from non-abstraction paths. If random sentences produce a similar ratio of creative solutions, MUSE's abstraction structure is not the cause of the 75% versus 49% gain; if they do not, the effect is attributable to the FCG paths. A cheaper check is to manually audit 100 sampled NLI abstraction edges and measure precision against human abstraction judgments.","supporting_citations":[],"review_version":1}