{"id":"cb80a44c-59b1-438d-b662-d79ad545f92c","arxiv_id":"2608.07243","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Iterative LLM search can reach human-reference-level recipe creativity scores under LLM-based scoring, but in-loop evaluator choice matters more than more iterations or higher temperature.","lead":"This pilot study adapted the FunSearch evolutionary search algorithm to generate Pillsbury Bake-Off recipes with large language models, then scored creativity with LLM judges. It found that the size of the in-loop scorer affected final creativity scores more than iteration count or temperature, pointing to evaluator design as a key lever in creative search.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim is only established under unvalidated LLM evaluation, with Meta-17B serving as both in-loop scorer and final evaluator; a human-rating check is needed to rule out self-preference.","rationale":"The reader's weakest assumption correctly identifies the validity and independence of the LLM evaluators as the pivotal premise. The paper is transparent about the limitation, which is why a conditional verdict is appropriate rather than rejection. A human rating study is the decisive test because the central claim concerns subjective creativity, and LLM-only scores cannot establish a human-facing conclusion. The leave-one-out analysis would additionally test the specific overlap confound, but the human study subsumes it: if human judgments reproduce the 8B-over-17B pattern, the overlap concern is largely moot; if they do not, the central claim is unsupported regardless of internal consistency. The cross-experimental comparison of effect sizes is a secondary concern; the primary risk is the unvalidated outcome measure.","tokens_in":7879,"tokens_out":7811,"duration_ms":78793,"concrete_test":"Run a small human rating study: recruit three culinarily knowledgeable raters to independently score a stratified random sample of 30 final recipes (5 per condition from the 3x2 design) on the four TTCT dimensions using the same product-level definitions as the LLM evaluators, blind to condition. If the 8B-scorer condition is not significantly higher than the 17B-scorer condition on creativity, fluency, flexibility, and elaboration, the central claim fails to generalize to human judgment; if the human ratings replicate the LLM pattern, the concern is resolved and the conditional verdict can be upgraded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's strongest claim (Abstract and Conclusion) is that evaluator design matters more than additional search depth, supported by Appendix D Table 1, where the LargeEvaluator coefficient is significantly negative for creativity, fluency, flexibility, and elaboration. The outcome variable is the average of four LLM TTCT evaluators, one of which is Meta-17B, the same model used as the in-loop selection scorer. This creates a direct overlap between the manipulated factor and the measurement instrument. The benchmark check in Study Design and Evaluation validates only the in-loop rubric against the human competition; it does not validate the post hoc TTCT evaluators against human creativity judgments of generated recipes. No human ratings of the generated recipes are reported, and the paper explicitly concedes it can say more about behavior under LLM evaluation than about human judgment. If LLM evaluator scores mostly reflect self-preference or rubric alignment rather than perceived creativity, the observed 8B-over-17B advantage would not generalize to human judgment, and the comparative 'more than search depth' claim would lose its empirical basis. This is the load-bearing assumption on which the central result rests.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper adapts the FunSearch evolutionary generation loop to a subjective creative domain—recipe generation for the 2024 Pillsbury Bake-Off—and evaluates the resulting artifacts with four TTCT-style LLM evaluators. Two experiments manipulate iteration count, generator temperature, and in-loop selection-scorer model size. The authors report that iterative generation can reach creativity scores comparable to a human reference set, that additional iterations do not improve scores, that the smaller Meta-8B in-loop scorer yields significantly higher scores than the Meta-17B scorer on most TTCT dimensions, and that lower temperature mainly reduces originality. They conclude that evaluator design is a first-order design variable in subjective creative search.","tokens_in":8080,"tokens_out":3956,"duration_ms":42726,"significance":"If the main finding holds, it would be a valuable reframing for computational creativity: the design of the in-loop evaluator, rather than additional search depth or sampling stochasticity, may determine the creativity of iteratively generated artifacts. The paper has real strengths: it transfers a well-known program-search method to a non-objective domain, provides detailed appendices with the recipe skeleton, in-loop rubric, evaluation prompt structure, and full regression output, and explicitly acknowledges its main limitation—that Meta-17B serves as both the in-loop scorer and one of the final evaluators. The use of four LLM evaluators and a partial human benchmark is also a reasonable pilot-level design. However, the statistical support is weak and the evaluator-overlap issue is load-bearing, so the contribution is best assessed as a promising pilot rather than a fully established result.","major_comments":[{"comment":"The central negative LargeEvaluator coefficient in Appendix D, Table 1 is confounded by the fact that Meta-17B is both the in-loop selection scorer and one of the four final TTCT evaluators. The paper acknowledges this overlap in the Limitations section, but it does not address it quantitatively. Because the outcome is the average of four LLM evaluators, one of which is identical to the manipulated scorer, the observed 8B-over-17B difference could partly reflect self-preference or rubric alignment rather than perceived creativity. I request a re-analysis that excludes the Meta-17B evaluator from the outcome variable, or a small human-rating study on a subset of the generated recipes, before the evaluator-design claim is made.","section":"Study Design and Evaluation; Limitations and Future Work; Appendix D, Table 1"},{"comment":"The regression evidence is statistically fragile. Adjusted R2 values range from 0.006 to 0.037, meaning the manipulated factors explain almost none of the variance in the outcomes, and the five outcome models together involve fifteen significance tests with no multiple-comparison correction. The originality model, which is the basis for the claim that temperature mainly reduces originality, has adjusted R2 = 0.006. The paper should report standardized effect sizes, confidence intervals, and either corrected p-values or a pre-specified analysis plan before claiming that evaluator size 'matters most.' It should also include iteration count and scorer size in a single comparative model if the conclusion is to be that evaluator design matters more than search depth.","section":"Appendix D, Table 1; Results and Interpretation"},{"comment":"Experiment 1 reports mean final creativity scores of 3.921, 3.835, and 3.927 for 5, 15, and 30 iterations, respectively, but does not provide error bars, confidence intervals, or inferential tests for these means. The claim that additional iterations do not improve creativity, and that iterative search reaches scores comparable to the human benchmark, is therefore not supported by the statistics reported. The one-iteration baseline comparison (in-loop score dropping from 4.71 to 4.00, final creativity roughly 4.1) is also given only as point estimates in the text. I recommend reporting per-condition uncertainty and formal comparisons, and ideally placing Experiment 1 and Experiment 2 in a single model that allows a direct test of the relative importance of iteration count and evaluator size.","section":"Results and Interpretation, Experiment 1"},{"comment":"The benchmark check validates only the in-loop rubric against the human competition, and even there the agreement is partial: the human winning recipe is placed 'near the top' rather than at rank one, and the stories are absent from the comparison. The post hoc TTCT-style evaluators are not validated against human creativity judgments of the generated recipes. Since the paper's central comparative claim ('evaluator design matters more than search depth') is expressed in terms of creativity scores from these LLM evaluators, this validation gap is load-bearing. The authors explicitly concede that the paper can say more about behavior under LLM evaluation than about human judgment; I would like that qualification to be reflected in the abstract and conclusion, or the missing human ratings to be supplied.","section":"Study Design and Evaluation; Limitations and Future Work"}],"minor_comments":[{"comment":"The caption mentions a red cross marking the winning Pillsbury recipe, but the figure would be easier to read with explicit panel labels and significance annotations; error bars are shown as one standard deviation, but no statistical tests are indicated on the figure itself.","section":"Figure 3"},{"comment":"The one-iteration baseline values (in-loop weighted score 4.71 versus 4.00, final creativity roughly 4.1) appear only in the text; a small table or figure would improve reproducibility and make the comparison easier to verify.","section":"Results and Interpretation, Experiment 1"},{"comment":"The reference to Taylor et al. (2024) concerns clinical tasks and is not an obvious citation for the model-size effect in creative generation; a more directly relevant scaling-law or model-capability reference would strengthen the motivation in the Introduction.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is transparent and the limitations section is unusually honest, but the gap between the strength of the claims and the strength of the evidence is currently too large. The evaluator-overlap problem is acknowledged yet unresolved, and the regression results explain very little variance. A revision that re-runs the key analyses without the overlapping evaluator and adds even a small human-judgment component would make the central claim credible. The paper might also consider framing the contribution explicitly as a pilot, since the current abstract and conclusion assert a first-order design conclusion that the data do not yet fully support."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this if you work on LLM creativity or evaluator design. The paper does something genuinely new: it adapts FunSearch to a subjective creative domain (recipes) and runs a balanced comparison of iteration count, temperature, and in-loop scorer size. The design is sensible, the reporting is honest, and the central claim—that the in-loop evaluator matters more than search depth or stochasticity—is interesting and actionable.\n\nThe soft spots are real, though mostly acknowledged. The big one is that final creativity is measured by four LLM evaluators, one of which is the same Meta-17B used as the in-loop scorer. So the negative coefficient on the larger evaluator may reflect self-preference rather than a general truth about scorer size. The authors note this in Limitations, but it is load-bearing, not a side caveat. The regression R2 values are 0.006–0.037, so the model explains almost nothing; the significant coefficients are small. Experiment 1 reports means without error bars, and the “human-comparable” claim relies on a 30-recipe reference set with no human ratings of the generated recipes. The benchmark check showed only partial agreement with the official outcome, so we do not know whether the 8B-over-17B difference would survive human judgment. No code or data are released.\n\nThat said, this is a pilot, and the authors do not overclaim. They explicitly say the paper says more about behavior under LLM evaluation than about human judgment. The direction of the effect—smaller scorer beating larger—is plausible as an artifact of evaluator alignment, and the paper frames it as a question worth studying, not a settled fact. I would not cite it as evidence for the main claim yet, but I would bring it to a reading group as a case study in evaluator overlap and LLM-only evaluation.\n\nFor peer review: yes, send it out. The idea is novel, the execution is transparent, and the critique will be productive. A referee should push for human evaluation of at least a sample of recipes, multiple-comparison correction, error bars in Experiment 1, and ideally code/data release. With those, the central claim might become solid. As is, treat it as a promising pilot, not as an established result.","headline":"A transparent pilot with a plausible but unproven central claim: the in-loop evaluator seems to matter more than iteration or temperature, but the evidence rests on unvalidated LLM evaluation and a scorer/judge overlap.","tokens_in":8604,"tokens_out":2119,"would_cite":false,"duration_ms":22826,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The in-loop evaluator, not search depth or temperature, is what determines creativity scores in iterative LLM recipe search.","keywords":["iterative generation","LLM creativity","FunSearch","in-loop evaluator","TTCT evaluation","recipe generation","temperature","selection scorer"],"falsifier":"Have a panel of human cooks rate, blind, the same retained recipes from the small-scorer and large-scorer conditions; the central claim fails if average human creativity ratings do not favor the small-scorer condition, or if the official Bake-Off winner is not rated at or near the top of the human reference set.","tokens_in":7667,"feed_emoji":"🍳","tokens_out":8990,"duration_ms":72844,"temperature":0.7,"pith_summary":"This pilot study asks whether iterative generation-and-evaluation search improves the creativity of LLM-generated recipes, and which design factor matters most. It adapts the FunSearch program-search algorithm, originally built for objective tasks, to the 2024 Pillsbury Bake-Off: recipes are generated under a competition skeleton, an in-loop scorer keeps the best, and the survivors seed the next round. Across two experiments manipulating iteration count, generator temperature, and in-loop scorer size, the paper finds that iterative search can reach human-comparable creativity scores but extra iterations alone do not help. The in-loop evaluator is the decisive factor: a smaller scorer yields significantly higher creativity, fluency, flexibility, and elaboration scores than a larger one, while lower temperature mainly suppresses originality. The paper concludes that evaluator design is a first-order design variable in subjective creative search.","feed_headline":"Smaller AI judge, not more search, lifts recipe creativity","feed_subtitle":"Iterative recipe search scores highest with a small in-loop judge; extra cycles and low temperature barely matter.","key_machinery":"The load-bearing mechanism is the FunSearch loop adapted to a subjective domain. Semi-isolated islands preserve different recipe lineages; best-shot prompting rebuilds prompts from high-scoring candidates; and an in-loop selection scorer applies a weighted Pillsbury-style rubric (recipe content 70%, story 30%) to decide which candidates survive into the next generation. The final creativity analysis is kept separate from the search objective, using four post hoc LLM evaluators with TTCT-derived definitions of fluency, flexibility, originality, and elaboration. The controlled contrast that carries the argument is the in-loop scorer's model size, since the generator stays fixed: the same search loop is run with a smaller and a larger evaluator, and the smaller one wins on most dimensions.","core_discovery":"The central claim is that when creativity is judged subjectively, the in-loop evaluator shapes how much iterative search improves outputs more than search depth or sampling randomness does. The paper supports this by adapting FunSearch to recipe generation: multi-island search keeps separate lineages, high-scoring candidates seed later prompts, and an in-loop selection scorer applies a weighted Pillsbury-style rubric (70% recipe, 30% story) to admit or discard candidates. Final creativity is assessed separately by four LLM evaluators using TTCT-derived product-level dimensions: fluency, flexibility, originality, and elaboration. In the main comparison, the smaller 8B selection scorer produced significantly higher final scores on creativity, fluency, flexibility, and elaboration than the larger 17B scorer, whereas moving from 5 to 30 iterations produced no clear monotonic gain and lowering generator temperature only reduced originality. The paper interprets this as evidence that the evaluative ecology, not the amount of generation, determines whether iterative novelty accumulates in subjective creative domains.","pith_inferences":["A natural next test is whether the evaluator-size effect survives human judging: because the paper's own benchmark shows only partial agreement with the official competition outcome, a human panel could plausibly reverse the smaller-scorer advantage.","The paper's framing suggests varying the in-loop evaluator during search, for example rotating several scorers or changing rubrics between islands, as a way to escape convergence on conservative recipes, but this extension is not tested here.","Because the larger in-loop scorer also served as one of the final evaluators, a cleaner replication with fully disjoint generator, scorer, and final judge models is needed to rule out self-preference as the source of the size effect."],"forward_implications":["Adding more FunSearch iterations to a recipe-generation loop does not by itself raise final creativity scores; the near-one-shot baseline and the 5, 15, and 30 iteration conditions land at similar levels.","Switching the in-loop selection scorer from a larger to a smaller model reliably raises final creativity, fluency, flexibility, and elaboration scores under TTCT-style LLM evaluation.","Lower generator temperature mainly costs originality; it does not improve any other TTCT dimension at a significant level.","Iterative search with a well-chosen evaluator can produce recipe artifacts whose LLM-judged creativity is comparable to a human reference set from the same competition.","Creative LLM systems should treat the diversity, architecture, and incentive structure of in-loop evaluators as a central design variable rather than relying on more generation cycles."],"supporting_citations":[{"why":"Supplies the FunSearch algorithm that the paper adapts from objective program search to subjective recipe generation.","marker":"Romera-Paredes et al. 2024"},{"why":"Defines the TTCT dimensions (fluency, flexibility, originality, elaboration) that the final evaluators measure.","marker":"Torrance 1966"},{"why":"Provides the TTCT-style LLM evaluation approach for creativity that the paper uses for final scoring.","marker":"Zhao et al. 2024"},{"why":"Establishes the precedent for adapting TTCT-style dimensions to products rather than using them as direct psychometric measures.","marker":"Chakrabarty et al. 2024"},{"why":"Motivates the temperature manipulation by linking sampling temperature to novelty-coherence trade-offs in creative language generation.","marker":"Peeperkorn et al. 2024"},{"why":"Documents self-preference bias in LLM-as-a-judge, the evaluator-overlap concern the paper acknowledges as a limitation.","marker":"Wataoka, Takahashi, and Ri 2024"}],"fun_headline_variants":["Small in-loop judge tops extra search for recipe creativity","Recipe creativity hinges on evaluator size, not iteration count","LLM recipe search: smaller critic beats longer search","Originality rises with small selection scorer, not more loops","Evaluator choice, not search depth, drives recipe creativity"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the automated judges used to score creativity are a valid stand-in for human taste in recipes; if they mostly reward rubric compliance or their own preferences, the reported advantage of the smaller in-loop scorer would not carry over to human judgment.","fun_headline_variants_meta":{"raw":{"variants":["Small in-loop judge tops extra search for recipe creativity","Recipe creativity hinges on evaluator size, not iteration count","LLM recipe search: smaller critic beats longer search","Originality rises with small selection scorer, not more loops","Evaluator choice, not search depth, drives recipe creativity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000142,"raw_usage":{"total_tokens":1135,"prompt_tokens":879,"completion_tokens":256,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":495,"completion_tokens_details":{"reasoning_tokens":191}},"tokens_in":495,"tokens_out":256,"duration_ms":3732,"temperature":1.0,"reasoning_tokens":191,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T11:56:04.891892+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have a panel of human cooks rate, blind, the same retained recipes from the small-scorer and large-scorer conditions; the central claim fails if average human creativity ratings do not favor the small-scorer condition, or if the official Bake-Off winner is not rated at or near the top of the human reference set.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the FunSearch algorithm that the paper adapts from objective program search to subjective recipe generation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the TTCT dimensions (fluency, flexibility, originality, elaboration) that the final evaluators measure."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the TTCT-style LLM evaluation approach for creativity that the paper uses for final scoring."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the precedent for adapting TTCT-style dimensions to products rather than using them as direct psychometric measures."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Motivates the temperature manipulation by linking sampling temperature to novelty-coherence trade-offs in creative language generation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents self-preference bias in LLM-as-a-judge, the evaluator-overlap concern the paper acknowledges as a limitation."}],"review_version":1}