{"id":"d4109f66-1323-4348-bc80-facd04c386d2","arxiv_id":"2607.09709","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"Strict-launch-gated rejection-sampling self-distillation compounds cross-family clean Godot generation from 8.8% to 42.2% and full best-of-K coverage, while gold duplication and a lenient BUILD filter erase the gain.","lead":"Filtering a code model’s own game projects by whether they actually launch under a headless engine, then retraining on the keepers, compounds clean generation on unseen game genres from 8.8% to 42.2%. The result isolates that verifier precision—not more data or a fancier optimizer—sets the ceiling of self-improvement loops.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Strict-launch plus exec-ground still leave open whether accepted projects realize brief-specific game capability rather than engine-compliant scaffolding that merely launches and exhibits probeable state change.","rationale":"The reader’s weakest_assumption correctly isolates the load-bearing condition for the interpretive claim. The empirical core (8.8 %→42.2 % per-candidate, 18/25→25/25 coverage, each round significant; gold-dup regression; BUILD-gated erasure; quality+quantity split) is tightly controlled and statistically appropriate for N=25. §8 audits mitigate the stub concern but stop short of brief fidelity or playability, exactly as the reader notes. This soft spot does not invent an internal contradiction or overturn the within-testbed launch-rate results; it keeps the verdict CONDITIONAL pending stronger functional measures or public code/fuel for external audit. No more severe load-bearing flaw (leakage, optimizer confound, or statistical artifact) is present.","tokens_in":13764,"tokens_out":587,"duration_ms":24885,"concrete_test":"Blind-score one best clean candidate per held-out task from SFT versus RFT-r3 (25 pairs) with a short human checklist or non-gameable rubric for brief fidelity (family-typical mechanics present, input responsiveness, win/lose conditions) and basic playability; if RFT shows no significant fidelity/playability gain while retaining the launch-rate gain, the functional-transfer reading does not hold.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (compounding cross-family clean generation under a precise ungameable gate, with filter precision causal) rests on the accepted distribution teaching transferable functional game synthesis. §4–§5 define the gate as exit-0 with no parse/load/runtime error; §8 adds static richness (+43 % GDScript lines, more signals/files) and a rising exec-ground score (10→15/25 tasks grounded at ≥0.50). Yet exec-ground only requires runtime-error-free SceneTree state change under passive/synthetic/replay probes and is explicitly “a functional lower bound, not a measure of playability” (§10). No metric checks whether a clean horror/rhythm/etc. project implements the natural-language brief (family-typical mechanics, win conditions, input semantics beyond the probe). If the curriculum mainly teaches Godot project structure that launches and wiggles, the transfer is harness-level launchability rather than game-generation capability, so the “verifier is the curriculum” lesson for open-ended synthesis is narrower than claimed. The BUILD-swap and gold-dup controls cleanly isolate precision for the launch-rate outcome; they do not close this semantic gap.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper argues that post-training a code generator against a deterministic, judge-free execution filter (strict-launch: clean headless Godot launch with exit code 0 and no parse/load/runtime error) yields compounding cross-family gains under rejection-sampling self-distillation, whereas a learned visual judge is gameable. On GameCraft-Bench, a Qwen3-14B+LoRA model trained only on ten families improves clean generation on four held-out families from 8.8% to 42.2% per-candidate and best-of-K coverage from 18/25 to 25/25 over three rounds. Matched controls attribute the gain to verifier precision and fuel diversity/quality rather than data volume or the optimizer: exact gold-duplication regresses below the base model; a count-matched generator swap isolates comparable quality and quantity channels; and swapping only the filter for the near-vacuous BUILD check erases the first-round gain. A secondary execution-grounding probe and a static code audit are used to argue that accepted projects are functional rather than launch-but-empty stubs. The stated lesson is that in self-distillation the verifier is the curriculum.","tokens_in":14091,"tokens_out":1537,"duration_ms":26109,"significance":"If the result holds, the paper supplies unusually clean causal evidence that the precision and ungameability of an acceptance filter, not merely the self-distillation optimizer or added sample count, determine what capability transfers. The design is a genuine strength: budget- and task-matched controls, a filter-swap ablation, paired task-level cluster-permutation tests (N=25), leave-one-family-out checks on the key contrast, training-seed robustness for round 1, and an independent calibrated grounding signal. Framing game generation as a verifiable testbed for a falsifiable principle (precision and exploration conditions) is useful beyond the specific engine. The contribution is empirical and methodological rather than a new algorithm; its value is the controlled isolation of the verifier as the causal variable.","major_comments":[{"comment":"The central empirical claim about compounding clean-launch rates and the causal role of filter precision is well supported by Table 1 and the matched controls in §7. The interpretive claim that the transferred capability is functional game generation, however, rests on a thinner bridge. §8’s static audit (more GDScript lines, signals, files) and exec-ground (SceneTree state change under passive/synthetic/replay probes, score ≥0.50) rule out empty stubs and pure non-execution, but neither checks brief- or family-specific mechanics, win conditions, or input semantics beyond the probe. The paper itself calls exec-ground “a functional lower bound, not a measure of playability” (§10). As written, the curriculum may primarily teach engine-compliant project structure that launches and exhibits probeable state change. This does not invalidate the launch-rate results or the BUILD-swap isolation o","section":"§8, §10, abstract"},{"comment":"The general principle that “the ceiling of a self-improvement loop is set by the integrity of its verifier” (§11) is motivated by a single 14B LoRA setup, one engine, and four held-out families (N=25 tasks). §10 acknowledges broader families and larger models as next tests, and training-seed robustness is reported only for round 1 (§6). That is acceptable for a testbed paper, but the conclusion currently states the principle more strongly than the evidence warrants. Please qualify the conclusion so that the load-bearing claim remains the controlled filter-swap result on this benchmark, with the broader implication framed as a hypothesis supported by one instance rather than established.","section":"§6, §10, §11"}],"minor_comments":[{"comment":"Notation for p-values is inconsistent in the abstract and body (e.g., p<1e-4 vs. p <10 −4 with a space and superscript). Standardize to a single scientific notation throughout.","section":"abstract, §6"},{"comment":"Figure 1’s verification ladder is helpful; a one-line quantitative callout of the BUILD vs. strict-launch acceptance rates (99.9% vs. 8.2% on the supervised training-task pool, as stated in §4) on the figure itself would make the precision gap immediately readable.","section":"Figure 1, §4"},{"comment":"Algorithm 1 mentions “at least three files” as an acceptance side-condition; the main text should briefly justify this threshold (or note it as a stub filter) so readers do not treat it as an unmotivated hyperparameter.","section":"Algorithm 1, §5"},{"comment":"The companion diagnosis that the visual judge is gameable is important for motivation (§3, Appendix A) but is summarized from separate work. A short table of the asset-swap score deltas in the main text would make the “why not train against the judge” argument self-contained without forcing a full appendix read.","section":"§3, Appendix A"},{"comment":"Related work (§9) correctly positions the paper relative to STaR/ReST and reward-hacking literature. A sentence distinguishing this filter-swap design from concurrent runtime-verification game-generation work (already cited) would further clarify the novelty axis: not a new game generator, but a controlled study of verifier integrity under fixed optimizer.","section":"§9"}],"recommendation":"minor_revision","confidential_remarks":"The experimental controls are stronger than average for self-distillation / code-generation papers; the BUILD filter-swap is the right causal test. My main reservation is claim scope (launch/scaffolding vs. brief-faithful games), not internal inconsistency. I would not block on additional playability metrics if the authors tighten language, but I would push back if they keep the broader “game generation capability” framing without qualification. Fit for a methods/empirical AI venue is good; less so if the journal expects new optimizers or large multi-domain scaling."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing is the control set. Same rejection-sampling loop, same budget, same tasks: strict-launch compounds held-out clean generation 8.8\to42.2% and coverage to the gold 25/25 ceiling; exact gold duplication regresses below base; swapping only the filter for near-vacuous BUILD erases the gain. That is a sharp, falsifiable isolation of verifier precision as the curriculum, not data volume or the optimizer.\n\nWhat is new is not self-distillation or executable game checks—those exist—but the matched decomposition (quality vs quantity channels, both significant) plus the filter-swap that holds everything else fixed. Statistics are task-paired cluster-permutation (N=25), Wilson intervals, leave-one-family-out on the key contrast, and training-seed check for round 1. Static richness and the independent exec-ground probe (calibrated gold vs known-bad) address the stub worry and rise monotonically. Citations to STaR/ReST/ReST-EM and reward-hacking work are appropriate; the paper does not overclaim novelty on the method family.\n\nSoft spots are real but proportionate. N=25 on one engine/model, code/fuel not yet released, seed robustness only for r1. More importantly, the stress-test concern lands partially: strict-launch plus exec-ground certify launchability and probeable state change, not brief-specific playability or family mechanics. The authors say so in §10. That narrows the lesson for open-ended synthesis, but it does not break the within-testbed claim that a precise ungameable gate compounds while a lenient one does not. The BUILD and gold-dup controls cleanly isolate precision for the launch-rate outcome they actually measure.\n\nThis is for people working on RLVR, code self-training, or agentic coding who care about what the acceptance filter actually teaches. It deserves a serious referee. I would bring it to reading group and cite the control design.","headline":"Clean causal isolation of verifier precision in self-distillation: launch-gated RFT compounds cross-family Godot generation, gold-dup regresses, BUILD-swap erases the gain.","tokens_in":14701,"tokens_out":503,"would_cite":true,"duration_ms":4894,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Under a deterministic launch check, self-distillation compounds clean game generation on unseen families from 8.8% to 42.2% and full best-of-K coverage—because the verifier, not the optimizer, sets the curriculum.","keywords":["self-distillation","rejection sampling","game generation","verifier precision","reward hacking","execution verification","Godot","cross-family generalization"],"falsifier":"Rerun the same loop on the same split with only the acceptance filter changed to another high-throughput but low-precision check (or measure accepted projects with a stronger independent functional audit); if compounding still appears under a high false-positive gate, or if launch-clean projects systematically fail richer playability checks, the precision-as-curriculum claim fails.","tokens_in":14669,"feed_emoji":"🎮","tokens_out":967,"duration_ms":15865,"temperature":0.7,"pith_summary":"Post-training a code generator against a learned judge can reward cheap surface tricks that raise the score without making the artifact better. This paper trains against the opposite signal: a fixed, judge-free engine check that only accepts a generated Godot project if it launches cleanly headless, with no parse, load, or runtime error. On a from-scratch game-generation benchmark, three rounds of rejection-sampling self-distillation under that gate lift clean generation on four held-out game families from 8.8% to 42.2% per candidate and push best-of-K coverage to the gold ceiling of 25/25. Matched controls show the gain is not from adding more gold data or from the training loop itself: exact gold duplication hurts, and swapping only the filter for a near-vacuous build check erases the improvement. A second execution-grounding probe rises in step, arguing the accepted projects are functional rather than empty stubs. The lesson the authors press is simple and general: in self-distillation, what the verifier certifies is what the model learns.","feed_headline":"Launch gate lifts clean game gen from 8.8% to 42%","feed_subtitle":"Filter precision, not more gold data, drives the self-distillation gains on unseen families.","key_machinery":"Strict-launch: a binary, judge-free headless engine predicate (exit code 0 and no parse/load/runtime error) used as the acceptance filter in iterative self-distillation, so only projects that actually run become the next training distribution—the mechanism behind the claim that “the verifier is the curriculum.”","core_discovery":"When rejection-sampling self-distillation is gated by a deterministic, ungameable strict-launch check, out-of-family clean game generation compounds across rounds; the filter’s precision is causal, because an exactly matched gold-duplication control regresses below the base model and a lenient BUILD filter swap returns performance to base, isolating the signal rather than the optimizer or data volume.","pith_inferences":["The same precision-over-volume principle should matter for other conjunctive synthesis tasks—full web apps, infrastructure configs, or multi-file agents—where partial correctness is not enough to run.","If the ceiling of self-improvement is set by the verifier, investing in harder ungameable execution checks may raise capability more than scaling soft preference models alone.","Learned judges may remain useful as secondary scorers only when an ungameable execution gate already filters out non-running candidates from below."],"forward_implications":["Self-improvement loops for open-ended generation are limited by verifier integrity, not by data volume alone.","Exact duplication of gold references can regress capability relative to diverse self-generated clean fuel of the same size.","A binary ungameable execution gate can act as a coverage process that widens the set of working solutions across rounds.","Game generation is a measurable testbed for studying how filter precision shapes transfer under a fixed optimizer.","Replacing a gameable learned judge with a deterministic launch gate changes what post-training transfers from cosmetic score features to executable behavior."],"fun_headline_variants":["Strict-launch gate compounds clean game gen from 8.8% to 42%","Verifier precision, not gold data, drives cross-family game gains","Execution-gated self-distillation hits 25/25 best-of-K coverage","Ungameable launch filter is the curriculum for out-of-family games","Lenient BUILD swap erases gains; strict-launch is the causal signal"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That a clean headless launch, backed by a secondary execution-grounding score, is precise enough to certify functional transferable projects rather than harness-specific or launch-but-empty artifacts.","fun_headline_variants_meta":{"raw":{"variants":["Strict-launch gate compounds clean game gen from 8.8% to 42%","Verifier precision, not gold data, drives cross-family game gains","Execution-gated self-distillation hits 25/25 best-of-K coverage","Ungameable launch filter is the curriculum for out-of-family games","Lenient BUILD swap erases gains; strict-launch is the causal signal"]},"model":"grok-4.5","effort":"low","cost_usd":0.006748,"raw_usage":{"total_tokens":1792,"prompt_tokens":953,"num_sources_used":0,"completion_tokens":106,"cost_in_usd_ticks":67480000,"prompt_tokens_details":{"text_tokens":953,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":733,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":953,"tokens_out":106,"duration_ms":5254,"temperature":1.0,"reasoning_tokens":733,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T17:24:05.496786+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Rerun the same loop on the same split with only the acceptance filter changed to another high-throughput but low-precision check (or measure accepted projects with a stronger independent functional audit); if compounding still appears under a high false-positive gate, or if launch-clean projects systematically fail richer playability checks, the precision-as-curriculum claim fails.","supporting_citations":[],"review_version":1}