{"id":"ee713c73-4ebc-431d-a953-d57b7e36b828","arxiv_id":"2608.01000","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"LLMs under one-shot greedy decoding enumerate or materialize acceptable sets much worse than they judge membership, a gap that persists across scale, family, and generation and is dominated by omissions.","lead":"Language models judge whether a candidate fits a specification far better than they author the set that fits it, under one-shot greedy prompting. This matters because model-written answer keys, test suites, and reward functions are increasingly used to define correctness for other systems, where silent omissions are hard even for reviewers to spot.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Core scissors is solid, but practical reach rests on an unmeasured deployment premise: §3.7 shows reasoning closes the algorithmic gap, so the abstract's 'usually deployed with one-shot greedy' scope needs support.","rationale":"The reader's CONDITIONAL verdict already identifies the same load-bearing scope premise: one-shot greedy decoding with no test-time reasoning. My stress-test agrees with that assessment. The paper's internal evidence is strong: the authoring–execution gap is triangulated across four constructions with orthogonal confounds, the controls (JSON checkbox, prompt rephrasing, alignment, intensional predicate emission) localize the deficit, and the paper explicitly reports the reasoning-enabled result that closes the algorithmic gap. The concern is not that the measurements are wrong; it is that the abstract and RLVR motivation generalize from 'under this protocol' to 'the role is usually deployed with this protocol,' and that adoption premise is unmeasured. The concrete test—running the full protocol with reasoning enabled across scale and family—would settle whether the gap is a stable capability boundary or a decoding-protocol artifact. Because the reader already conditioned the verdict on this scope caveat (and on reproducibility), my read does not change the verdict; it reinforces it. No ad hominem, no manufactured flaw: the soft spot is real but already flagged, and the central claim's scientific core remains intact within its stated scope.","tokens_in":48296,"tokens_out":11944,"duration_ms":140847,"concrete_test":"Re-run the full authoring protocol with test-time reasoning enabled at all open-weight scales (Qwen2.5/Qwen3 3B–72B) and frontier models, on the complete algorithmic set (n=240) and all 164 HumanEval+ problems (not the 80-problem subsample), scoring with the same statistics as Tables 2 and 5. Pre-register an equivalence threshold, e.g., the bootstrap CI of the authoring–execution gap within ±0.02. If the gap closes for a majority of model×construction cells, the headline should be re-scoped as a claim about greedy decoding rather than about the examiner role generally; if it persists, the deployment premise survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is valid as scoped: under greedy one-shot decoding, authoring lags execution on all four constructions, with disjoint CIs, prompt/format controls, an intensional control, and an honestly falsified repair prediction. The load-bearing soft spot is the scope premise stated in the abstract and first paragraph—that one-shot greedy authoring with no test-time reasoning is 'the protocol the role is usually deployed with.' That premise is asserted, not measured. §3.7 shows the protocol boundary is doing real work: GPT-5.1 algorithmic authoring rises from F1 0.673 to 0.976 and the gap falls from +0.259 to +0.008 with a CI covering zero, and Opus's code gap shrinks from +0.274 to +0.096. The manuscript's own §7.2 lists decoding sweeps, constrained decoding, and tool-augmented pipelines among controls not run. If actual production verifier authoring uses reasoning, sampling, or tools, the measured deficit and the RLVR tax (keys authored one-shot) may shrink materially. This is not an internal inconsistency—the paper explicitly scopes its claim—but it is a genuine load-bearing premise about the paper's practical generalization. The EvalPlus oracle premise is secondary: the complete-truth arms already carry the core comparison.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper asks whether LLMs can author the acceptable sets that a correctness specification denotes, and compares that authoring interface against the same models' pointwise membership judging. Four reference constructions are used: mechanically decidable string and numeric predicates over presented lists (complete truth), HumanEval+/MBPP+ with an EvalPlus-hardened executable oracle, and WordNet synsets. Under one-shot greedy decoding, authoring F1 lags execution F1 in every tested model–construction cell across a 3B–72B scale range, six families, and two generations; the gap does not close. Supplementary results isolate the deficit as enumeration rather than specification (predicate-emission F1 ≈ 0.99), show omission is 6–7× harder to detect than over-inclusion in the review protocol, falsify a registered sample-then-verify repair, introduce a known-correct-probe gate that cuts code false rejection to ≤5%, and measure an RLVR reward tax of −1.9 points on complete truth and up to −32 points WordNet-relative. The paper explicitly scopes its claims to the one-shot regime and reports a large set of controls, negative results, and instrumentation defects.","tokens_in":48545,"tokens_out":10823,"duration_ms":111922,"significance":"If the results hold, they identify a robust and consequential asymmetry: models can judge membership far better than they can materialise the membership set, and the resulting omissions are structurally resistant to local review and cost accuracy when wired into RLVR. The central measurement is unusually well grounded: the complete-truth constructions have no incompleteness confound; code correctness is decided by execution; no LLM judge grades the construction measurements; bootstrap CIs are reported; and the design includes a predicate-emission control, prompt/format ablations, cross-family and cross-generation replications, and a falsified registered prediction. The paper also reports its own instrumentation defects and retracted sub-predictions, which strongly supports credibility. The exploratory OLS scale fit is explicitly not used for the headline. The main qualification is scope: the practical reach of the headline depends on the premise that one-shot greedy authoring is the usual deployment protocol, a premise that is asserted rather than measured and that §3.7 shows is doing real work.","major_comments":[{"comment":"The abstract and §1 assert that one-shot greedy authoring with no test-time reasoning is 'the protocol the role is usually deployed with.' No evidence is cited for this deployment claim, and §3.7 shows the protocol boundary is load-bearing: with reasoning enabled, GPT-5.1's algorithmic authoring F1 rises from 0.673 to 0.976 and the gap falls to +0.008 with a CI covering zero; Opus's code gap falls from +0.274 to +0.096. §7.2 also lists decoding sweeps and tool-augmented pipelines among controls not run. The central measurement is internally valid as scoped, but the practical generalization in the abstract depends on an unmeasured premise. Please either provide evidence of deployment practice or re-scope the rhetoric (e.g., 'under one-shot greedy decoding, a protocol in current use') and temper the abstract's 'usually deployed.'","section":"Abstract, §1, §3.7"},{"comment":"The abstract states that 'a production deployment of 43,227 scored items fails omission-first at 10:1' without the caveats that §6 itself insists on. §6 says the 10:1 ratio is evaluator-diagnosed, that the evaluator re-affirms only 25% of open-format fail verdicts, and that the ratio 'is directional, not a measured rate.' Because the paper's central claim does not depend on the field evidence, this is fixable by moving the number behind the same qualification in the abstract or removing it from the abstract; as written, a reader taking only the abstract will read a rate the paper itself says cannot be read as one.","section":"Abstract, §6"}],"minor_comments":[{"comment":"References [20] and [57] are the same Konstantinou et al. paper with the same title and arXiv identifier. Please merge into one entry and update citations accordingly.","section":"References"},{"comment":"The limitation paragraph says the complete-truth RLVR tax is 'a much smaller 2.3 points,' but §5.1 and Appendix G report the corrected estimate as 1.91 (≈1.9) points. Update this number to match the revised evaluation.","section":"§7.3"},{"comment":"In the sentence after Eq. (3), 'whichever of recall or precision is , which is what makes' appears to be missing the word 'lower' or 'smaller' after 'is.'","section":"§3.1"},{"comment":"Table 3 shows authoring F1, execution F1, and gap [95% CI], but only the gap column has a CI. Since the text claims 'disjoint CIs' between authoring and execution, please provide the marginal CIs for both arms or explicitly state that the gap CI is the evidence of separation.","section":"Table 3"},{"comment":"In the caption of Table 6, 'False rejection falls from 0.58–0.92 to ≤0.05' is supported by the table, but the adjacent text says '7–10% of Llama's' yield while the table lists 0.213 and 0.134 (i.e., 21.3% and 13.4%). Please reconcile the yield percentages in the prose with the table.","section":"§3.5 / Table 6"}],"recommendation":"major_revision","confidential_remarks":"This is a strong paper with unusually honest reporting: a registered prediction is falsified, instrumentation defects are disclosed, and the exploratory scale fit is explicitly not used for the headline. The revision should focus on abstract-level scoping and internal numeric consistency. The two major comments are both fixable without new experiments, but the deployment-premise issue is substantive enough that I would want to see the revised text before accepting."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper earns its pay: it shows, on judge-free constructions, that LLMs execute membership far better than they author the acceptable set in one greedy pass, that the error is omission-dominated, and that omission is witness-poor to review. The triangulation is real—complete truth on strings and numbers, executable truth on HumanEval+/MBPP+, lexical truth on WordNet—and the conventions are careful: no LLM judge grades the central measurements, CIs are bootstrapped, a falsified sample-then-verify prediction is reported, and a technical defect they found in their own instrumentation (double BOS) is disclosed with before/after numbers. The intensional control (predicate emission at F1≈0.99) genuinely localizes the deficit to one-shot enumeration, and the RLVR tax is paired-seed significant on both constructions. That is a solid contribution.\n\nThe soft spot is the scope premise, not the measurement. The abstract says \"the protocol the role is usually deployed with\" is one-shot greedy authoring with no test-time reasoning. That is asserted, not measured. And the paper's own §3.7 shows the boundary does real work: GPT-5.1 with reasoning closes the algorithmic gap (gap +0.008, CI covers zero) and Opus's code gap shrinks from +0.274 to +0.096. §7.2 also lists constrained decoding, decoding sweeps, and tool-augmented pipelines as untested. If real verifier-authoring pipelines use reasoning or tools, the practical deficit may shrink materially. The paper is explicit about this in places, but the headline and abstract lean harder than the evidence supports. I would ask for a softer deployment claim or a measured survey of actual practice.\n\nSecond soft spot: reproducibility. No public code or data link is visible in the text. The appendix references a released artifact and a receipt, but no URL or commit hash. For a paper whose central measurements depend on exact prompts, seeds, and parsing, that is a real brake.\n\nThe field evidence is honestly labeled observational and judge-relative; the 10:1 ratio is directional, as the authors say. The EvalPlus oracle premise is a premise, but the complete-truth arms carry the core comparison, so it is a secondary concern.\n\nWho is this for: people building RLVR verifiers, test generators, answer keys, rubrics; anyone taking \"LLM-as-examiner\" seriously. It deserves a serious referee. I would send it to review, with a clear request for the artifact and a rewrite of the deployment-scope wording. The core result will stand.","headline":"Core finding is solid and well-triangulated; the 'usually deployed one-shot' scope is asserted, not measured, and the paper's own data show reasoning closes the algorithmic gap.","tokens_in":49072,"tokens_out":2573,"would_cite":true,"duration_ms":28554,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLMs promoted to examiners judge membership well but do not reliably author the acceptable sets those roles demand; the omission-dominated gap persists across scales and costs accuracy inside RLVR rewards.","keywords":["LLM-as-judge","verifier authoring","acceptable sets","silent omission","judging-enumerating gap","test-suite generation","RLVR rewards","answer-key quality"],"falsifier":"Give reviewers the drafted keys with the full candidate universe and the correct cardinality of the oracle set supplied, then measure whether planted omissions are recovered at the same rate as planted over-inclusions, net of clean-key false alarms: Observation 2 predicts omissions stay undiscoverable because recovering a witness is itself the authoring operation, so a comparable recovery rate would falsify the witness-poor mechanism and the directional-review consequence built on it. A second, cheaper check: if any one-shot greedy model at any scale authors the complete-truth sets at F1 match","tokens_in":48153,"feed_emoji":"✂️","tokens_out":18650,"duration_ms":165084,"temperature":0.7,"pith_summary":"This paper tries to establish that language models, increasingly promoted from examinees to examiners, are weak at the very capability the promotion assumes: materializing the acceptable set a correctness specification denotes. Under one-shot greedy authoring—the protocol by which answer keys, test suites, and reward functions are actually produced at scale—models judge membership of a presented candidate well but enumerate the set or write the suite that induces it poorly, and the gap persists across a $24\\times$ parameter range without closing. A control locates the deficit precisely: asked to state the rule as an executable predicate rather than list its extension, the same models reach F1 ≈ 0.99, so the failure is not knowledge or specification but one-shot materialization of the induced region. The dominant error is silent omission, which resists audit because a missing member, unlike an over-inclusion, is an absence no reviewer can name; planted over-inclusions are caught 6–7$\\times$ more often than planted omissions. Wired into RLVR, an authored key costs 1.9 points of accuracy against a complete oracle and 18.5 WordNet-relative points, in every one of six paired seeds.","feed_headline":"24× scale range: LLM answer-key authoring never catches up to judging","feed_subtitle":"Models judging at F1 0.90 author suites admitting only 19–42% of correct code; omissions are 6–7× harder to audit.","key_machinery":"The central object is the paired 'two interfaces to the same weights': execution, a one-pass membership function $\\hat{m}(x)\\in\\{0,1\\}$ on a presented candidate (Eq. 1), versus authoring, autoregressive emission of a set $\\hat{S}$ under a causal mask with a learned stopping rule (Eq. 2), both scored by set F1 against a constructed oracle $S^*$. Three devices carry the argument: Observation 1 (the measured gap is a set-emission penalty); Observation 2 (omission is witness-blind to local review—even a perfect membership oracle cannot certify recall over candidates the reviewer never names, because finding a missing member is itself the authoring operation); Observation 3 (the authored reward i","core_discovery":"On the paper's own terms: language models execute membership far better than they author artifacts whose induced acceptance region matches the target set $S^*$, except where authoring reduces to restating a compact rule already in the prompt. Measured on four judge-free constructions (mechanical predicates, HumanEval+/MBPP+ execution, WordNet), the gap is +0.34 to +0.29 F1 over a $24\\times$ scale range on complete truth; on code, models judging at F1 0.74–0.90 author suites admitting only 19–42% of oracle-correct solutions. The residual error is omission-dominated and witness-poor: planted over-inclusions are caught 6–7$\\times$ more often than planted omissions, and subtractive review pushes","pith_inferences":["If the asymmetry is as structural as the paper argues, the same directional bias should be measurable in human-authored specifications: the paper's Observation 2 is an information-theoretic statement, not an LLM-specific one, so recall auditing against constructed ground truth—rather than more review passes—is the transferable prescription for any specification pipeline.","The frontier reasoning result (GPT-5.1's algorithmic gap falling to +0.008 with a CI covering zero at effort=low) suggests the deficit is as much a property of the one-shot greedy decoding protocol as of the weights; a natural test is whether small reasoning budgets close the gap across models and scales, and whether the RLVR tax shrinks correspondingly when keys are authored with reasoning.","The trace-repair decomposition—models choose discriminating inputs well but compute expected outputs badly—isolates the intensional half of suite authoring as the tractable target for fine-tuning or tool augmentation; one could test whether feeding the candidate universe or giving execution feedback during generation closes the enumeration gap on the open-set constructions.","Taken with the sample-then-verify falsification, the paper implies that model self-checking cannot certify its own authoring, shifting the burden to external oracles; deployments that cannot exhibit a known-correct behavior (the paper's own gate precondition) face an unsolved verification gap that warrants direct study."],"forward_implications":["Model-authored answer keys, rubrics, and test suites produced one-shot without test-time reasoning are systematically incomplete in the omission direction, so any pipeline that treats them as ground truth inherits under-acceptance: over-rejection of correct behavior rather than permissiveness.","Review and self-critique do not repair this; they make it worse. Subtractive review is a directional filter that removes visible over-inclusions while omissions persist, walking the artifact monotonically toward higher precision and no higher recall.","A known-correct-probe gate is the cheapest real mitigation: discarding any authored verifier that rejects a verified-correct solution cuts false rejection from 58–92% to ≤5%, at a yield of 5–39%; most of the discarded majority is recoverable by rewriting wrong expected values to what reference execution returns (3.3–10.6× yield).","Where a compact executable rule exists, the deficit is escapable: asking for the predicate rather than the roster is the difference between F1 ≈ 0.26 and ≈ 0.99, and test-time reasoning closes the algorithmic gap at the frontier.","As an RLVR reward, the authored key costs measured accuracy (1.9 points against a complete oracle, 18.5 WordNet-relative), a two-sided channel that withholds reward from correct outputs and grants it to incorrect ones."],"supporting_citations":[{"why":"Supplies the hardened executable oracle (EvalPlus 'plus' suites over HumanEval/MBPP) that defines oracle-correct code for the executable construction, the gate, and the repair arm.","marker":"[23]"},{"why":"Supplies WordNet synsets, the third reference construction; its recall side is incompleteness-robust and its precision is treated as a lower bound.","marker":"[26]"},{"why":"The Qwen2.5 family whose 3B–72B scale ladder carries the central claim that the authoring–execution gap persists over a 24× range.","marker":"[29]"},{"why":"Frames RLVR with noisy and imperfect verifiers as a channel-parameter problem; the paper supplies the direction and magnitude of machine-authored verifier noise.","marker":"[3]"},{"why":"Frames autoregressive emission of a set under a causal mask with a stopping rule, the mechanism the paper identifies as the authoring failure.","marker":"[37]"},{"why":"Documents incomplete decoding and stopping failures in sequence models, the enumeration-without-lookahead defect bundled into the emission term.","marker":"[39]"},{"why":"Program-aided prompting: the technique the intensional control (§3.11) uses, where emitting an executable predicate instead of a roster reaches F1 ≈ 0.99 and localizes the deficit to enumeration.","marker":"[58]"},{"why":"Documents benchmark suites rejecting valid solutions, motivating the discipline of gating any authored verifier on known-correct probes (§3.5).","marker":"[7]"},{"why":"The GRPO objective used in the RLVR propagation arms, the optimizer whose training loop inherits the authored key's starvation and pollution.","marker":"[32]"},{"why":"Reports that LLM-generated test oracles encode actual rather than expected behavior, independently corroborating the paper's modal code failure: over-specification by invented requirements.","marker":"[20]"}],"fun_headline_variants":["LLMs judge better than they author: omissions resist audit","Answer-key authoring gap: 0.29–0.34 F1 behind judging at 24x scale","LLM authors admit only 19–42% of correct code, omissions hidden","Omissions 6–7× harder to catch than over-inclusions in LLM sets","Judging vs authoring: LLM suite yield lags despite 24x scale"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise is the scope condition stated first in the paper—that verifiers are authored one-shot with greedy decoding and no test-time reasoning (the paper's own §3.7 shows frontier reasoning closes the algorithmic gap, authoring F1 rising 0.673 to 0.976, so pipelines that reason, ensemble, fine-tune, or use tools may escape the measured deficit); a second acknowledged premise is that the EvalPlus/HumanEval+ oracle correctly defines code behavior.","fun_headline_variants_meta":{"raw":{"variants":["LLMs judge better than they author: omissions resist audit","Answer-key authoring gap: 0.29–0.34 F1 behind judging at 24x scale","LLM authors admit only 19–42% of correct code, omissions hidden","Omissions 6–7× harder to catch than over-inclusions in LLM sets","Judging vs authoring: LLM suite yield lags despite 24x scale"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000296,"raw_usage":{"total_tokens":1665,"prompt_tokens":964,"completion_tokens":701,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":708,"completion_tokens_details":{"reasoning_tokens":590}},"tokens_in":708,"tokens_out":701,"duration_ms":8321,"temperature":1.0,"reasoning_tokens":590,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T00:34:10.714847+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give reviewers the drafted keys with the full candidate universe and the correct cardinality of the oracle set supplied, then measure whether planted omissions are recovered at the same rate as planted over-inclusions, net of clean-key false alarms: Observation 2 predicts omissions stay undiscoverable because recovering a witness is itself the authoring operation, so a comparable recovery rate would falsify the witness-poor mechanism and the directional-review consequence built on it. A second, cheaper check: if any one-shot greedy model at any scale authors the complete-truth sets at F1 match","supporting_citations":[{"cited_title":"Is your code generated by ChatGPT really correct? rigorous evaluation of large language models for code generation","cited_arxiv_id":null,"evidence_quote":"Supplies the hardened executable oracle (EvalPlus 'plus' suites over HumanEval/MBPP) that defines oracle-correct code for the executable construction, the gate, and the repair arm."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies WordNet synsets, the third reference construction; its recall side is incompleteness-robust and its precision is treated as a lower bound."},{"cited_title":"Order matters: Sequence to sequence for sets","cited_arxiv_id":null,"evidence_quote":"Frames autoregressive emission of a set under a causal mask with a stopping rule, the mechanism the paper identifies as the authoring failure."},{"cited_title":"Con- sistency of a recurrent language model with respect to incomplete decoding","cited_arxiv_id":null,"evidence_quote":"Documents incomplete decoding and stopping failures in sequence models, the enumeration-without-lookahead defect bundled into the emission term."},{"cited_title":"Jimenez, John Yang, Kevin Liu, and Aleksander Madry","cited_arxiv_id":null,"evidence_quote":"Documents benchmark suites rejecting valid solutions, motivating the discipline of gating any authored verifier on known-correct probes (§3.5)."}],"review_version":1}