{"id":"baf00056-ccf2-4223-938a-a00dd1d14584","arxiv_id":"2607.17044","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"In a production agent, scaffolding and specialists deliver most of the +7-15pp reliability uplift; the verification loop's isolated contribution is +1.5pp on SpreadsheetBench yet rescues 6 tasks that separate a mid-pack from a near-top result.","lead":"One production enterprise agent's reliability gains over its base model are measured across three public benchmarks and decomposed layer by layer: most of the uplift comes from scaffolding, routing, and trained specialist models, while the verification loop itself adds only a small but positionally decisive slice. This is a vendor-authored evaluation of the vendor's own system, with public run data but proprietary model weights, so the numbers carry an asterisk pending indepe","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"GAIA layer decomposition (Table 5) rests on unrecorded internal single-run estimates (~60% base, ~70% planner–executor, ~74% routing); if controlled ablations move these materially, the cross-benchmark claim that scaffolding/routing—not verification—dominates is unsupported outside SpreadsheetBench.","rationale":"Reader's weakest assumption and mine align. The SpreadsheetBench decomposition is the paper's best evidence: full system 91.25% vs prompt+scaffold 89.75% gives loop +1.5 pp from six rescues, supported by per-task confusion matrix and released run records. The BullshitBench result establishes large whole-firewall uplift but no layer decomposition. The GAIA numbers are the only source for routing and specialist contributions, and they rest on internal estimates the paper itself flags as unrecorded and indicative. If those estimates are off, the paper's headline cross-benchmark conclusion overstates the evidence; the central claim would reduce to 'on spreadsheets, most uplift is scaffolding, verification adds a small decisive margin.' That is still a meaningful finding, but not the cross-benchmark decomposition announced in the title and abstract. I therefore keep the reader's CONDITIONAL verdict: accept the well-instrumented component, require the controlled GAIA ablation (and ideally a same-harness SpreadsheetBench baseline and matched same-model scaffold baselines) before the generalized claim is treated as established.","tokens_in":14592,"tokens_out":7621,"duration_ms":71291,"concrete_test":"Run a pre-registered GAIA validation ablation on the same 165-item split with the frozen production harness: (a) base frontier model + minimal tool harness, (b) +planner–executor, (c) +routing, (d) +verification loop; three complete runs per tier, pass@1 selection rule fixed in advance, trajectories and selection logs published. Compare measured tier increments to Table 5's ~60→70→74→75.2. If the bare-tool base differs by >5 pp, any tier delta differs by >3 pp, or the verification-only increment is not ~1 pp, the cross-benchmark decomposition as published is not supported; if the tiers reproduce within those bounds, the concern is settled.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing weak point is the GAIA layer decomposition that carries the cross-benchmark version of the central claim. Table 5 and §7 report GAIA uplift as: ~60% bare-tool base → ~70% with planner–executor → ~74% with routing → 75.2% full, with the verification loop adding ~+1 pp. The paper explicitly states that these GAIA structure tiers come from 'internal ablation runs whose selection rules were not recorded' (Table 5, §7) and that the ~60% base is 'an internal single run ... should be read as an estimate' (§6.1); the controlled GAIA ablation is 'forthcoming' (§7). The central claim says most reliability comes from scaffolding, routing, and specialist models rather than verification, and routing/specialist contributions are quantified only on GAIA. If the true bare-tool baseline were higher, or the planner–executor/routing deltas were materially smaller, the 'most comes from structure/routing' conclusion would survive only for SpreadsheetBench, where it is well instrumented (prompt+scaffold 89.75% vs full 91.25%, §7). The paper's own caveat that the GAIA loop-isolated figure is 'indicative only' means the +15 pp GAIA headline and the routing contribution are not yet evidence. This is a missing-evidence concern, not an internal inconsistency; the SpreadsheetBench decomposition and its released run records remain credible.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies Leni, a production enterprise agent whose reliability layer consists of verification loops (execute, observe, compare, correct) staffed by small post-trained specialists. On three public benchmarks it reports total uplifts over a frontier base model: +11.0 pp on SpreadsheetBench Verified (91.25% vs. 80.25%), +7–10 pp on BullshitBench, and ~+15 pp on GAIA validation (75.2% pass@1 vs. ~60% estimate). The central contribution is a decomposition attributing most of the uplift to scaffolding, routing, and specialist models, with the verification loop contributing a small (+1.5 pp on SpreadsheetBench), positionally decisive increment. The paper instruments the loop end-to-end, yielding a task-level verifier confusion matrix (c≈0.20, r≈0.75, f=0), proposes a compounding-reliability model, reports a valid-premise control bounding firewall over-rejection at ≲3.6%, and gives preliminary specialist-swap ablations suggesting that observer independence matters. It is candid about limitations: vendor evaluation, unpaired GAIA baselines, unrecorded internal ablation tiers, and missing controlled comparisons.","tokens_in":14905,"tokens_out":7056,"duration_ms":65864,"significance":"If the decomposition holds, this is a valuable contribution: it shifts the reliability discussion from base-model choice to architecture, gives empirically estimable loop parameters (c, r, f), and states a falsifiable chain-length prediction. The SpreadsheetBench instrumentation is genuinely strong — 397 loop-triggering tasks, a confusion matrix, run-to-run sensitivity bounds, and a released run record that includes unfavorable and superseded runs. The GAIA re-grading script reproducing all previously stored grades is a concrete reproducibility asset. The paper's reporting discipline is above the norm for vendor evaluations. However, the significance is conditional: the GAIA layer decomposition, which carries the cross-benchmark version of the central claim, rests on internal single-run estimates whose selection rules were not recorded. The stress-test concern about this point lands. The full run records and audit scripts are a real strength, and the paper explicitly enumerates the missing experiments needed to strengthen its claims.","major_comments":[{"comment":"The cross-benchmark version of the central claim is not currently supported by the evidence. The GAIA layer decomposition is built from an internal single-run bare-tool base (~60%) and internal ablation tiers (~70% planner–executor, ~74% routing) whose selection rules were not recorded; §7 itself calls the GAIA loop-isolated increment (~+1 pp) 'indicative only.' Because the routing/specialist contributions are quantified only on GAIA, the abstract's claim that 'most of it comes from scaffolding, routing, and specialist models' is, outside SpreadsheetBench, an estimate rather than a measured decomposition. This is a missing-evidence problem, not an internal inconsistency: the SpreadsheetBench decomposition and its run records are credible. The revision should either run the controlled GAIA ablation with recorded selection rules, per-tier runs, and confidence intervals, or explicitly restr","section":"§7 / Table 5 / §6.1"},{"comment":"The headline SpreadsheetBench uplift and the +9.5 pp 'structure' share both depend on a baseline that is not harness-matched. The 80.25% figure is the leaderboard's Claude Opus 4.6 entry with a minimal three-line prompt, not a run of the same model under Leni's production prompt, sandbox, and evaluation harness; the paper itself says the comparison is unpaired. Since the total uplift is defined as the full system minus this external number, any difference between leaderboard conditions and the production harness changes the decomposition. A same-harness bare-model rerun is needed to make the +11.0 pp and the structure share quantitative; at minimum, a sensitivity analysis over plausible base rates should be reported.","section":"§5.1 / Table 3"},{"comment":"The specialist-swap evidence is currently too weak to carry the 'observer independence' conclusion. The swaps are single runs, cover only Cell-S and Triage-S, and omit the condition that would separate independence from specialization: an independent frontier verifier from a different provider. With this design, the drop from six rescues to two could be due to the verifier's post-training or its smaller size rather than to independence. The paper acknowledges this in §10, but the abstract and §7 still present the swap result as supporting the 'who observes matters' claim. The revision should either add the missing verifier condition or present the claim strictly as a preliminary hypothesis and soften the abstract and conclusion accordingly.","section":"§7 / §10"}],"minor_comments":[{"comment":"The phrase 'word8/13-gram containment' appears to be a typo; it should read 'word / 8-gram / 13-gram containment.'","section":"§5.4"},{"comment":"The sentence beginning 'The GAIA figures correct an earlier company report whose 77.6%...' is grammatically awkward; clarify that the earlier report mixed selection rules across tiers, not that the figure itself was a selection rule.","section":"§6.3"},{"comment":"The denominator in 'catch rate c = 8/40 = 0.20' is not immediately obvious from the table. Add one sentence noting that 40 = 32 missed errors + 8 flagged errors, so the denominator is all erroneous artifacts.","section":"Table 4 / §6.2"},{"comment":"The statement that '+1.5 pp matches (1−p)cr within rounding' should be explicitly labeled as a consistency check rather than a fit of Eq. (1), since p is derived from the same data. This is already implied but could be made explicit.","section":"§6.2"}],"recommendation":"major_revision","confidential_remarks":"This is an unusually transparent vendor evaluation, and the conflict of interest is fully disclosed. The released artifacts (run records, audit scripts, re-grader) are a genuine asset. My recommendation of major revision is driven by the gap between the title's 'cross-benchmark decomposition' and the evidence: the GAIA layers are explicitly unrecorded internal estimates. I would ask the authors to either provide the controlled GAIA ablation or narrow the claim. I do not see evidence of selective reporting; the paper includes a self-correction of its own earlier GAIA figure and preserves unfavorable runs. The remaining contamination-sweep gaps (workbook-trace and tool-trajectory corpus components) are on attestation; reviewers should verify the published pseudonymous internal-account labels if possible."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: the SpreadsheetBench instrumentation is the real contribution, and it's solid. The verifier confusion matrix (c≈0.20, r≈0.75, f≈0) from a production deterministic loop is genuinely new; the layer decomposition on that benchmark (structure +9.5pp, loop +1.5pp from 6 rescues) is carefully done and backed by released run records. The paper also deserves credit for correcting its own earlier GAIA number, disclosing a degraded run, and publishing audit scripts.\n\nThe soft spot is exactly where the stress-test puts it: the GAIA layer decomposition. The ~60% base and the ~70%/~74% structure tiers are internal single runs with unrecorded selection rules; the controlled ablation is 'forthcoming.' That means the cross-benchmark claim that scaffolding and routing dominate the uplift rests on evidence that is currently absent. The SpreadsheetBench decomposition stands on its own, but the GAIA numbers are context, not evidence. The specialist-swap results are also single-run and lack the independent-generalist control that would separate 'independence' from 'specialization' — the paper says this plainly, and it's why the headline claim about observers should be treated as a hypothesis, not a finding.\n\nOne more caveat: the SpreadsheetBench baseline is an unpaired leaderboard entry, not a harness-matched rerun. That matters for the +11pp headline, though the internal scaffold comparison (89.75% vs 91.25%) doesn't depend on it.\n\nWhat this paper is for: anyone building enterprise agents will get concrete guidance — structure before verification, cheap independent observers, instrument the loop. Researchers studying self-correction will get a useful empirical data point and a clean decomposition methodology, even if the proprietary weights limit replication.\n\nRecommendation: yes, send it to peer review. The core decomposition is measureable and measured; the GAIA gaps are missing evidence, not internal incoherence. Referees should ask for the controlled GAIA ablations and structured loop telemetry before accepting the broader claims, but the SpreadsheetBench result should survive.","headline":"The SpreadsheetBench decomposition is the real contribution; the GAIA layer numbers are not yet evidence.","tokens_in":15467,"tokens_out":2120,"would_cite":true,"duration_ms":20194,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper decomposes the reliability gain of a production enterprise agent and finds that most of it comes from scaffolding, routing, and specialist models, not from the verification loop itself, whose isolated contribution is small but po","keywords":["agent reliability","verification loops","scaffolding","routing","specialist models","benchmark decomposition","self-correction","enterprise AI"],"falsifier":"Run a controlled ablation that holds base model, tools, and token budget fixed and varies only the verification loop's presence; if the loop's isolated contribution on SpreadsheetBench exceeds roughly +1.5 points, or if removing scaffolding does not eliminate most of the uplift, the decomposition is wrong. A second check: swapping in an independent generalist verifier from a different provider would test whether independence rather than specialization drives the rescues.","tokens_in":14398,"feed_emoji":"🧩","tokens_out":4644,"duration_ms":40769,"temperature":0.7,"pith_summary":"The paper asks where a production enterprise agent's reliability actually comes from. Evaluating one deployed system on three benchmarks that stress different failure modes, it finds large total gains over a bare frontier model. The central finding is a decomposition: most of the uplift comes from architecture—prompting, planning, routing, and lightweight specialist models—rather than from the verification loop's isolated contribution, which measures about +1.5 percentage points on the spreadsheet benchmark. That small contribution is concentrated at the top of the score distribution, converting otherwise-failing tasks into passes, and preliminary ablations suggest the loop's value depends on the observer being independent of the generator. If correct, teams should invest in structure first, staff observation with small independent specialists, and treat verification as tail-end infrastructure rather than the primary reliability lever.","feed_headline":"Scaffolding adds 9.5 points; verification loops add 1.5","feed_subtitle":"In a production enterprise agent, most reliability gains come from architecture; the loop rescues only the last few failures.","key_machinery":"The four-stage verification loop (execute, observe, compare, correct), with observation as the load-bearing stage, staffed by lightweight post-trained specialists in the 0.5–4B parameter range. The loop is formalized in a compounding-reliability model with catch rate c, fix rate r, false-alarm rate f, and breakage rate b; the paper reports what it describes as the first empirical production estimate of (c, r, f) ≈ (0.20, 0.75, 0). The verification oracle class—deterministic re-execution, self-reflective triage, or planner-mediated typed artifacts—sets the reliability ceiling, while the independence of the observer sets the floor.","core_discovery":"The durable contribution is the decomposition: across three unrelated failure modes, total architectural uplift is +7 to roughly +15 points over the bare base model, of which the verification step itself contributes only about +1.5 points on SpreadsheetBench (measured) and roughly +1 point on GAIA (estimated against internal tiers). The loop's value is positional: it rescues tasks at the top of the distribution (6 of 400 on SpreadsheetBench), which is the difference between a mid-pack and a near-top result. An instrumented verifier confusion matrix shows catch rate about 0.20, fix rate 0.75, and false-alarm rate about 0, and a specialist-swap ablation indicates that replacing the small train","pith_inferences":["The positional-concentration result suggests a general principle: verification loops are tail-end infrastructure, valuable only when the base architecture is already strong; teams with weak scaffolding should not expect loops to rescue them.","The untested chain-length prediction of the reliability model—that loop value grows with the number of dependent steps—implies verification loops will matter most for long-horizon enterprise workflows; a controlled experiment varying chain length would confirm or refute this directly.","If the independence hypothesis holds, a plausible extension is to use a verifier from a different model family or provider as a cheap drop-in replacement for specialist post-training; this is testable with the same swap-ablation design.","The measured false-alarm rate of about zero on this distribution may not hold under adversarial valid premises; the paper's own matched-style control would be the natural stress test."],"forward_implications":["For a fixed base model, architectural structure (planning, routing, typed interfaces, prompting) is the largest lever on reliability, worth roughly +9.5 points on spreadsheets before verification is added.","The verification loop is worth adding even when its marginal gain is small, because it operates at the top of the score distribution where the remaining failures live; at a leaderboard or SLA tail, it can decide pass versus fail.","The loop's value depends on who observes: a small trained specialist that did not generate the artifact outperforms the generating frontier model at catching errors, consistent with documented self-assessment bias.","Instrumenting the loop turns design from folklore into measurement: catch, fix, and false-alarm rates are estimable quantities that pinpoint where the next reliability gain will come from (raising catch rate is worth up to +8 points; fix rate is nearly saturated).","A commodity deterministic oracle (headless spreadsheet recalculation) closes most of the gap to a custom neurosymbolic runtime, localizing the remaining cost in LLM-mediated comparison rather than re-execution."],"fun_headline_variants":["Scaffolding adds 9.5 points; verification loops add just 1.5","Reliability comes from scaffolding, not verification loops","Verification loops rescue top tasks; scaffolding provides the rest","Agent reliability: decomposition shows scaffolding dominates","Scaffolding vs verification: 9.5 vs 1.5 points in agent reliability"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The GAIA decomposition rests on internal, unrecorded baseline estimates: the roughly 60% bare-tool-use figure and the roughly 70%/74% structure tiers come from single internal runs whose selection rules were not recorded, so the GAIA layer split and the roughly +15-point uplift could be off if those estimates are not representative.","fun_headline_variants_meta":{"raw":{"variants":["Scaffolding adds 9.5 points; verification loops add just 1.5","Reliability comes from scaffolding, not verification loops","Verification loops rescue top tasks; scaffolding provides the rest","Agent reliability: decomposition shows scaffolding dominates","Scaffolding vs verification: 9.5 vs 1.5 points in agent reliability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001146,"raw_usage":{"total_tokens":4671,"prompt_tokens":902,"completion_tokens":3769,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":646,"completion_tokens_details":{"reasoning_tokens":3679}},"tokens_in":646,"tokens_out":3769,"duration_ms":23979,"temperature":1.0,"reasoning_tokens":3679,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T19:10:51.884896+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a controlled ablation that holds base model, tools, and token budget fixed and varies only the verification loop's presence; if the loop's isolated contribution on SpreadsheetBench exceeds roughly +1.5 points, or if removing scaffolding does not eliminate most of the uplift, the decomposition is wrong. A second check: swapping in an independent generalist verifier from a different provider would test whether independence rather than specialization drives the rescues.","supporting_citations":[],"review_version":1}