{"id":"89779612-2ea5-4c31-8e2b-380315db6e18","arxiv_id":"2607.13347","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Reference-free LLM judge scores failed to select better table-extraction outputs over eight regeneration iterations on FinTabNet and OmniDocBench; keeping the first output was safest.","lead":"This paper finds that LLM-as-a-judge scores are too weak and unstable to guide iterative table-recognition pipelines: no score-based selection policy beat simply keeping the first output, on two benchmarks. The result matters because LLM judges are increasingly used not just to evaluate but to drive self-improvement loops.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The §4.4 mechanism claim is vulnerable to the paper's own documented run-to-run non-replication; with multi-backend nondeterminism unconfirmed, the B/C severe-loss reduction may be a temporal artifact.","rationale":"The reader's weakest_assumption identifies the causal decomposition's reliance on near-deterministic temperature-0 generation. My review confirms this is the most load-bearing concern because the paper's own Limitations document non-replication of run-level means across re-forking runs, directly threatening the B/C contrast in §4.4 and the no-feedback control in §4.1. The selection-failure central claim does not depend on this decomposition: the high tie rates, low TEDS rank correlations (≤0.13), non-reproducible rankings (Kendall's W=0.36), and negative recovery under most policies stand independently of generation determinism. However, the paper's headline narrative includes the mechanism claim ('target-preservation failure'), and if the B/C reduction is a temporal artifact, the contribution shrinks to a cautionary negative result about selection, which the reader already judged as CONDITIONAL. I therefore agree with the reader's assessment and recommend no change to the verdict. The concrete test (interleaved repeated B/C runs plus determinism measurement) would settle whether the mechanism claim survives; absent that, CONDITIONAL remains appropriate given the unresolved serving confound and the missing verification package.","tokens_in":30442,"tokens_out":6509,"duration_ms":74203,"concrete_test":"Re-run the frozen-iter0 B/C contrast on the full 476 FinTabNet tables with B and C interleaved per table (alternating order) within a single batch, and repeat the entire re-fork three times, reporting severe-loss rates per run. If the C−B difference remains ≈ −2.7pp with similar McNemar p across all repeats, the mechanism claim holds. In parallel, quantify generation determinism: generate the same prompt (image + previous HTML + generic instruction) 20 times at temperature 0 via a single pinned backend, and compute the byte-identical output rate. If that rate is below ~90%, the near-deterministic assumption underpinning iteration-to-iteration attribution is violated, and the B/C contrast must be re-analyzed as a between-run comparison with temporal confound.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central selection-failure result is robust, but the paper's causal decomposition—that degradation stems from target-preservation failure rather than judge feedback—rests on paired contrasts executed across runs in a setting where run-level means do not replicate. The Limitations admit the unconstrained no-feedback condition's mean net effect was −0.0122 in the original 99-table paired experiment but +0.0000 and +0.0007 in later re-forking runs, and that multi-backend serving is a possible but unconfirmed cause. If routing or temporal drift shifts generation behavior between conditions, the B/C frozen-iter0 severe-loss reduction (17/476 to 4/476, p=0.0023) could be confounded: B and C may have been run at different times, and the paper does not report interleaving, randomization, or backend pinning for this contrast. The same concern weakens the claim that 'feedback content is not the source of net degradation' from the 99-table no-feedback control, since that comparison also spans time. This matters because the abstract and conclusion elevate target-preservation failure as a proximate mechanism; if that decomposition is unsettled, the paper's robust contribution narrows to the selection-failure negative result, which is independently supported by tie rates, weak rank correlations, and non-reproducibility of judge rankings.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies whether reference-free LLM-as-a-judge scores can serve as reliable optimization signals in closed-loop table recognition. Using a Gemini model as generator and judge, it runs eight-iteration regeneration loops on FinTabNet and OmniDocBench, with TEDS computed post hoc from ground truth as a deterministic audit metric. The paper reports three main findings: (1) judge scores are weak selection signals—frequent ties, low or negative rank correlations, non-reproducible rankings, negative or near-zero recovery relative to the first output, and no tested score policy improving over iter0 on both datasets; (2) severe degradation occurs even without judge feedback, with target-preservation failure under unconstrained regeneration as a proximate mechanism; (3) a structure-preserving instruction reduces the severe-loss rate in a frozen-iter0 paired contrast, significantly on FinTabNet and directionally on OmniDocBench. The paper concludes that evaluation-style evidence for LLM judges does not establish closed-loop optimization utility and that deterministic structural verification is needed.","tokens_in":30730,"tokens_out":7019,"duration_ms":88589,"significance":"If the central negative result holds, it is a valuable, well-scoped contribution: the task provides a deterministic external metric, and the authors run multiple independent diagnostics—tie rates, rank correlations, repeated-scoring repeatability, an independent TEDS implementation cross-check, a no-feedback control, an independent-candidate control, and artifact audits. These controls make the selection-failure result substantially more convincing than a single point estimate. The paper is also appropriately cautious in many places: it does not dispute LLM-as-a-judge as an evaluator, it discloses the development-set selection of the calibrated prompt, it reports the post hoc margin selection, and it includes a Limitations section that candidly documents non-replication. The main weakness is that the mechanistic claims about target-preservation failure and feedback content are built on across-run contrasts in a setting the paper itself shows to be temporally unstable. The selection-failure negative result is likely robust; the causal decomposition is not yet established at the same standard.","major_comments":[{"comment":"The proximate-mechanism claim rests on contrasts that are not protected against the paper's own documented run-to-run non-replication. The Limitations state that the original 2×2 unconstrained-condition mean (−0.0122) did not recur in later re-forking runs (+0.0000 and +0.0007), and that multi-backend serving is an unconfirmed possible cause. The full-n frozen-iter0 B/C severe-loss contrast (17/476 vs 4/476, two-sided exact McNemar p=0.0023) is paired by table, but the text does not report whether B and C were interleaved or randomized within the run, or whether backend allocation was pinned. If B and C were executed in separate time blocks, temporal drift could produce or inflate the severe-loss difference. The same concern applies to the no-feedback control in §4.1/Appendix B.6 (paired 99-table difference +0.0001), where the main loop and the no-feedback loop are separate runs. Please","section":"§4.4 and Limitations"},{"comment":"The central selection-failure claim is presented as a general property of the loop, but the policy comparisons in Table 1 are computed on a single main-run candidate stream. The paper's own Limitations record that the unconstrained-condition mean did not recur in later re-forking runs (means +0.0000 and +0.0007), and the full-n B condition in §4.4 has mean −0.0030 on FinTabNet versus −0.0200 in the main run. Thus 'no tested judge score policy improved on the first output on both datasets' is, strictly, a claim about one realization. The tie-rate, rank-correlation, and repeatability diagnostics make the selection-failure result plausible, but they do not by themselves establish that the recovery numbers in Table 1 would survive replication on a re-forked run. Please report a replication of the selection policies on a re-forked candidate stream, or explicitly restrict the claim to the obse","section":"§4.1 and Table 1; Limitations"},{"comment":"The guarded-acceptance result is used to support the overall conclusion, but the OmniDocBench improvement at δ=20 is selected post hoc from a grid, as the Limitations state. The FinTabNet δ=20 result is negative (−0.0020) and the margin rule mostly rejects revisions. Given that the post hoc selection is not accounted for in the confidence interval, the sentence in §4.2 that 'the best margin (δ=20, selected post hoc from the grid) improved mean TEDS by +0.0116' should be explicitly labeled exploratory, and the abstract's summary of the guarded-policy finding should not rely on this value. A simple multiple-comparison acknowledgment (e.g., reporting all six margins, as Appendix J does) is already present, but the main text should not present the selected margin as a confirmed policy without a fixed-analysis qualification.","section":"§4.4 and Appendix J"}],"minor_comments":[{"comment":"'Feedback content is not the source of net degradation' is too strong given the cross-run nature of the comparison; suggest rewording to 'not the source in this paired comparison' or adding the temporal-drift caveat where the sentence first appears.","section":"§4.1"},{"comment":"The 'Better'/'Worse' column compares against random-among-8, not against iter0. Since several policies are simultaneously 'better than random' and worse than iter0, the table would benefit from a footnote making this distinction explicit in each verdict cell.","section":"Table 1"},{"comment":"The statement 'Full determinism does not hold' appears only in an appendix footnote. Given that the paper's causal interpretation depends on near-determinism, this qualification should appear in the main-text experimental setup (Section 3.1) alongside the temperature-0 statement.","section":"§I.1"},{"comment":"The multi-backend routing is mentioned in the Limitations and Appendix I.1 but not in Section 3.1. Consider moving a one-sentence mention into the main text so the reader is not surprised by the later caveat.","section":"§3.1 / Appendix I.1"},{"comment":"The blinded author validation reports Cohen's κ=0.648 and a relabeling of several REAL items to GT-inconsistent. The main-text Section 4.3 cites the LLM-assisted labels without noting this revision; a pointer to the sensitivity analysis near the main claim would improve transparency.","section":"Appendix B.4"}],"recommendation":"major_revision","confidential_remarks":"The paper's own Limitations contain strong evidence that the run-level means do not replicate, and the main text's abstract and conclusion still elevate the mechanism claim. I think the selection-failure negative result is close to publishable, but the causal decomposition needs either a properly interleaved replication or an explicit demotion to 'hypothesis.' The authors seem aware of the issue; the fix is likely achievable within the manuscript's scope, so major revision rather than rejection is appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The main finding holds up: in this table-recognition loop, the LLM judge's scores were too weak and unstable to improve on the first output, and no policy—pointwise, pairwise, or margin-guarded—robustly beat the trivial iter0 baseline. That negative result is supported by converging evidence: high tie rates, low or negative rank correlations with deterministic TEDS, non-reproducible rankings across repeated scoring, and near-zero or negative recovery relative to oracle headroom. The no-feedback paired control and the independent-candidate contrast are careful, and the paper is unusually candid about its own limitations.\n\nWhat's genuinely new here is the controlled instantiation: a deterministic external metric used to audit every judge decision at instance level, across two datasets, with several judge families and scoring formats. This goes beyond the prior general warnings about judge alignment not implying selection utility (Landesberg, Pan, Zhou). The finding that feedback content contributes little to net degradation, while unconstrained regeneration itself causes structural breakage, is a useful boundary condition—even if the mechanism claim is not as secure as the selection-failure result.\n\nThe soft spots are real but not load-bearing. The §4.4 causal decomposition rests on a re-forked B/C contrast that, as the stress-test notes, is not documented as interleaved or backend-pinned. The paper itself flags run-level non-replication and multi-backend routing as unresolved. So the severe-loss reduction under copy-preserving instruction could be partly a temporal artifact. That doesn't sink the paper, because the central conclusion does not depend on this contrast—it's the secondary mechanism story. Also missing: the promised code/data artifacts, the post hoc δ=20 margin, and the dev-set selection of calibrated_v2. All are acknowledged in the text, but they add friction to full acceptance.\n\nWho gets value: anyone building closed-loop LLM systems with reference-free judges, and anyone doing table recognition with LLM-based refinement. The paper is a useful cautionary data point, not a definitive proof about all judges or all domains.\n\nRecommendation: send to peer review, and press the authors to ship the verification package and report run-interleaving/backend details for the B/C contrast. The main result deserves a serious referee; the mechanism claim needs to be pinned down or softened further.","headline":"The central negative result—reference-free judge scores don't drive selection in closed-loop table recognition—is solid and worth taking seriously; the secondary mechanism claim is softer but honestly hedged.","tokens_in":31231,"tokens_out":2287,"would_cite":true,"duration_ms":25013,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A judge LLM that can evaluate table outputs cannot reliably guide closed-loop regeneration: no tested score policy beat keeping the first output.","keywords":["LLM-as-a-judge","table recognition","closed-loop refinement","TEDS","selection utility","self-correction","reward hacking","iterative optimization"],"falsifier":"A judge configuration with near-perfect scoring reproducibility (e.g., exact agreement above 95% across repeated calls) and near-zero tie rates on TEDS-separated pairs, whose best-of-n or margin-guarded selection then beats keeping the first output on both benchmarks, would overturn the central claim.","tokens_in":30283,"feed_emoji":"📊","tokens_out":8259,"duration_ms":83156,"temperature":0.7,"pith_summary":"This paper asks whether an LLM judge that can evaluate outputs can also guide their iterative improvement—an assumption behind many closed-loop pipelines. In a table-recognition setting where a deterministic structural metric (TEDS) audited every decision, the answer is no: across two benchmarks, judge scores tied frequently, rankings were not reproducible, and no tested selection or margin-guarded policy improved on simply keeping the first output. Better candidates existed, but judge policies recovered at most a small fraction of the available improvement. Severe breakage also occurred without judge feedback, indicting regeneration that fails to preserve the target structure; a structure-preserving instruction reduced the worst failures but never raised mean quality. If right, this means evaluation-style evidence alone is not enough to justify using judge scores as optimization signals, and iterative refinement needs a deterministic structural verification signal, not just LLM judgments.","feed_headline":"No judge-score policy beat the first output in table recognition","feed_subtitle":"Scores tied, rankings shifted, and iterative refinement damaged tables; a deterministic verifier is the missing piece.","key_machinery":"The central object is a closed-loop regeneration pipeline whose every decision is audited by TEDS, a deterministic tree-edit-distance similarity between the generated HTML table and fixed ground truth. TEDS converts the otherwise fuzzy notion of table quality into a fixed benchmark objective, so the paper can quantify, per table and per iteration, how far each judge decision deviates from the target. The argument is carried by two crossed contrasts, both forking from the same stored initial output: a feedback-presence contrast (judge error list injected vs. generic instruction only) and a structure-preservation contrast (unconstrained regeneration vs. copy-preserving regeneration instructed","core_discovery":"An eight-iteration closed loop regenerates HTML tables from images: a reference-free LLM judge scores each candidate and lists errors; the generator revises from that feedback; ground truth never enters; the deterministic tree-edit metric TEDS audits every decision. The judge signal is weak: scores tie on tens of percent of distinct candidate pairs, rankings are not reproducible, and no judge configuration, scoring format, or tie rule beats the trivial always-iter0 baseline on both datasets; margin-guarded acceptance does not help. Removing feedback leaves net degradation unchanged and severe losses persist, indicting target-preservation failure under unconstrained regeneration. A structure-","pith_inferences":["Editorial: If this pattern holds beyond the two benchmarks, evaluation-score alignment with humans should not be taken as evidence that the same scores can steer optimization; applications should first audit within-instance ranking reproducibility (tie rates, repeated-scoring agreement) before building a loop around a judge.","Editorial: The convention-conflict mechanism suggests a testable extension: the breakage rate should scale with the divergence between the target dataset's ground-truth HTML conventions and the markup conventions the judge prefers; measuring that divergence could predict which tables are at risk.","Editorial: The exploratory generator trend (larger oracle headroom associated with positive loop benefit) suggests a practical decision rule for practitioners: measure baseline error on a small labeled audit and enable closed-loop refinement only for generators/inputs with large headroom.","Editorial: A deterministic structural guard could be tested directly: add a checker that rejects any revision whose row/column count or merge structure differs from the previous candidate, and measure whether the loop's tail losses disappear."],"forward_implications":["Judge scores should not gate deployment in reference-free iterative extraction: keeping the first output was the safest baseline on both benchmarks.","Guarded margin rules are not a safety net; the margins that avoided degradation accepted almost no revisions, so they merely replicated the first-output policy.","A neutral average loop outcome can hide severe tail damage: 4–12% of tables suffered large losses, so mean metrics alone are insufficient for deciding whether to deploy a refine loop.","Regeneration itself, not the content of judge feedback, is the main source of damage; preventing structural changes during revision reduces the worst failures.","Iterative refinement needs a deterministic structural verification signal (e.g., checking row/column counts and merge structure) rather than judge scores alone."],"fun_headline_variants":["LLM judge scores can't drive table refinement","Table loop needs deterministic verifier, not judge","Weak LLM judge scores sink closed-loop tables","Iterative table repair with LLM judge fails"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The causal decomposition assumes that temperature-0 generation and scoring are near-deterministic; the paper concedes full determinism does not hold and that using more than one serving backend may account for part of the scoring nondeterminism, so if routing or sampling noise is substantial, attribution of changes to feedback weakens.","fun_headline_variants_meta":{"raw":{"variants":["LLM judge scores can't drive table refinement","Table loop needs deterministic verifier, not judge","Weak LLM judge scores sink closed-loop tables","Iterative table repair with LLM judge fails"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000545,"raw_usage":{"total_tokens":2462,"prompt_tokens":778,"completion_tokens":1684,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":522,"completion_tokens_details":{"reasoning_tokens":1625}},"tokens_in":522,"tokens_out":1684,"duration_ms":11957,"temperature":1.0,"reasoning_tokens":1625,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T05:24:51.809919+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A judge configuration with near-perfect scoring reproducibility (e.g., exact agreement above 95% across repeated calls) and near-zero tie rates on TEDS-separated pairs, whose best-of-n or margin-guarded selection then beats keeping the first output on both benchmarks, would overturn the central claim.","supporting_citations":[],"review_version":1}