{"id":"cd0b2216-57c5-4846-bb30-919d8ae208e5","arxiv_id":"2505.14131","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A controlled study shows that for small models, the best table representation depends on table size and question complexity, and the proposed FRES selection rule improves accuracy by about 10 points on average.","lead":"This paper compares whether table question answering systems should read tables as text or as images, and finds the best choice depends on model size, table size, and question difficulty. The authors propose a simple selection rule, FRES, that chooses the input format automatically and improves average exact-match accuracy by about 10 points on tested benchmarks.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 10% FRES gain in Table 4 is driven by TableLlaVA's unusually poor image+text baseline; against the best single-representation baseline the average gain is only ~2.4 EM points.","rationale":"The paper's empirical analysis (Figure 1, Table 9) is carefully controlled, and the directional findings about representation and model choices are supported across six small models; I do not object to that part. The vulnerable part is the FRES quantitative claim, which is the headline number in the abstract and Section 4. The reader's weakest-assumption (the question classifier) is real but less threatening: classifier noise would tend to dilute FRES's advantage rather than create it, and the paper's own error analysis attributes only 19% of FRES errors to classification. In contrast, the 10% figure is arithmetically true only because TableLlaVA's combined-input baseline collapses relative to its text baseline; for Pixtral the gain is about 1.1 EM points. This makes the central quantitative claim dependent on a single model's handling of an input format that may be out-of-distribution for it. It does not invalidate the study, but it means the paper should be conditional on a re-analysis of Table 4 with fair baselines and per-model reporting. Since the reader's verdict is already CONDITIONAL, my read does not change that verdict, but it identifies a different and more load-bearing reason for the condition.","tokens_in":13515,"tokens_out":15047,"duration_ms":133100,"concrete_test":"Recompute the Table 4 evaluation with two changes: (1) use the best single-representation baseline max(EM_text, EM_image) per model and dataset instead of the t,i baseline; (2) for TableLlaVA, inspect the released MMTab/TableLlaVA code to verify whether image+text was ever a training input, and if not, run the t,i condition with a prompt format consistent with its training. If the average FRES advantage remains about 10 EM points after these corrections, the concern is resolved; if it drops to about 2-3 points, the 10% claim should be rephrased as model-specific and the paper should report per-model gains.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4 reports 'an average of 10% of EM gain' for FRES versus the 'both representations' (t,i) baseline, averaged over Pixtral and TableLlaVA on WTQ, TabFact, HiTab, and WikiSQL. This average is dominated by TableLlaVA: TL-FRES reaches 38.8, 49.7, and 48.4 on WTQ, HiTab, and WikiSQL, while TL-t,i reaches only 17.2, 16.3, and 29.5, far below TL-t (34.4, 44.3, and 47.8). For Pixtral, the same FRES-versus-t,i differences are +1.9, -0.5, +1.8, and +1.3 EM points, averaging about +1.1. The large TL-t,i deficit suggests TableLlaVA, a fine-tuned MLLM, was not trained to consume image and text simultaneously, so the t,i condition may be an unfair baseline for this model. If that condition is removed or re-prompted consistently, the headline improvement drops to roughly 2.4 average EM points over the best single-representation baseline. Thus the paper's central quantitative claim rests on one model's poor handling of a possibly out-of-distribution input format rather than on a robust property of FRES.","agreement_with_reader":"disagree"},"referee_report":null,"author_rebuttal":null,"desk_editor":null,"rs_alignment":null,"lean_confirmation":null,"pith_extraction":null,"created_at":"2026-08-07T15:39:33.825989+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":null,"supporting_citations":[],"review_version":1}