{"id":"fb4ec173-87e6-4787-997c-4b19af02f421","arxiv_id":"2607.10400","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":5,"one_line_summary":"A controlled synthetic long-document VQA benchmark surfaces three VLM failure modes—length degradation, middle-position sensitivity, and long-context chart collapse—that real-document benchmarks cannot isolate.","lead":"SynthDocBench is a fully synthetic, factorially controlled benchmark for long multi-page visual document understanding. It isolates length, layout, modality, and question type so failures of frontier VLMs can be attributed rather than confounded.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified that would overturn the central diagnostic claims.","rationale":"The paper's central contribution is a controlled diagnostic instrument, not a claim that absolute ACC numbers are definitive. The three concurrent failure modes are supported by factorial stratification (difficulty, position, modality), multi-model consistency, OCR/vision asymmetry, and cross-judge rank stability. The Gemini rendering-familiarity confound and Claude's leniency affect score levels and one model's ranking margin, but they do not erase the positional U-shape/decline, the L5 drop for non-Gemini models, or the chart-in-document collapse relative to saturated isolated-chart benchmarks. Code, data, and generation pipeline are released, enabling independent re-render and re-judge checks. Residual abstract/body count mismatches (e.g., \"five of six\" vs eight models evaluated) are presentation issues, not load-bearing. Therefore the reader's ACCEPT verdict and identification of the weakest assumption are both correct; no stronger objection lands that would move the verdict.","tokens_in":26184,"tokens_out":526,"duration_ms":6817,"concrete_test":"Re-score the full 1,788-question set for all models under Gemini-as-judge (already partially done) and recompute Table 4 position buckets + Table 3 L1–L5 ACC; if middle-third hardness and L5 drop remain for ≥4 of the top models and chart ACC still collapses relative to isolated ChartQA-style settings, the three failure modes are judge- and style-robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest assumption (rendering-style familiarity + single-judge absolute scores) is real but already partially stress-tested by the paper itself and does not underwrite the three failure modes as a package. Positional sensitivity (middle-third hardest for most models; Early\to Late declines) and L1\to L5 degradation appear across open-weight and proprietary families, not only Gemini; the OCR vs vision split isolates chart-reading as a genuine visual bottleneck independent of HTML/D3.js style; and Gemini-as-judge recovers GPT-5 rankings within ~3.5 ACC points (r≥0.94). Absolute score levels can shift under a lenient judge, but relative patterns and the qualitative existence of the three modes survive. The abstract's numerical phrasing of the Early\to Late trend is slightly looser than Table 4, yet the body evidence is consistent enough that the strongest claim holds.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces SynthDocBench, a fully synthetic, factorially controlled benchmark for long-context visual document understanding. Documents are generated end-to-end via an LLM pipeline over six layout archetypes (with a 40% random layout override), dual-layer D3.js charts plus hidden structured metadata for deterministic ground truth, and three QA families (chart-reading, cross-modal, complex multi-hop) at difficulty levels L1–L5. The released corpus comprises 200 multi-page reports (avg. 51.1 pages, ~20.6k words, 16.7 charts) and 1,788 questions. Evaluating eight frontier VLMs under a vision-only protocol with GPT-5 as primary judge (cross-validated against Gemini), the authors report three concurrent failure modes: degradation with reasoning depth (L1→L5), positional sensitivity (middle third often hardest; several models show Early→Late declines), and collapse of precise chart reading once charts sit in long multi-page contexts—supported by OCR baselines, bootstrap CIs, and rendering/prompting ablations.","tokens_in":26493,"tokens_out":1689,"duration_ms":27620,"significance":"If the diagnostic claims hold, the work fills a clear gap: existing chart and long-document VQA benchmarks either isolate charts or confound length, layout, modality, and question difficulty, so failures cannot be attributed. The dual-layer generation design, combinatorial axes, deterministic answers, public code/dataset, cross-judge checks (GPT-5 vs Gemini r≥0.94, ΔACC≤3.5), OCR vs vision split, and systematic ablations (pages-per-strip, DPI, prompting) are concrete methodological strengths that make the failure modes falsifiable and reusable. The positional and chart-in-context results are especially useful for the community, because they surface deficits that saturated single-page DocVQA/ChartQA scores and heterogeneous real-document suites do not isolate. The contribution is primarily empirical and diagnostic rather than algorithmic, but that is appropriate for a benchmark paper and is executed at a high standard.","major_comments":[{"comment":"Abstract vs body mismatch on the first failure mode and on reported statistics. The abstract states three modes including “sharp degradation with document length” and “middle third … hardest for five of six models … negative Early-to-Late trend (steepest decline: 8.3 percentage points),” and “seven frontier VLMs.” Contribution 3 and §5 instead emphasize degradation with evidence complexity/reasoning depth (L1→L5; Table 3), Table 2 evaluates eight models, and Table 4 reports middle-third hardness for 5 of 8 models with Early→Late Δ as large as −11.7 pp (Claude) and −16.0 pp (Qwen3.5-VL-122B). The Conclusion again uses different counts (“four of six” / “three of six”). Please align abstract, intro, and conclusion with the tables, and either (i) add a direct stratification of ACC by document page length (pages range 24–91 in Table 1) under controlled other factors, or (ii) rephrase the firs","section":"Abstract; §1; §5 Table 3–4; Conclusion"},{"comment":"The claim that length is varied as an independent diagnostic axis is only weakly realized in the reported results. §3 and the contributions assert combinatorial control of document length, yet the main empirical sections do not report ACC vs page count (or word count) while holding modality and question type fixed; the closest analyses are positional thirds within documents (Table 4) and pages-per-strip presentation ablations (Table 5), which measure context packing, not document length. Without a length-stratified result, the abstract’s “degradation with document length” and the “independently variable length” framing overclaim relative to the evidence. A short length-bucket table (or explicit statement that length is controlled for generation diversity but not the primary reported axis) is needed for the central diagnostic narrative to be load-bearing as written.","section":"§3.1–3.3; §5; Table 1 vs Tables 2–4"},{"comment":"Rendering-familiarity and judge absolute-score sensitivity remain residual threats to interpreting absolute ACC levels, especially Gemini’s large lead. §5 already flags HTML/D3.js familiarity as a possible confound for the Gemini–Qwen gap and shows Claude-as-judge is systematically more lenient (+11–16 ACC overall; Appendix C). Relative rankings and the qualitative existence of positional / L1→L5 / chart-in-context modes are reasonably robust (Gemini-as-judge recovers rankings; OCR isolates visual chart bottleneck). Still, for the stronger claim that models “may be overfitting to benchmark artifacts,” the paper should either (a) report at least one alternative rendering backend or non-web chart style on a subset, or (b) clearly scope the claim to relative patterns under this rendering distribution rather than absolute competence. This is fixable within the manuscript’s scope but currentl","section":"§5 (Gemini dominance paragraph); Appendix C Tables 8–9"}],"minor_comments":[{"comment":"Figure 1 caption and landscape discussion refer to “Top/Bottom” layout that is hard to parse in the text-only rendering; ensure the published figure clearly separates the scatter of length vs tokens from the tabular comparison.","section":"Figure 1"},{"comment":"Table 2 header says “Prior Benchmarks” with DocVQA and MMLongBench-Doc, but several models lack published MMLongBench-Doc numbers (—); a footnote on score provenance and access date (already partly in footnote 1) would help.","section":"Table 2"},{"comment":"Hard-failure analysis text oscillates between “all six models” and “all eight models” (Figure 6 caption vs §5 and Appendix M Table 21). Unify the model set used for the hard-failure definition.","section":"§5; Figure 6; Appendix M"},{"comment":"Notation: difficulty is ℓ in the glossary (Table 15) but L1–L5 in the main text; pick one convention. Also clarify whether τ=6 is inherited solely from MMLongBench-Doc or re-validated on this corpus (Appendix N.4 helps but is buried).","section":"§4 Eq. (1); Appendix L.1; Appendix N.4"},{"comment":"Ethics/reproducibility statements are thorough; consider adding approximate API cost or token volume alongside the stated 150 GPU-hours / 200 API-hours for completeness.","section":"Ethics Statement; Reproducibility Statement"},{"comment":"Minor copy-edits: “adual-layer” spacing in §3.1; “SYNTHDOCBENCH” capitalization is inconsistent with “SynthDocBench” in the abstract and URLs; “five of six” leftovers after model-set expansion.","section":"Throughout"}],"recommendation":"minor_revision","confidential_remarks":"Solid benchmark paper for COLM-level venues. The abstract/body numerical drift and the length-vs-difficulty framing are the main things that would bother a careful reader; they look like revision residue after the model set grew from six to eight and after the failure-mode wording was refined. I would not reject over them. Ecological-validity objections to synthetic data are expected but the paper already positions the work correctly as diagnostic. No integrity or scope concerns beyond ordinary benchmark Goodhart risk, which the authors discuss."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a useful instrument paper, not a theory paper. What is new is the factorial synthetic design for multi-page visual documents: length, six layout archetypes (with 40% random override), 24 chart types, dual-layer D3.js + hidden metadata for deterministic ground truth, and three question families (chart, cross-modal, complex L1–L5). That combination is not in DocVQA, ChartQA, or MMLongBench-Doc, and it lets them isolate failures that real corpora confound.\n\nThey do the empirical work carefully. 200 reports, 1,788 questions, eight models, bootstrap CIs, OCR vs vision split (complex is text-recoverable; chart-reading is not), pages-per-strip / DPI / prompt ablations, and cross-judge check (GPT-5 vs Gemini r≥0.94, ΔACC ≤3.5). Rankings are stable. The three modes hold up in the body: L1→L5 degradation for most models, middle-third hardness for most models with several Early→Late declines (Claude −11.7 pp on chart-reading), and chart collapse once charts sit inside long documents. The OCR baseline and multi-family pattern make the “just Gemini likes HTML” story incomplete.\n\nSoft spots are real but secondary. The abstract’s Early→Late numbers are looser than Table 4 (and model counts drift between abstract and body). Absolute judge scores move under Claude’s leniency. Gemini’s lead may partly reflect rendering familiarity—the authors say so in §5. Free parameters (τ=6, concat-num=5, DPI=144, L1–L5 taxonomy) are design choices, not hidden fitting. None of that overturns the qualitative claims or the value of the released code and HF dataset.\n\nWho it is for: people building or evaluating long-context VLMs and document AI who need attribution, not another saturated leaderboard. I would bring it to reading group, cite the benchmark and the positional/chart findings, and send it to peer review. Minor cleanup on abstract consistency and a second rendering backend would strengthen it; the core contribution is already solid.","headline":"Solid controlled diagnostic for long-context visual docs: three real failure modes, released code/data, and confounds the authors already flag.","tokens_in":27115,"tokens_out":523,"would_cite":true,"duration_ms":6616,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Frontier vision-language models fail on long multi-page documents in three controlled ways existing benchmarks hide.","keywords":["vision-language models","long-context document understanding","synthetic benchmark","chart reading","cross-modal reasoning","positional bias","visual document VQA"],"falsifier":"Re-render the same documents and questions with a non-web chart backend (e.g., vector or scientific plotting) and re-score with multiple independent judges; if the three failure modes shrink or reverse while real-document rankings stay stable, the central diagnostic claim weakens.","tokens_in":27115,"feed_emoji":"📄","tokens_out":872,"duration_ms":10304,"temperature":0.7,"pith_summary":"Real document benchmarks mix length, layout, charts, and question difficulty so tightly that a wrong answer rarely tells you which factor broke. SynthDocBench is a fully synthetic long-context suite that varies those factors independently across 200 multi-page reports and 1,788 questions spanning chart reading, cross-modal grounding, and multi-hop reasoning. When seven frontier models are evaluated under a vision-only protocol, three failures appear together: accuracy falls sharply as evidence complexity and reasoning depth rise from L1 to L5; the middle third of a document is hardest for most models and most models lose accuracy from early to late evidence; and precise chart-value reading collapses once charts sit inside long documents rather than as isolated images. The paper argues these patterns show models may be overfitting to the artifacts of existing short or uncontrolled benchmarks rather than achieving genuine long-context visual document understanding.","feed_headline":"Long-document VLMs fail three ways short tests hide","feed_subtitle":"Middle pages, deep multi-hop questions, and embedded charts all break frontier models under controlled conditions.","key_machinery":"SynthDocBench: an end-to-end LLM generation pipeline that produces dual-layer documents (rendered D3.js charts plus hidden structured metadata for deterministic ground truth) across six layout archetypes with a 40% random override, yielding three controlled question families (chart, cross-modal, complex multi-hop) at difficulty levels L1–L5.","core_discovery":"By holding document factors independent in a synthetic long-context suite, the authors show that frontier VLMs exhibit three concurrent, previously unobservable failure modes: sharp degradation with evidence complexity and reasoning depth, systematic positional sensitivity that makes the middle of a document hardest and produces a negative early-to-late trend for most models, and collapse of precise chart-reading accuracy once charts are embedded in multi-page contexts.","pith_inferences":["Training curricula that never place dense charts deep inside long page sequences may leave a permanent blind spot that isolated chart benchmarks cannot detect.","The middle-third hardness pattern is consistent with known long-context retrieval biases and may be mitigated by explicit positional encoding or hierarchical document memory.","If rendering-style familiarity proves real, community benchmarks will need multi-backend rendering to prevent silent distribution overfitting.","The OCR-versus-vision asymmetry implies hybrid pipelines remain necessary until pure VLMs close the precise chart-value gap inside long documents."],"forward_implications":["Progress claims based only on DocVQA, ChartQA, or uncontrolled multi-page suites will systematically overstate long-context visual robustness.","Model development must separately target middle-of-document retrieval, pixel-level chart decoding under long context, and multi-hop cross-modal alignment.","Future long-document benchmarks can isolate single axes of difficulty instead of confounding them, making failure attribution routine.","Vision-only evaluation with deterministic synthetic ground truth becomes a practical alternative to costly real-document annotation for diagnosis."],"fun_headline_variants":["VLMs collapse on length, mid-page bias, and charts in long docs","SynthDocBench reveals three VLM failures short tests miss","Middle pages hardest as VLMs show length and chart breakdown","Controlled synth docs expose VLM positional sensitivity and collapse","Frontier VLMs fail length, mid-document, and embedded chart tasks"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That the measured gaps mainly reflect true long-context visual reasoning limits rather than models’ familiarity with the HTML/D3.js rendering style or quirks of the primary automatic judge.","fun_headline_variants_meta":{"raw":{"variants":["VLMs collapse on length, mid-page bias, and charts in long docs","SynthDocBench reveals three VLM failures short tests miss","Middle pages hardest as VLMs show length and chart breakdown","Controlled synth docs expose VLM positional sensitivity and collapse","Frontier VLMs fail length, mid-document, and embedded chart tasks"]},"model":"grok-4.5","effort":"low","cost_usd":0.00538,"raw_usage":{"total_tokens":1500,"prompt_tokens":809,"num_sources_used":0,"completion_tokens":71,"cost_in_usd_ticks":53800000,"prompt_tokens_details":{"text_tokens":809,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":620,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":809,"tokens_out":71,"duration_ms":7048,"temperature":1.0,"reasoning_tokens":620,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T12:01:07.865813+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-render the same documents and questions with a non-web chart backend (e.g., vector or scientific plotting) and re-score with multiple independent judges; if the three failure modes shrink or reverse while real-document rankings stay stable, the central diagnostic claim weakens.","supporting_citations":[],"review_version":1}