{"id":"85d2f106-e3e9-404e-a53a-c24ce94871a7","arxiv_id":"2505.00063","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A document-intelligence benchmark decouples visual and reasoning complexity, and a parameter-freezing fine-tuning method improves an 8B model without catastrophic forgetting.","lead":"This paper introduces GDI-Bench, a benchmark of 2,300 document images and about 3,660 test cases that separates visual difficulty from reasoning difficulty. It also introduces LW-AFT, a fine-tuning method that freezes most model parameters to avoid forgetting while improving document reasoning, and a resulting 8B model that beats several larger systems on the new benchmark.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The V/R decoupling is not validated: V labels are thresholded model performance and Table 5 contradicts the R difficulty ordering (every model scores higher on R2V0 than R1V0), so the 3x3 grid may not measure independent visual and reasoning difficulty.","rationale":"The reader's weakest assumption identifies the circularity of V labels, which I share. I also find a second, independent validity threat to the decoupling claim: the R dimension is not an ordered difficulty scale in the reported results, and the R1/R2 comparison is confounded with answer format and scoring metric. Table 4 further weakens the LW-AFT generalization claim by contradicting the text, but that is secondary; the benchmark's diagnostic validity is the load-bearing issue. These problems are addressable with a focused validation study, so the reader's CONDITIONAL verdict is appropriate. Acceptance should require the proposed matched-metric and structural-label re-analysis rather than treating the 3x3 grid as self-evidently decoupled.","tokens_in":16763,"tokens_out":9967,"duration_ms":104406,"concrete_test":"Construct a matched factorial validation subset from existing GDI-Bench images: same page/category at each V level, with R1 and R2 items in identical answer format (all multiple-choice, or all NED-scored), and re-score the six models. Independently re-assign V labels from layout-structural features (column count, figure area, region density) instead of OmniDocBench model thresholds and re-run the analysis. The decoupling claim survives only if the R1<R2 ordering is stable under matched metrics and the V gradient and LW-AFT weakness localization persist under structural labels.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that GDI-Bench decouples visual and reasoning complexity and thereby localizes weaknesses requires both axes to be valid, independent difficulty scales. That condition is not established. (1) Section 3.2.2 assigns V1/V2 by thresholding OmniDocBench end-to-end edit-distance scores (0.142) achieved by current SOTA models and pipeline tools, rather than by document structure. This re-encodes the performance of the same model families the benchmark is meant to diagnose; the threshold also disagrees with Fig. 3, where textbook (0.102) is below the threshold yet is called V2 in the text. (2) The R axis is not monotonic in Table 5: at V0, every evaluated model scores substantially higher on R2 than on R1 (e.g., InternVL3-8B: R2V0=0.89 vs R1V0=0.35; Qwen2.5-VL-72B: 0.90 vs 0.60). Because R1 extraction is scored by 1-NED while R2 is mostly single-choice, the apparent R1<R2 ordering may be a metric artifact rather than a reasoning-difficulty ordering. If either axis is invalid, the conclusion that a 1%-parameter sparse update fixes the R1/R2 weakness is not supported. Appendix A.5 lists only single-image scope as a limitation and does not address this validity threat.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces GDI-Bench, a document intelligence benchmark containing 3,660 test cases over 2.3k images across 9 scenarios and 19 tasks, with a claimed decoupling of visual complexity (V0-V2) and reasoning complexity (R0-R2). It also proposes LW-AFT (Layer-wise Adaptive Freeze-Tuning), which identifies a sparse subset of parameters for updating during SFT, and a GDI-Model fine-tuned from InternVL3-8B. The experiments report that GDI-Model improves on the base model and on LoRA fine-tuning for GDI-Bench while roughly preserving performance on existing benchmarks (Table 3), and the authors claim that GDI-Bench's difficulty decoupling enables weakness localization. The paper's central claims are that the benchmark's two difficulty axes are valid and independent, that the R1/R2 weakness of the base model is precisely identified, and that updating about 1% of parameters suffices to fix it without catastrophic forgetting. The submitted manuscript does not fully establish the validity of the difficulty axes and contains internal inconsistencies in the cross-domain/cross-task experiments.","tokens_in":17082,"tokens_out":7702,"duration_ms":71993,"significance":"If the main claims hold, GDI-Bench would be a useful diagnostic tool for document MLLMs, and LW-AFT would offer a simple, parameter-efficient way to improve targeted weaknesses while avoiding forgetting. The paper has several strengths: it covers a broad set of realistic document tasks, includes a human-verification pipeline for the annotations, evaluates multiple open and closed models, and reports results on standard benchmarks showing that LW-AFT largely preserves the base model's accuracy. The proposed method is simple and reproducible in principle, and the authors state an intention to open-source the benchmark and models. However, the validity of the V/R decoupling is not established, the R-axis is non-monotonic in the reported results, the V-axis is built partly from model-performance thresholds, and the cross-domain/cross-task experiments contradict the accompanying text. These issues are load-bearing because the paper's contribution is precisely the diagnostic and localization value of the benchmark and the claimed effectiveness of LW-AFT for cross-domain generalization.","major_comments":[{"comment":"The visual-complexity taxonomy is inconsistent with the stated threshold and is partially circular. Section 3.2.2 says that domains with end-to-end edit-distance scores above 0.142 are categorized as V2, but Fig. 3 reports textbook at 0.102, below the threshold, while §3.1.1 explicitly designates textbook as V2. Furthermore, the 0.142 threshold is derived from the performance of current SOTA models and pipeline tools on OmniDocBench, and those same model families are then evaluated on GDI-Bench; this means the V-axis may re-encode the performance of the models the benchmark is meant to diagnose. Please replace the model-performance threshold with structural, content-based, or human-annotated visual-complexity criteria (or provide a convincing argument that model performance is a stable proxy for intrinsic visual complexity), and resolve the textbook/threshold inconsistency.","section":"§3.1.1, §3.2.2, Fig. 3"},{"comment":"The reasoning-complexity ordering is not supported by the reported data. In Table 5, every evaluated model scores higher on R2V0 than on R1V0 (e.g., InternVL3-8B: R2V0=0.89 vs. R1V0=0.35; Qwen2.5-VL-72B: 0.90 vs. 0.60). Because R1 is scored with 1−NED on free-form extraction while R2 is mostly single-choice exact match, the apparent R1<R2 ordering may be a metric artifact rather than a true ordering of reasoning difficulty. Without an independent validation of the R-axis (for example, human difficulty ratings or a task design that equalizes the output format across R1 and R2), the central claim that R0, R1, and R2 measure increasing reasoning complexity is not established, and the conclusion that the base model has a specific reasoning weakness in R1 is not grounded.","section":"§3.1.2, Table 5"},{"comment":"The 'theoretical proposition' in Eq. (3) is stated as a formal claim but is not proven; it is essentially an adaptation of the Lottery Ticket Hypothesis to MLLMs with an unspecified domain transformation ψ. As written, the paper asks the reader to accept the existence of a sparse task-salient subnetwork and a transformation ψ that makes the subnetwork comparable to the full model, but ψ is never defined, instantiated, or empirically tested. Please re-label this as a motivating hypothesis and provide supporting evidence (e.g., verify for multiple models and tasks that updating only the selected top-parameter subset matches full-model updates), or remove the theorem-like presentation.","section":"§4, Eq. (3)"},{"comment":"The results in Table 4 contradict the accompanying text. The text states that LW-AFT 'demonstrates strong cross-domain and cross-task capabilities, significantly outperforming the LoRA fine-tuning,' but Table 4 shows LoRA outperforming LW-AFT on all four transfer tasks (e.g., T2: 0.473 vs. 0.365; T4 date: 0.093 vs. 0.010), and LW-AFT is also worse than the base model on T2 and T3. This discrepancy directly undermines the generalization claim for LW-AFT. Please report the correct interpretation of Table 4, and either supply additional evidence for cross-domain/cross-task transfer or revise the conclusion to reflect the actual results.","section":"§5.1.3, Table 4"},{"comment":"The R1/R2 question-answer pairs are generated by GPT-4o, and GPT-4o is then evaluated on the same benchmark (Table 5). The authors filter out cases answerable without the image using DeepSeek-R1 and state that PhD-level annotators review the data, but the paper does not report the number of annotators, the fraction of cases revised, or any inter-annotator agreement. If the human review did not independently verify all ground-truth answers, the benchmark may inherit GPT-4o's annotation errors, which would inflate or distort model comparisons. Please add annotation quality statistics and a contamination analysis for the models that were used in the annotation loop.","section":"§3.2.2, §5.2"}],"minor_comments":[{"comment":"The legend includes a 'random' baseline, but the text never explains how this baseline is computed. Please clarify what the random score is (e.g., 25% chance for four-option questions) and how it is applied to the R0 and R1 non-choice tasks.","section":"Fig. 2"},{"comment":"The caption calls these 'visual complexity scores,' but the values are actually end-to-end edit distances from OmniDocBench. Please label the axis accordingly and mark the 0.142 threshold used in Section 3.2.2 so the reader can see which domains fall above and below it.","section":"Fig. 3"},{"comment":"The freeze-rate ablation reports GDI-Bench score and several other benchmarks, but no confidence intervals or significance tests are provided. The differences between adjoining freeze rates are small, so please add error bars or statistical tests to support the claim that 99% freezing is optimal.","section":"§5.1.1, Fig. 9"},{"comment":"Please specify the number of annotators and the inter-annotator agreement for the human-correct step. This is important for establishing the reliability of the benchmark's ground truth, especially since part of it is generated by GPT-4o.","section":"§3.2.2"},{"comment":"The metric P(·) in Eq. (3) is never defined, and the domain transformation ψ is not used again anywhere in the paper. Please define both symbols or remove them if they are not operational.","section":"§4, Eq. (3)"},{"comment":"Eq. (5) allocates h_ℓ unfrozen parameters per layer, but the text does not state how h_ℓ is rounded to an integer or how ties are handled in the TopK selection. Please clarify these implementation details.","section":"Eqs. (5)-(6)"},{"comment":"There are several typos and stylistic issues, including the run-on sentence in Section 4 (\"the model performs well ... To address the model's weaknesses ... constructed supervised fine-tuning\"), the inconsistent use of \"GDIBench\" vs. \"GDI-Bench\", and the hyphenation of \"intelligence-preserving.\" I recommend a careful proofreading pass.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper has a genuinely useful benchmark contribution and a simple, interesting training strategy, but the validity of the decoupling axes is not established, and the cross-domain/cross-task results in Table 4 contradict the text. The discrepancy between the text and Table 4 is serious enough that I would ask the editor to have the authors confirm the data and protocol, and to require the code and benchmark data be released before acceptance. I also recommend that the authors be asked to address the R-axis monotonicity issue directly, as it affects the core diagnostic claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nGDI-Bench is a serious attempt to bring difficulty grading to document intelligence, and the LW-AFT fine-tuning recipe is a clean, cheap variation on sparse fine-tuning. The 3x3 grid over visual and reasoning complexity is genuinely new for this area—Table 2 shows no prior doc benchmark does it—and the evaluations cover a sensible range of open and closed models. The ablations on alpha and freeze rate are useful, and the result that freezing 99% of parameters keeps most general abilities while improving the target tasks is consistent with the broader sparse-tuning literature. I think the paper deserves a real referee, but it has some validity problems that need to be fixed before the central claims can be trusted.\n\nThe biggest issue is that the decoupling is not validated. The V1/V2 labels come from thresholding OmniDocBench edit-distance performance (0.142), so the \"visual complexity\" axis is partly defined by how well current models do, not by document structure. The threshold also disagrees with Fig. 3: textbook (0.102) should be V1 by that threshold, yet the text calls it V2. The R axis is worse: in Table 5, every model scores much higher on R2V0 than R1V0 (e.g., InternVL3-8B: 0.89 vs 0.35). Since R1 is scored by 1−NED and R2 is mostly single-choice, the apparent ordering is likely a metric artifact. If the reasoning axis is not monotonic, the benchmark's core diagnostic claim—that you can localize reasoning-level weaknesses—is not supported.\n\nSecond, the cross-domain results in Table 4 directly contradict the text. LW-AFT is the worst or near-worst method on all four cross-domain tasks (T1–T4), yet the text says it \"significantly outperforms LoRA fine-tuning.\" That is not a minor typo; it undermines the generalization claim that LW-AFT is supposed to provide.\n\nThere are also smaller issues: Eq. 3 is called a \"theoretical proposition\" but is asserted without proof—fine as a heuristic, but not a proposition. The \"SOTA on previous benchmarks\" claim is not substantiated by the comparisons shown, which are only against their own fine-tuned variants. And the benchmark/code are promised but not shipped, which is a real problem for a benchmark paper.\n\nNet: the benchmark idea and the LW-AFT recipe are worth engaging, and the paper is clearly written and reproducible in spirit. But as it stands, the validity threats and the Table 4 contradiction mean the main conclusions are not established. I'd send it to peer review, but I'd expect major revision: validate or re-label the V/R axes, fix the cross-domain reporting, and make the data/code available.\n\nBest.","headline":"A useful benchmark idea and a cheap fine-tuning trick, but the decoupling isn't validated and the cross-domain results contradict the paper's own claims.","tokens_in":17663,"tokens_out":5432,"would_cite":false,"duration_ms":50893,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A two-axis benchmark isolates whether document AI fails at seeing or at reasoning, and a 1%-parameter fix lifts weak areas without forgetting others.","keywords":["document intelligence","benchmark","multimodal large language models","complexity decoupling","visual complexity","reasoning complexity","catastrophic forgetting","parameter freezing"],"falsifier":"Recompute the V0/V1/V2 labels for the nine document categories using a substantially improved OCR/parsing pipeline and check whether the categories assigned as V2 change: if, say, textbook and exam-paper domains no longer have edit distances above the 0.142 threshold, the visual axis is tracking current model difficulty rather than stable visual complexity, undermining the claimed decoupling. A second check would be to verify with human annotators that V2 documents are genuinely harder to parse visually than V1 documents after controlling for content length and language.","tokens_in":16528,"feed_emoji":"📄","tokens_out":2934,"duration_ms":32274,"temperature":0.7,"pith_summary":"This paper introduces GDI-Bench, a benchmark that labels document tasks by visual complexity (V0–V2) and reasoning complexity (R0–R2), forming a 3x3 grid over 2.3k images and 19 tasks. The authors argue that this decoupling lets evaluators pinpoint whether a model's failure comes from perception or from higher-level reasoning, which ordinary single-score benchmarks cannot do. They also propose LW-AFT, a training method that freezes about 99% of a model's parameters and updates only a small domain-sensitive subset, reducing catastrophic forgetting during supervised fine-tuning. Using this method on the InternVL3-8B model, they produce a GDI-Model that keeps the base model's strong R0 performance while improving its R1 and R2 reasoning scores. If the claims hold, the benchmark offers a practical diagnostic grid for document intelligence and a cheap recipe for targeted model improvement.","feed_headline":"Two-axis document-AI benchmark isolates seeing from reasoning","feed_subtitle":"A 3x3 grid exposes why models fail, and updating just 1% of parameters fixes weak spots without forgetting old skills.","key_machinery":"The central object is the 3x3 difficulty grid of GDI-Bench, where each test case is labeled by visual complexity (V0/V1/V2) and reasoning complexity (R0/R1/R2). The visual labels are derived from a data-driven rule: domains whose OmniDocBench end-to-end edit-distance scores exceed 0.142 are classed as V2, and the rest as V1. The reasoning labels come from task design, with R0 being verbatim page extraction, R1 requiring selective information retrieval, and R2 requiring logical or multi-element inference. The training method LW-AFT is the carrying mechanism for the repair claim: it first fine-tunes a small expert model on a mini-set to measure per-parameter absolute changes, then computes per-layer average change magnitudes, allocates a global unfrozen-parameter budget H across layers proportionally to those magnitudes, and finally masks gradients so that only the top h_l parameters in each layer update. This parameter-freezing mask is what preserves R0 skills while allowing R1 and R2 improvements, directly connecting the benchmark's diagnostic output to a concrete optimization strategy.","core_discovery":"The paper's central claim is that decoupling document understanding into a visual-complexity axis and a reasoning-complexity axis exposes weaknesses that single-score benchmarks hide. Concretely, GDI-Bench assigns each task a V level (V0: plain text, V1: formal representations like tables and equations, V2: explanatory representations such as charts and complex layouts) and an R level (R0: full-page structured extraction, R1: information extraction, R2: reasoning). Evaluation on this grid shows, for example, that InternVL3-8B is strong at R0 but degrades sharply at R1 and R2. The paper further claims that this weakness localization is actionable: by analyzing parameter changes during full-parameter fine-tuning, they find over 95% of parameters barely move while a sparse 5% subset changes significantly, and they leverage this to freeze 99% of the model and update only the top 1% of sensitive parameters per layer (LW-AFT). The resulting GDI-Model maintains the base model's R0 accuracy, improves R1 and R2, and outperforms much larger models such as Qwen2.5-VL-72B at higher reasoning levels. Thus the paper simultaneously offers a diagnostic benchmark and a training method that turns diagnosis into targeted repair.","pith_inferences":["A natural extension is to build analogous V×R grids for other multimodal domains, such as medical imaging, UI screenshots, or video frames, where perception and reasoning failures are also confounded; the same decoupling logic should expose weakness patterns there.","The visual-complexity labels depend on current OCR pipeline performance, so the benchmark's V axis may drift as OCR improves; one testable consequence is that a substantially better OCR engine would re-classify some V2 domains as V1, which would change the reported weakness landscape.","The sparse-update observation that '95% of parameters barely move' suggests a broader hypothesis: for many SFT tasks, only a small task-salient subnetwork needs to be adjusted; if true, sensitivity-based masking could replace heavier continual-learning methods across a range of fine-tuning scenarios.","Because the training data for R1/R2 is explicitly sourced from domains disjoint from GDI-Bench, the benchmark could be used to measure how well a fine-tuned model transfers to unseen document types — a property the current experiments only partially probe."],"forward_implications":["GDI-Bench can serve as a diagnostic that maps a document model's failure mode onto one of six coordinates (V-level times R-level), guiding developers to either improve visual encoders or strengthen reasoning layers.","LW-AFT shows that updating only about 1% of parameters suffices to repair reasoning-level weaknesses without destroying existing extraction skills, offering a data-efficient and compute-light alternative to full fine-tuning.","The GDI-Model, at 8B parameters, matches or exceeds the reasoning performance of the 72B Qwen2.5-VL model on GDI-Bench, suggesting that targeted adaptation of a smaller base model can rival much larger general-purpose models on document-specific tasks.","The benchmark's task filtering pipeline — removing questions solvable without the image via DeepSeek-R1 and human review — gives a template for constructing vision-grounded QA data that tests genuine multimodal understanding."],"supporting_citations":[{"why":"OmniDocBench supplies the nine document domains and the edit-distance scores that define the V1/V2 visual complexity threshold, so the benchmark's V axis is built directly on it.","marker":"[19]"},{"why":"DeepSeek-R1 is used to filter out synthetic questions that can be answered without the document image, providing the no-visual-input check that grounds the vision dependence of GDI-Bench tasks.","marker":"[2]"},{"why":"InternVL3-8B is the base model for all fine-tuning experiments; its strong R0 and weaker R1/R2 profile is the demonstration case for GDI-Bench's weakness localization.","marker":"[53]"},{"why":"The Lottery Ticket Hypothesis motivates the paper's proposition that a sparse task-salient subnetwork can achieve comparable performance, which directly justifies the LW-AFT freezing strategy.","marker":"[52]"},{"why":"Elastic Weight Consolidation is cited as the canonical catastrophic-forgetting baseline and frames the problem that LW-AFT aims to solve.","marker":"[13]"},{"why":"LoRA is the parameter-efficient fine-tuning baseline that LW-AFT is compared against in the forgetting and cross-task experiments.","marker":"[45]"},{"why":"DocVQA is one of the held-out benchmarks used to measure whether LW-AFT preserves general document QA ability after fine-tuning.","marker":"[54]"},{"why":"ChartQA is another held-out benchmark used to test the generalization-versus-forgetting trade-off of the freezing method.","marker":"[16]"},{"why":"VisualSimpleQA is the prior decoupling benchmark that the paper contrasts with, grounding the novelty of GDI-Bench's complexity-decoupling approach.","marker":"[50]"}],"fun_headline_variants":["Document AI benchmark decouples vision and reasoning to expose gaps","Two-axis benchmark isolates seeing from reasoning in document AI","New benchmark finds document AI weak spots, then 1% tuning fixes them","GDI-Bench grades document tasks by visual and reasoning difficulty","Sparse 1% parameter update repairs document AI after diagnostic test"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's visual-complexity labels are assigned from current model and pipeline performance (edit-distance above 0.142 on OmniDocBench means V2), not from any intrinsic measure of document structure, so the V axis may already encode the strengths and weaknesses of the very models being evaluated.","fun_headline_variants_meta":{"raw":{"variants":["Document AI benchmark decouples vision and reasoning to expose gaps","Two-axis benchmark isolates seeing from reasoning in document AI","New benchmark finds document AI weak spots, then 1% tuning fixes them","GDI-Bench grades document tasks by visual and reasoning difficulty","Sparse 1% parameter update repairs document AI after diagnostic test"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000807,"raw_usage":{"total_tokens":3596,"prompt_tokens":1050,"completion_tokens":2546,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":666,"completion_tokens_details":{"reasoning_tokens":2458}},"tokens_in":666,"tokens_out":2546,"duration_ms":16855,"temperature":1.0,"reasoning_tokens":2458,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:55:00.477987+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the V0/V1/V2 labels for the nine document categories using a substantially improved OCR/parsing pipeline and check whether the categories assigned as V2 change: if, say, textbook and exam-paper domains no longer have edit distances above the 0.142 threshold, the visual axis is tracking current model difficulty rather than stable visual complexity, undermining the claimed decoupling. A second check would be to verify with human annotators that V2 documents are genuinely harder to parse visually than V1 documents after controlling for content length and language.","supporting_citations":[{"cited_title":"ChartQA: A benchmark for question answering about charts with visual and logical reasoning","cited_arxiv_id":null,"evidence_quote":"ChartQA is another held-out benchmark used to test the generalization-versus-forgetting trade-off of the freezing method."},{"cited_title":"Visualsimpleqa: A benchmark for decoupled evaluation of large vision-language models in fact-seeking question answering, 2025","cited_arxiv_id":null,"evidence_quote":"VisualSimpleQA is the prior decoupling benchmark that the paper contrasts with, grounding the novelty of GDI-Bench's complexity-decoupling approach."}],"review_version":1}