{"id":"9dca775c-0ff5-435c-81d2-8bbeb47440e4","arxiv_id":"2607.05614","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":1,"one_line_summary":"BaFCo introduces the first fine-grained Bangla form benchmark (26 entity types, relationships) and demonstrates that current MLLMs struggle especially with granular layout localization.","lead":"BaFCo is a carefully annotated benchmark of 200 multi-page Bangladeshi government forms for Document Layout Analysis and Key Information Extraction in Bangla. It shows that flagship multimodal LLMs still fail at precise localization of fine-grained form entities, limiting real-world document AI for a major low-resource language.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The strongest claim is a measured performance gap, not a universal generalization. The numbers in Tab. 4 (granular mAP 0.1177 vs. coarse 0.2646) and Tab. 6 (KIE F1 0.848) are the evidence; the English-control experiment (Tab. 5) further shows the DLA gap is largely language-agnostic while KIE is more language-sensitive. The reader's concern about representativeness is the natural soft spot for any modest-scale curated benchmark, yet the paper already frames BaFCo as a high-quality, purpose-built resource rather than a comprehensive sample of all Bangla forms, and the Limitations section explicitly flags scale and domain focus. Because the claim is scoped to the released set and the evaluation protocol is fully specified (prompts, IoU thresholds, validator), the representativeness issue does not undermine the contribution or the reported numbers. No stronger load-bearing flaw (e.g., metric misdefinition, non-reproducible matching, or contradictory English/Bangla results) appears. Therefore the ACCEPT verdict stands without adjustment.","tokens_in":22100,"tokens_out":567,"duration_ms":4835,"concrete_test":"Independently re-run the high-reasoning zero-shot DLA evaluation for Gemini 3 Pro on the released BaFCo granular split (or a 50-page random subset) using the exact prompt template in Appendix D and the same IoU@0.3 matching protocol; if mAP remains within ±0.02 of 0.1177 the central empirical claim is confirmed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is an empirical observation about current flagship MLLMs on a carefully constructed Bangla form benchmark: best granular DLA mAP@0.3 is only 0.1177 (Gemini 3 Pro, high-reasoning zero-shot), while KIE reaches F1 0.848. That claim is directly supported by the reported tables (Tab. 4–6), the dual granular/coarse taxonomy, the English-control comparison, and the public release. The reader's weakest assumption (representativeness of the 200 curated government forms) is real but already scoped and acknowledged in Sec. 3.1–3.2 and the Limitations section; it does not falsify the measured gaps on this set. Annotation reliability (Cohen's κ 0.974), multi-domain coverage (15 domains), and difficulty stratification further reduce the risk that the numbers are artifacts of a single layout style. No internal inconsistency or unstated assumption that would reverse the headline result is present.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces BaFCo, a curated benchmark of 200 multi-page Bangladeshi government forms (316 pages) for Document Layout Analysis (DLA) and Key Information Extraction (KIE) in Bangla. It defines a 26-class fine-grained form-entity taxonomy plus a 5-class coarse mapping, with 16,382 entities, 8,771 relationships, and 1,926 key–value pairs, annotated by trained annotators with expert review (Cohen’s κ = 0.974). Flagship MLLMs (GPT-5.2, Gemini 3 Pro, Claude Opus 4.6, Qwen 3.6-Plus, Kimi K2.5) are evaluated zero-shot and with CoT under low/high reasoning. The central empirical claim is that current MLLMs remain weak at localizing highly granular form entities (best granular mAP@0.3 = 0.1177 for Gemini 3 Pro) while KIE is substantially stronger (Bangla F1 up to 0.848), with language effects small for DLA and larger for KIE.","tokens_in":22331,"tokens_out":1252,"duration_ms":18223,"significance":"If the measured gaps hold, BaFCo fills a clear low-resource gap: no prior public Bangla form benchmark supports fine-grained entities, key–value linking, multi-page government layouts, and dual DLA/KIE evaluation (Table 1). The dual granular/coarse taxonomy, English-control comparison (Tables 5–6), difficulty stratification, multi-domain coverage (15 domains), and public release of data and code make the result actionable for both model developers and the low-resource document-AI community. Annotation reliability (κ = 0.974) and standard external metrics (IoU-based mAP, NES) strengthen credibility. The work is primarily empirical rather than theoretical, but the resource and the documented MLLM failure modes (spatial shift, miscategorization, hallucination) are of practical significance for ECCV-style document understanding and for real-world Bangla public-service automation.","major_comments":[{"comment":"Sec. 4.2, Eqs. (1)–(2): invalid model outputs are discarded by the schema validator V(·) before scoring, but exclusion rates (by model, prompt, reasoning level, and entity set) are not reported. Because free-form MLLM bbox generation often fails format constraints, unreported filtering can inflate or deflate mAP relative to true layout competence and makes cross-model comparison harder to interpret. Please report the fraction of pages/predictions rejected and, if non-negligible, provide a secondary score that treats schema failures as empty predictions.","section":null},{"comment":"Sec. 4.1 and Tables 4–6: the evaluation is restricted to off-the-shelf flagship MLLMs with no specialized document layout or KIE baseline (e.g., LayoutLMv3-style, DocLLM, or a strong OCR+rule pipeline) and no human upper bound. The authors justify the MLLM focus, but without at least one calibrated reference system it is difficult to separate “Bangla form hardness” from “MLLM bbox generation weakness.” A single specialized or OCR-based baseline on the same splits would substantially strengthen the claim that the observed gaps are task-level rather than interface-level.","section":null},{"comment":"Sec. 3.1–3.2 and Limitations: the 200 forms are carefully filtered public government forms with a quality-over-quantity design. The headline claim about “current MLLMs’ ability in comprehending Bangla forms” is therefore scoped to this curated distribution. The paper already acknowledges limited scale; still, the main text should more explicitly bound generalization (e.g., private-sector, handwritten-heavy, or non-portal forms) so that the measured mAP/F1 numbers are not over-read as language-wide.","section":null}],"minor_comments":[{"comment":"Table 5: absolute counts of English DLA forms/pages used in the Bangla–English comparison are not stated as clearly as the KIE ratios in Sec. 3.2; please add them for reproducibility.","section":null},{"comment":"Fig. 1 and Appendix figures: several qualitative panels are dense; larger crops or color-coded legends for predicted vs. ground-truth boxes would improve readability in print.","section":null},{"comment":"Sec. 3.2 difficulty rubric: “easy/medium/hard” criteria are described narratively; a short checklist table (presence of tables, checkboxes, nested cells, etc.) would make the stratification fully reproducible.","section":null},{"comment":"Throughout: minor consistency issues (e.g., “Gemini 3 Pro” vs “Gemini-3 Pro”, “Qwen 3.6 Plus” vs “Qwen 3.6-Plus”) should be normalized.","section":null},{"comment":"Related Work: BaDLAD is correctly positioned as coarse-layout only; a one-sentence note on whether any of its government pages overlap BaFCo would help readers assess independence.","section":null},{"comment":"Appendix D prompts: the full prompt templates are valuable; consider also releasing them as plain-text files in the public code repo for exact reproduction.","section":null}],"recommendation":"minor_revision","confidential_remarks":"Solid, well-scoped benchmark paper with unusually careful annotation protocol for a low-resource setting. The three major points are fixable without new data collection and do not undermine the central empirical claim. Fit for a document-understanding / multimodal venue is good; I would not block on the MLLM-only design if exclusion rates and one reference baseline are added. No integrity or novelty concerns."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a clean dataset paper that fills a real hole. BaFCo is the first public Bangla form resource with a 26-class entity taxonomy, explicit key–value and table relationships, multi-page coverage, and easy/medium/hard stratification. Prior Bangla work (BaDLAD) stopped at coarse layout; FUNSD/XFUND never covered Bangla. That is the actual novelty, and it is earned.\n\nWhat they did well is the annotation process. Seventeen trained annotators plus expert review, Cohen’s κ 0.974, explicit guidelines for the usual Bangla-form ambiguities (inline vs. form keys, guided tables, etc.), and a public release with code. The dual granular/coarse evaluation is useful: best granular DLA mAP@0.3 is only 0.1177 (Gemini 3 Pro, high-reasoning zero-shot) while coarse jumps to ~0.26 and KIE reaches F1 0.848. Language effects are small for DLA and larger for KIE, which is a clean empirical observation. Metrics follow Form-NLU / OmniDocBench conventions; no circular scoring.\n\nSoft spots are real but proportionate. Two hundred carefully filtered government forms is small, and the taxonomy is derived from that same set, so representativeness beyond public Bangladeshi admin forms is an open question—they acknowledge it. They deliberately skip OCR and specialized document models, which is a defensible scope choice for an MLLM-focused benchmark but leaves the absolute performance floor unmeasured. CoT and higher reasoning effort help little for localization; that is reported honestly rather than spun.\n\nWho it is for: anyone working on low-resource document AI, Bangla NLP, or form understanding who needs a hard, fine-grained test set. The numbers are reproducible enough to be useful baselines. I would send it to peer review without hesitation; it is the kind of careful resource paper that belongs at a major venue. Worth citing if you touch multilingual document understanding in the next year, and worth a reading-group slot if your group cares about low-resource multimodal work.","headline":"Solid first Bangla form benchmark with careful 26-class annotations and clear MLLM gaps; modest scale and MLLM-only scope are real but already scoped.","tokens_in":23003,"tokens_out":527,"would_cite":true,"duration_ms":6104,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Flagship multimodal models still fail at fine-grained layout of complex Bangla government forms, even when key-value extraction is strong.","keywords":["document layout analysis","key information extraction","Bangla","multimodal large language models","form understanding","low-resource NLP","government forms","benchmark dataset"],"falsifier":"A larger or differently sampled collection of Bangla forms (or a re-annotation under a coarser/finer taxonomy) on which the same flagship models achieve granular layout mAP comparable to their coarse or English results would falsify the claim that current MLLMs are systematically limited on fine-grained Bangla form layout.","tokens_in":23025,"feed_emoji":"📄","tokens_out":857,"duration_ms":7300,"temperature":0.7,"pith_summary":"This paper introduces BaFCo, the first public benchmark built for complex Bangla form understanding. It collects 200 multi-page Bangladeshi government forms from many administrative domains and annotates them with a 26-class layout taxonomy (plus a 5-class coarse grouping) and key-value relationships. The authors then test current flagship multimodal models with zero-shot and chain-of-thought prompts under low and high reasoning budgets. The central finding is that these models remain weak at precise geometric localization of fine-grained form entities, while they perform substantially better at extracting the textual values once the keys are known. A sympathetic reader cares because government forms are how citizens interact with public services, and without reliable layout and extraction tools the benefits of document AI stay out of reach for a major low-resource language.","feed_headline":"Flagship models still miss fine-grained Bangla form layouts","feed_subtitle":"New BaFCo benchmark shows strong key extraction but weak bounding-box localization on complex government forms","key_machinery":"BaFCo: a curated set of 200 multi-page Bangladeshi government forms annotated with 26 fine-grained layout entity types (and a mapped 5-type coarse set), entity relationships, and 1,926 key-value pairs, evaluated under controlled zero-shot / chain-of-thought and low / high-reasoning regimes.","core_discovery":"Current off-the-shelf multimodal large language models cannot yet accurately localize highly granular form entities on complex Bangla government forms (best reported granular mAP@0.3 is 0.1177), whereas key-information extraction on the same documents reaches much higher F1 scores (up to 0.848). Layout difficulty is driven more by spatial granularity and hierarchical structure than by language itself.","pith_inferences":["The same evaluation protocol could be applied to other low-resource South Asian scripts that share dense multi-column government-form conventions, testing whether the geometric bottleneck is script-general.","Because chain-of-thought and higher reasoning effort yield only mixed or negligible gains on layout, future gains are more likely to come from better visual-spatial grounding or layout-specific post-training than from longer textual reasoning.","The long-tailed entity distribution and hierarchical table structures in BaFCo make it a natural stress test for any future multimodal architecture that claims hierarchical region decomposition."],"forward_implications":["Document AI pipelines for Bangla public services must still treat fine-grained layout analysis as an unsolved bottleneck rather than a solved precursor to extraction.","Coarse 5-class layouts are markedly easier for current models than 26-class taxonomies, so practical systems may need hierarchical or progressive localization.","Language itself is not the dominant barrier for layout detection; the same models show similar DLA scores on matched English forms, so layout failure modes are primarily geometric.","Key-information extraction can already be useful on Bangla forms even when full layout understanding remains unreliable.","Open-source models close much of the KIE gap but still lag proprietary leaders on precise bounding-box localization."],"fun_headline_variants":["MLLMs still miss granular entity boxes on Bangla forms","BaFCo: strong KIE yet weak mAP on complex Bangla layouts","Flagship models fail fine-grained Bangla form localization","High entity F1 but low mAP@0.3 on Bangla government forms","Spatial granularity not language drives BaFCo layout failures"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The 200 carefully filtered public government forms and the 26-class taxonomy derived from them are assumed to be representative enough of real-world Bangla form diversity that the measured model gaps will generalize beyond this curated set.","fun_headline_variants_meta":{"raw":{"variants":["MLLMs still miss granular entity boxes on Bangla forms","BaFCo: strong KIE yet weak mAP on complex Bangla layouts","Flagship models fail fine-grained Bangla form localization","High entity F1 but low mAP@0.3 on Bangla government forms","Spatial granularity not language drives BaFCo layout failures"]},"model":"grok-4.5","effort":"low","cost_usd":0.00487,"raw_usage":{"total_tokens":1373,"prompt_tokens":794,"num_sources_used":0,"completion_tokens":75,"cost_in_usd_ticks":48700000,"prompt_tokens_details":{"text_tokens":794,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":504,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":794,"tokens_out":75,"duration_ms":4036,"temperature":1.0,"reasoning_tokens":504,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T04:55:54.085353+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A larger or differently sampled collection of Bangla forms (or a re-annotation under a coarser/finer taxonomy) on which the same flagship models achieve granular layout mAP comparable to their coarse or English results would falsify the claim that current MLLMs are systematically limited on fine-grained Bangla form layout.","supporting_citations":[],"review_version":1}