REVIEW 2 major objections 1 minor 4 references
CFMME benchmark shows top LVLMs reach only 66.11% accuracy on Chinese financial question answering.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Introduces CFMME benchmark and reports that top LVLMs reach 66.11% accuracy on financial QA and 77.18 average on detection, recognition, and extraction tasks.
T0 review reviewed 2026-06-29 challenge →
load-bearing objection CFMME is a new Chinese financial multimodal benchmark with clear evaluation numbers, but its claim to cover the full workflow rests on an unverified coverage assumption. the 2 major comments →
Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
CFMME comprises 6052 instances spanning eight primary financial image modalities and four core multimodal tasks that together cover the perception-to-cognition demands of the Chinese financial business workflow. Thorough evaluation of representative LVLMs on this benchmark shows the state-of-the-art model attaining 66.11 percent overall accuracy on the question-answering task and an average score of 77.18 on the detection, recognition, and information-extraction tasks, with accompanying analyses of error causes, cross-modal capabilities, and multi-orientation settings.
What carries the argument
CFMME, the Chinese Financial Multimodal Evaluation dataset that organizes 6052 instances across eight image modalities and four tasks to measure multimodal financial capabilities.
Load-bearing premise
The 6052 instances and eight image modalities are assumed to comprehensively represent the perception, understanding, reasoning, and cognition demands across the entire financial business workflow in Chinese contexts.
What would settle it
A model that scores above 95 percent on every CFMME task while also matching or exceeding its performance on established non-financial multimodal benchmarks, or the identification of major financial scenarios absent from the eight modalities.
If this is right
- Current LVLMs require targeted improvements to reach reliable performance on financial multimodal reasoning and extraction.
- Error analyses supplied with the benchmark identify concrete failure modes that future training regimes can address.
- The dataset enables consistent tracking of progress on cross-modal financial capabilities over successive model releases.
- Multi-orientation evaluation settings in CFMME can be reused to test robustness under varied prompt and input conditions.
Where Pith is reading between the lines
- A parallel benchmark for non-Chinese financial documents would reveal whether the observed performance gap is language-specific or domain-specific.
- If models trained on CFMME also improve on general financial text-only tasks, the dataset may be capturing transferable financial reasoning structures.
- Deployment of improved models on real financial workflows could be measured by direct correlation between CFMME scores and downstream accuracy on live document streams.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces CFMME, a Chinese financial multimodal evaluation benchmark consisting of 6,052 instances across eight image modalities and four tasks (question answering, detection, recognition, and information extraction). It evaluates representative LVLMs and reports that the state-of-the-art model achieves 66.11% accuracy on QA and an average score of 77.18 on the other tasks, concluding that there is substantial room for improvement in current LVLMs. Additional analyses on error causes, cross-modal capabilities, and multi-orientation settings are provided.
Significance. If the dataset construction and coverage claims hold, CFMME would be a valuable addition as a domain-specific multimodal benchmark in Chinese financial contexts, providing concrete performance baselines and insights that could guide LVLM development. The multi-task, multi-modal design and focus on real-world financial applications are strengths.
major comments (2)
- [Abstract and §3] Abstract and §3 (Dataset): The central claim that the 6,052 instances comprehensively represent perception, understanding, reasoning, and cognition demands across the entire financial business workflow rests on an unverified coverage assumption. No explicit mapping from instances to a defined workflow ontology, sampling frame from real documents, or expert validation metric is described, which is load-bearing for interpreting the 66.11% and 77.18 scores as evidence of 'substantial room for improvement' in the domain.
- [§4 and §5] §4 (Annotation) and §5 (Experiments): No details are provided on the annotation process, inter-annotator agreement, or statistical significance testing for the reported accuracy and average scores. This prevents verification of the reliability of the headline results and the claim that current LVLMs have substantial room for improvement.
minor comments (1)
- [Figures and Tables] Figure captions and tables could more explicitly link results to specific modalities or tasks for easier cross-referencing.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback. We address each major comment below, providing clarifications based on the manuscript and committing to revisions where the points identify gaps in documentation.
read point-by-point responses
-
Referee: [Abstract and §3] Abstract and §3 (Dataset): The central claim that the 6,052 instances comprehensively represent perception, understanding, reasoning, and cognition demands across the entire financial business workflow rests on an unverified coverage assumption. No explicit mapping from instances to a defined workflow ontology, sampling frame from real documents, or expert validation metric is described, which is load-bearing for interpreting the 66.11% and 77.18 scores as evidence of 'substantial room for improvement' in the domain.
Authors: Section 3 describes the construction of CFMME by selecting instances from real Chinese financial documents across eight modalities and four tasks, spanning academic knowledge to applied scenarios. The claim of comprehensive coverage is grounded in this sampling from representative financial contexts rather than a formal ontology. However, we agree that an explicit mapping table linking tasks to workflow stages and details on the expert-guided sampling frame would strengthen verifiability. We will add this as a new subsection or appendix in the revision. revision: yes
-
Referee: [§4 and §5] §4 (Annotation) and §5 (Experiments): No details are provided on the annotation process, inter-annotator agreement, or statistical significance testing for the reported accuracy and average scores. This prevents verification of the reliability of the headline results and the claim that current LVLMs have substantial room for improvement.
Authors: We will revise §4 to detail the annotation pipeline, including annotator qualifications, guidelines, and inter-annotator agreement metrics such as Cohen's kappa. In §5, we will include statistical significance testing (e.g., bootstrap confidence intervals or McNemar's test) for the key accuracy figures to support the interpretation of results. These additions will be made in the revised manuscript. revision: yes
Circularity Check
No significant circularity; empirical benchmark with direct evaluation results
full rationale
The paper introduces a new dataset (CFMME) and reports direct empirical accuracy scores from running existing LVLMs on its tasks. No derivations, fitted parameters renamed as predictions, or load-bearing self-citations appear in the evaluation chain. The coverage claim for the 6052 instances is an explicit modeling assumption rather than a derived result that reduces to its own inputs. All reported numbers (66.11% QA accuracy, 77.18 average on other tasks) are measured outputs, not constructed by definition from the benchmark itself.
Axiom & Free-Parameter Ledger
Cite this review
Pith. "Pith review of Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset." pith.science (2026). https://pith.science/paper/PXCU7MIB
@misc{pith2026260529462,
author = {Pith},
title = {Pith review of: Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset},
year = {2026},
howpublished = {\url{https://pith.science/paper/PXCU7MIB}},
note = {Machine review of arXiv:2605.29462}
}
read the original abstract
The emergence of Large Vision-Language Models (LVLMs) has substantially expanded model capabilities beyond text-only understanding, enabling unified inference across both visual and textual modalities and supporting a broader range of real-world applications. To comprehensively evaluate the perception, understanding, reasoning, and cognition capabilities of LVLMs throughout the entire financial business workflow in Chinese contexts, we introduce CFMME, a novel Chinese financial multimodal evaluation benchmark. CFMME comprises 6,052 instances spanning from fundamental academic knowledge to complex real-world applications, covering eight primary financial image modalities and four core multimodal tasks. On CFMME, we conduct a thorough evaluation of representative LVLMs. The results show that the state-of-the-art model attains an overall accuracy of 66.11\% on the question answering task and an average score of 77.18 on the detection, recognition, and information extraction tasks, indicating substantial room for improvement in current LVLMs. In addition, we conduct detailed analyses of error causes, cross-modal capabilities, and multi-orientation settings, yielding valuable insights for future research. We hope that CFMME will spur further progress in LVLMs, especially by improving their performance on multiple multimodal tasks in the financial domain.
Figures
Reference graph
Works this paper leans on
-
[1]
arXiv preprint arXiv:2401.14011
Cmmu: A benchmark for chinese multi-modal multi-type question understanding and reasoning. arXiv preprint arXiv:2401.14011. Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Stein- hardt. 2021. Measuring massive multitask language understanding.Proceedings of the International Con- ference on Learning Representatio...
-
[2]
When FLUE meets FLANG: Benchmarks and large pretrained language model for financial domain,
When flue meets flang: Benchmarks and large pre-trained language model for financial do- main.arXiv preprint arXiv:2211.00083. Baidu ERNIE Team. 2025. Ernie 4.5 technical report. Kimi Team, Angang Du, Bohong Yin, Bowei Xing, Bowen Qu, Bowen Wang, Cheng Chen, Chenlin Zhang, Chenzhuang Du, Chu Wei, and 1 others. 2025a. Kimi-vl technical report.arXiv preprin...
-
[3]
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Cdm: A reliable metric for fair and accurate formula recognition evaluation.arXiv e-prints, pages arXiv–2409. Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, Long Cui, Xingguang Wei, Zhaoyang Liu, Linglin Jing, Shenglong Ye, Jie Shao, and 1 others. 2025. In- ternvl3. 5: Advancing open-source multimodal mod- els in versatility, reasoning, and efficiency.a...
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[4]
bbox_2d": [x_min, y_min, x_max, y_max],
IEEE. Xu Zhong, Elaheh ShafieiBavani, and Antonio Ji- meno Yepes. 2020. Image-based table recognition: data, model, and evaluation. InEuropean conference on computer vision, pages 564–580. Springer. Jie Zhu, Junhui Li, Yalong Wen, and Lifan Guo. 2024. Benchmarking large language models on cflue–a chinese financial language understanding evaluation dataset...
2020
This paper was first reviewed by grok-4.3 on June 29, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.