Pith. sign in

REVIEW 2 major objections 1 minor 4 references

CFMME benchmark shows top LVLMs reach only 66.11% accuracy on Chinese financial question answering.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Introduces CFMME benchmark and reports that top LVLMs reach 66.11% accuracy on financial QA and 77.18 average on detection, recognition, and extraction tasks.

T0 review reviewed 2026-06-29 challenge →

load-bearing objection CFMME is a new Chinese financial multimodal benchmark with clear evaluation numbers, but its claim to cover the full workflow rests on an unverified coverage assumption. the 2 major comments →

arxiv 2605.29462 v1 pith:PXCU7MIB submitted 2026-05-28 cs.CV cs.AI

Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset

classification cs.CV cs.AI
keywords Chinese financial multimodal benchmarkLVLM evaluationquestion answeringinformation extractionfinancial image modalitiesmultimodal tasksmodel performance analysiserror cause analysis
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces CFMME, a benchmark dataset designed to test large vision-language models on perception, understanding, reasoning, and cognition across the full Chinese financial workflow. It contains 6052 instances drawn from eight image modalities and four tasks that range from academic knowledge to real-world applications. When representative LVLMs are evaluated, the strongest model scores 66.11 percent overall accuracy on the question-answering task and 77.18 on average across detection, recognition, and information extraction. The results are accompanied by error-cause analyses and cross-modal capability breakdowns that point to specific weaknesses. The authors present CFMME as a tool to drive measurable progress in financial-domain multimodal models.

Core claim

CFMME comprises 6052 instances spanning eight primary financial image modalities and four core multimodal tasks that together cover the perception-to-cognition demands of the Chinese financial business workflow. Thorough evaluation of representative LVLMs on this benchmark shows the state-of-the-art model attaining 66.11 percent overall accuracy on the question-answering task and an average score of 77.18 on the detection, recognition, and information-extraction tasks, with accompanying analyses of error causes, cross-modal capabilities, and multi-orientation settings.

What carries the argument

CFMME, the Chinese Financial Multimodal Evaluation dataset that organizes 6052 instances across eight image modalities and four tasks to measure multimodal financial capabilities.

Load-bearing premise

The 6052 instances and eight image modalities are assumed to comprehensively represent the perception, understanding, reasoning, and cognition demands across the entire financial business workflow in Chinese contexts.

What would settle it

A model that scores above 95 percent on every CFMME task while also matching or exceeding its performance on established non-financial multimodal benchmarks, or the identification of major financial scenarios absent from the eight modalities.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Current LVLMs require targeted improvements to reach reliable performance on financial multimodal reasoning and extraction.
  • Error analyses supplied with the benchmark identify concrete failure modes that future training regimes can address.
  • The dataset enables consistent tracking of progress on cross-modal financial capabilities over successive model releases.
  • Multi-orientation evaluation settings in CFMME can be reused to test robustness under varied prompt and input conditions.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A parallel benchmark for non-Chinese financial documents would reveal whether the observed performance gap is language-specific or domain-specific.
  • If models trained on CFMME also improve on general financial text-only tasks, the dataset may be capturing transferable financial reasoning structures.
  • Deployment of improved models on real financial workflows could be measured by direct correlation between CFMME scores and downstream accuracy on live document streams.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The paper introduces CFMME, a Chinese financial multimodal evaluation benchmark consisting of 6,052 instances across eight image modalities and four tasks (question answering, detection, recognition, and information extraction). It evaluates representative LVLMs and reports that the state-of-the-art model achieves 66.11% accuracy on QA and an average score of 77.18 on the other tasks, concluding that there is substantial room for improvement in current LVLMs. Additional analyses on error causes, cross-modal capabilities, and multi-orientation settings are provided.

Significance. If the dataset construction and coverage claims hold, CFMME would be a valuable addition as a domain-specific multimodal benchmark in Chinese financial contexts, providing concrete performance baselines and insights that could guide LVLM development. The multi-task, multi-modal design and focus on real-world financial applications are strengths.

major comments (2)
  1. [Abstract and §3] Abstract and §3 (Dataset): The central claim that the 6,052 instances comprehensively represent perception, understanding, reasoning, and cognition demands across the entire financial business workflow rests on an unverified coverage assumption. No explicit mapping from instances to a defined workflow ontology, sampling frame from real documents, or expert validation metric is described, which is load-bearing for interpreting the 66.11% and 77.18 scores as evidence of 'substantial room for improvement' in the domain.
  2. [§4 and §5] §4 (Annotation) and §5 (Experiments): No details are provided on the annotation process, inter-annotator agreement, or statistical significance testing for the reported accuracy and average scores. This prevents verification of the reliability of the headline results and the claim that current LVLMs have substantial room for improvement.
minor comments (1)
  1. [Figures and Tables] Figure captions and tables could more explicitly link results to specific modalities or tasks for easier cross-referencing.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive feedback. We address each major comment below, providing clarifications based on the manuscript and committing to revisions where the points identify gaps in documentation.

read point-by-point responses
  1. Referee: [Abstract and §3] Abstract and §3 (Dataset): The central claim that the 6,052 instances comprehensively represent perception, understanding, reasoning, and cognition demands across the entire financial business workflow rests on an unverified coverage assumption. No explicit mapping from instances to a defined workflow ontology, sampling frame from real documents, or expert validation metric is described, which is load-bearing for interpreting the 66.11% and 77.18 scores as evidence of 'substantial room for improvement' in the domain.

    Authors: Section 3 describes the construction of CFMME by selecting instances from real Chinese financial documents across eight modalities and four tasks, spanning academic knowledge to applied scenarios. The claim of comprehensive coverage is grounded in this sampling from representative financial contexts rather than a formal ontology. However, we agree that an explicit mapping table linking tasks to workflow stages and details on the expert-guided sampling frame would strengthen verifiability. We will add this as a new subsection or appendix in the revision. revision: yes

  2. Referee: [§4 and §5] §4 (Annotation) and §5 (Experiments): No details are provided on the annotation process, inter-annotator agreement, or statistical significance testing for the reported accuracy and average scores. This prevents verification of the reliability of the headline results and the claim that current LVLMs have substantial room for improvement.

    Authors: We will revise §4 to detail the annotation pipeline, including annotator qualifications, guidelines, and inter-annotator agreement metrics such as Cohen's kappa. In §5, we will include statistical significance testing (e.g., bootstrap confidence intervals or McNemar's test) for the key accuracy figures to support the interpretation of results. These additions will be made in the revised manuscript. revision: yes

Circularity Check

0 steps flagged

No significant circularity; empirical benchmark with direct evaluation results

full rationale

The paper introduces a new dataset (CFMME) and reports direct empirical accuracy scores from running existing LVLMs on its tasks. No derivations, fitted parameters renamed as predictions, or load-bearing self-citations appear in the evaluation chain. The coverage claim for the 6052 instances is an explicit modeling assumption rather than a derived result that reduces to its own inputs. All reported numbers (66.11% QA accuracy, 77.18 average on other tasks) are measured outputs, not constructed by definition from the benchmark itself.

Axiom & Free-Parameter Ledger

0 free parameters · 0 axioms · 0 invented entities

This is an empirical benchmark paper; no mathematical derivations, fitted parameters, or new postulated entities are introduced.

reviewed 2026-06-29 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset." pith.science (2026). https://pith.science/paper/PXCU7MIB

@misc{pith2026260529462,
  author       = {Pith},
  title        = {Pith review of: Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PXCU7MIB}},
  note         = {Machine review of arXiv:2605.29462}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

The emergence of Large Vision-Language Models (LVLMs) has substantially expanded model capabilities beyond text-only understanding, enabling unified inference across both visual and textual modalities and supporting a broader range of real-world applications. To comprehensively evaluate the perception, understanding, reasoning, and cognition capabilities of LVLMs throughout the entire financial business workflow in Chinese contexts, we introduce CFMME, a novel Chinese financial multimodal evaluation benchmark. CFMME comprises 6,052 instances spanning from fundamental academic knowledge to complex real-world applications, covering eight primary financial image modalities and four core multimodal tasks. On CFMME, we conduct a thorough evaluation of representative LVLMs. The results show that the state-of-the-art model attains an overall accuracy of 66.11\% on the question answering task and an average score of 77.18 on the detection, recognition, and information extraction tasks, indicating substantial room for improvement in current LVLMs. In addition, we conduct detailed analyses of error causes, cross-modal capabilities, and multi-orientation settings, yielding valuable insights for future research. We hope that CFMME will spur further progress in LVLMs, especially by improving their performance on multiple multimodal tasks in the financial domain.

Figures

Figures reproduced from arXiv: 2605.29462 by Chi Zhang, Feng Chen, Lifan Guo, Qian Chen, Xianyin Zhang, Yanzhi Liu.

Figure 1
Figure 1. Figure 1: Overview diagram of our benchmark. Category Subcategory / Task Size Description Knowledge Subject 562 43 textbooks across diverse subjects Certification 2,256 20 types of qualification certification examinations Application Seal Detection 756 company seals and personal seals Information Extraction 875 constrained-category and open-category Seal Recognition 213 shapes of oval, round, square, triangle and rh… view at source ↗
Figure 2
Figure 2. Figure 2: Visualization of error cases in LVLMs for our [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Prompt used in the knowledge assessment. [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Prompt used for the question answering task in the application assessment. [PITH_FULL_IMAGE:figures/full_fig_p011_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Prompt used for the seal recognition task in the application assessment. [PITH_FULL_IMAGE:figures/full_fig_p012_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Prompt used for the formula recognition task in the application assessment. [PITH_FULL_IMAGE:figures/full_fig_p012_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Prompt used for the table recognition task in the application assessment. [PITH_FULL_IMAGE:figures/full_fig_p012_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Prompt used for the detection task in the application assessment. [PITH_FULL_IMAGE:figures/full_fig_p012_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Prompt used for the information extraction task in the application assessment. [PITH_FULL_IMAGE:figures/full_fig_p013_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: An example in subject subset of the knowledge assessment. [PITH_FULL_IMAGE:figures/full_fig_p014_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: An example in the certification subset of the knowledge assessment. [PITH_FULL_IMAGE:figures/full_fig_p015_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: An example of the question answering task in the application assessment. [PITH_FULL_IMAGE:figures/full_fig_p016_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: An example of the seal recognition task in the application assessment. [PITH_FULL_IMAGE:figures/full_fig_p016_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: An example of the formula recognition task in the application assessment. [PITH_FULL_IMAGE:figures/full_fig_p016_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: An example of the table recognition task in the application assessment. [PITH_FULL_IMAGE:figures/full_fig_p017_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: An example of the detection task in the application assessment. [PITH_FULL_IMAGE:figures/full_fig_p018_16.png] view at source ↗
Figure 17
Figure 17. Figure 17: An example of the information extraction task in the application assessment. [PITH_FULL_IMAGE:figures/full_fig_p018_17.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

4 extracted references · 3 canonical work pages · 1 internal anchor

  1. [1]

    arXiv preprint arXiv:2401.14011

    Cmmu: A benchmark for chinese multi-modal multi-type question understanding and reasoning. arXiv preprint arXiv:2401.14011. Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Stein- hardt. 2021. Measuring massive multitask language understanding.Proceedings of the International Con- ference on Learning Representatio...

  2. [2]

    When FLUE meets FLANG: Benchmarks and large pretrained language model for financial domain,

    When flue meets flang: Benchmarks and large pre-trained language model for financial do- main.arXiv preprint arXiv:2211.00083. Baidu ERNIE Team. 2025. Ernie 4.5 technical report. Kimi Team, Angang Du, Bohong Yin, Bowei Xing, Bowen Qu, Bowen Wang, Cheng Chen, Chenlin Zhang, Chenzhuang Du, Chu Wei, and 1 others. 2025a. Kimi-vl technical report.arXiv preprin...

  3. [3]

    InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

    Cdm: A reliable metric for fair and accurate formula recognition evaluation.arXiv e-prints, pages arXiv–2409. Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, Long Cui, Xingguang Wei, Zhaoyang Liu, Linglin Jing, Shenglong Ye, Jie Shao, and 1 others. 2025. In- ternvl3. 5: Advancing open-source multimodal mod- els in versatility, reasoning, and efficiency.a...

  4. [4]

    bbox_2d": [x_min, y_min, x_max, y_max],

    IEEE. Xu Zhong, Elaheh ShafieiBavani, and Antonio Ji- meno Yepes. 2020. Image-based table recognition: data, model, and evaluation. InEuropean conference on computer vision, pages 564–580. Springer. Jie Zhu, Junhui Li, Yalong Wen, and Lifan Guo. 2024. Benchmarking large language models on cflue–a chinese financial language understanding evaluation dataset...

This paper was first reviewed by grok-4.3 on June 29, 2026.