Pith. sign in

REVIEW 6 cited by

MME-Finance: A Multimodal Finance Benchmark for Expert-level Understanding and Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.03314 v1 pith:SLOKBCKM submitted 2024-11-05 cs.CV cs.CL

classification cs.CVcs.CL
keywords financialmodelschartsmultimodalgeneralbenchmarkbenchmarksdevelopment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years, multimodal benchmarks for general domains have guided the rapid development of multimodal models on general tasks. However, the financial field has its peculiarities. It features unique graphical images (e.g., candlestick charts, technical indicator charts) and possesses a wealth of specialized financial knowledge (e.g., futures, turnover rate). Therefore, benchmarks from general fields often fail to measure the performance of multimodal models in the financial domain, and thus cannot effectively guide the rapid development of large financial models. To promote the development of large financial multimodal models, we propose MME-Finance, an bilingual open-ended and practical usage-oriented Visual Question Answering (VQA) benchmark. The characteristics of our benchmark are finance and expertise, which include constructing charts that reflect the actual usage needs of users (e.g., computer screenshots and mobile photography), creating questions according to the preferences in financial domain inquiries, and annotating questions by experts with 10+ years of experience in the financial industry. Additionally, we have developed a custom-designed financial evaluation system in which visual information is first introduced in the multi-modal evaluation process. Extensive experimental evaluations of 19 mainstream MLLMs are conducted to test their perception, reasoning, and cognition capabilities. The results indicate that models performing well on general benchmarks cannot do well on MME-Finance; for instance, the top-performing open-source and closed-source models obtain 65.69 (Qwen2VL-72B) and 63.18 (GPT-4o), respectively. Their performance is particularly poor in categories most relevant to finance, such as candlestick charts and technical indicator charts. In addition, we propose a Chinese version, which helps compare performance of MLLMs under a Chinese context.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A new benchmark shows LLMs' financial calculation accuracy collapses without explicit formulas and degrades further when they must generate multi-metric tables.

  2. FinGAIA: A Chinese Benchmark for AI Agents in Real-World Financial Domain

    cs.CL 2025-07 conditional novelty 6.0 of 10

    FinGAIA is a 407-task Chinese financial agent benchmark where the best agent, ChatGPT DeepResearch, scores 48.9%, far below financial experts at 84.7%.

  3. CFBenchmark-MM: Chinese Financial Assistant Benchmark for Multimodal Large Language Model

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A 9,356-pair Chinese multimodal financial benchmark reveals that state-of-the-art multimodal LLMs, including GPT-4V, still score below 53% on objective and 39% on subjective financial chart tasks.

  4. FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning Evaluation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    FinMME is a new 11,099-sample financial chart benchmark where top AI models average around 50% and FinScore adds penalties for guessing.

  5. FinLMM-R1: Enhancing Financial Reasoning in LMM through Scalable Data and Reward Design

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A two-stage RL framework with length, image-selection, and adversarial rewards, trained on 89,378 ASP-built financial image-question pairs, improves multimodal reasoning over LMM-R1.

  6. AI Trading: Evaluating Large Language Models for Technical Market Analysis

    cs.LG 2026-07 reject novelty 4.0 of 10

    A comparative evaluation claims GPT-4 Turbo and FinGPT outperformed the S&P 500 in a 2023 simulated backtest, but flawed baselines and missing code/data undermine the result.

Pith tools