Pith. sign in

REVIEW 6 cited by

FAMMA: A Benchmark for Financial Domain Multilingual Multimodal Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.04526 v4 pith:2VUJUEF5 submitted 2024-10-06 cs.CL cs.AI

classification cs.CLcs.AI
keywords reasoningunderlinebenchmarkdatafammamodelsquestionsquestion
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this paper, we introduce FAMMA, an open-source benchmark for \underline{f}in\underline{a}ncial \underline{m}ultilingual \underline{m}ultimodal question \underline{a}nswering (QA). Our benchmark aims to evaluate the abilities of large language models (LLMs) in answering complex reasoning questions that require advanced financial knowledge. The benchmark has two versions: FAMMA-Basic consists of 1,945 questions extracted from university textbooks and exams, along with human-annotated answers and rationales; FAMMA-LivePro consists of 103 novel questions created by human domain experts, with answers and rationales held out from the public for a contamination-free evaluation. These questions cover advanced knowledge of 8 major subfields in finance (e.g., corporate finance, derivatives, and portfolio management). Some are in Chinese or French, while a majority of them are in English. Each question has some non-text data such as charts, diagrams, or tables. Our experiments reveal that FAMMA poses a significant challenge on LLMs, including reasoning models such as GPT-o1 and DeepSeek-R1. Additionally, we curated 1,270 reasoning trajectories of DeepSeek-R1 on the FAMMA-Basic data, and fine-tuned a series of open-source Qwen models using this reasoning data. We found that training a model on these reasoning trajectories can significantly improve its performance on FAMMA-LivePro. We released our leaderboard, data, code, and trained models at https://famma-bench.github.io/famma/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DrafterBench: Benchmarking Large Language Models for Tasks Automation in Civil Engineering

    cs.AI 2025-07 conditional novelty 6.0 of 10

    DrafterBench is a new benchmark of 1,920 PDF drawing-revision tasks; on it, the best model (OpenAI o1) averages about 80/100, and all tested models fail hard on incomplete instructions and plan execution.

  2. Understanding Financial Reasoning in AI: A Multimodal Benchmark and Error Learning Approach

    cs.AI 2025-04 conditional novelty 6.0 of 10

    A new multimodal financial reasoning benchmark and a retrieval-based error feedback prompting method that improves model accuracy, with the improvement partly confounded by information leakage.

  3. FinLMM-R1: Enhancing Financial Reasoning in LMM through Scalable Data and Reward Design

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A two-stage RL framework with length, image-selection, and adversarial rewards, trained on 89,378 ASP-built financial image-question pairs, improves multimodal reasoning over LMM-R1.

  4. Multimodal Financial Foundation Models (MFFMs): Progress, Prospects, and Challenges

    cs.CE 2025-05 conditional novelty 4.0 of 10

    A position and survey paper argues that multimodal financial foundation models are the next step beyond text-only financial AI, and it organizes current data, benchmarks, models, and challenges around that claim.

  5. Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering

    cs.CV 2025-02 conditional novelty 4.0 of 10

    A structured survey of 3D Scene Question Answering that categorizes datasets, methods, and metrics and finds a common encoder-fusion-prediction pipeline across approaches.

  6. ROMAS: A Role-Based Multi-Agent System for Database monitoring and Planning

    cs.AI 2024-12 reject novelty 4.0 of 10

    A role-based multi-agent framework with a monitor that triggers re-planning is reported to outperform other LLM agent systems on two QA benchmarks, but no code, data, or error bars are provided.

Pith tools