Pith. sign in

REVIEW 2 cited by

Multimodal Multi-Hop Question Answering Through a Conversation Between Tools and Efficiently Finetuned Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.08922 v1 pith:6NWK3NSQ submitted 2023-09-16 cs.CL

classification cs.CL
keywords languagelargemodelsmulti-hopmultimodalquestiontoolscomplex
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We employ a tool-interacting divide-and-conquer strategy enabling large language models (LLMs) to answer complex multimodal multi-hop questions. In particular, we harness the power of large language models to divide a given multimodal multi-hop question into unimodal single-hop sub-questions to be answered by the appropriate tool from a predefined set of tools. After all corresponding tools provide the LLM with their answers, the LLM generates the next relevant unimodal single-hop question. To increase the reasoning ability of LLMs, we prompt chatGPT to generate a tool-interacting divide-and-conquer dataset. This dataset is then used to efficiently finetune the corresponding LLM. To assess the effectiveness of this approach, we conduct an evaluation on two recently introduced complex question-answering datasets. The experimental analysis demonstrate substantial improvements over existing state-of-the-art solutions, indicating the efficacy and generality of our strategy

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FCMR: Robust Evaluation of Financial Cross-Modal Multi-Hop Reasoning

    cs.CL 2024-12 conditional novelty 6.0 of 10

    FCMR is a finance-domain benchmark where every question requires combining text, tables, and charts; state-of-the-art MLLMs reach only 30.4% accuracy on the hardest tier.

  2. How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey

    cs.CV 2024-12 conditional novelty 4.0 of 10

    A survey that categorizes pre-trained-model-based vision-language methods into four challenge-driven paradigms, with performance tables and a discussion of risks.

Pith tools