Pith. sign in

REVIEW 1 cited by

Evaluating LLMs with Multiple Problems at once

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.10786 v3 pith:LZVKZ2CB submitted 2024-06-16 cs.AI cs.CL

classification cs.AIcs.CL
keywords multiplellmsproblemshandlingsingleevaluationmulti-problemproblem
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper shows the benefits and fruitfulness of evaluating LLMs with multiple problems at once, a paradigm we call multi-problem evaluation (MPE). Unlike conventional single-problem evaluation, where a prompt presents a single problem and expects one specific answer, MPE places multiple problems together in a single prompt and assesses how well an LLM answers all these problems in a single output. Leveraging 6 classification and 12 reasoning benchmarks that already exist, we introduce a new benchmark called ZeMPE (Zero-shot Multi-Problem Evaluation), comprising 53,100 zero-shot multi-problem prompts. We experiment with a total of 13 LLMs from 5 model families on ZeMPE to present a comprehensive and systematic MPE. Our results show that LLMs are capable of handling multiple problems from a single data source as well as handling them separately, but there are conditions this multiple problem handling capability falls short. In addition, we perform in-depth further analyses and explore model-level factors that may enable multiple problem handling capabilities in LLMs. We release our corpus and code to facilitate future research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Resurrecting saturated LLM benchmarks with adversarial encoding

    cs.LG 2025-02 conditional novelty 4.0 of 10

    Pairing questions and adding distractor options reliably lowers LLM scores across three benchmarks, and the authors use this to create a harder 'resurrected' version of MMLU.

Pith tools