Pith. sign in

REVIEW 1 cited by

MARPLE: A Benchmark for Long-Horizon Inference

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.01926 v1 pith:K5H7VTRJ submitted 2024-10-02 cs.LG

classification cs.LG
keywords inferencebenchmarkevidencehumanlong-horizonmodelswhatagent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reconstructing past events requires reasoning across long time horizons. To figure out what happened, we need to use our prior knowledge about the world and human behavior and draw inferences from various sources of evidence including visual, language, and auditory cues. We introduce MARPLE, a benchmark for evaluating long-horizon inference capabilities using multi-modal evidence. Our benchmark features agents interacting with simulated households, supporting vision, language, and auditory stimuli, as well as procedurally generated environments and agent behaviors. Inspired by classic ``whodunit'' stories, we ask AI models and human participants to infer which agent caused a change in the environment based on a step-by-step replay of what actually happened. The goal is to correctly identify the culprit as early as possible. Our findings show that human participants outperform both traditional Monte Carlo simulation methods and an LLM baseline (GPT-4) on this task. Compared to humans, traditional inference models are less robust and performant, while GPT-4 has difficulty comprehending environmental changes. We analyze what factors influence inference performance and ablate different modes of evidence, finding that all modes are valuable for performance. Overall, our experiments demonstrate that the long-horizon, multimodal inference tasks in our benchmark present a challenge to current models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Breakpoint generates code-repair benchmarks by corrupting real GitHub functions and shows that frontier AI coding agents solve easy repairs but fail completely on tasks requiring coordinated, system-wide changes.

Pith tools