Pith. sign in

REVIEW 3 cited by

MultiHiertt: Numerical Reasoning over Multi Hierarchical Tabular and Textual Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.01347 v1 pith:2Q5GZVKG submitted 2022-06-03 cs.AI cs.LG

classification cs.AIcs.LG
keywords reasoningdatamultihierttfactshierarchicalnumericaltablesexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Numerical reasoning over hybrid data containing both textual and tabular content (e.g., financial reports) has recently attracted much attention in the NLP community. However, existing question answering (QA) benchmarks over hybrid data only include a single flat table in each document and thus lack examples of multi-step numerical reasoning across multiple hierarchical tables. To facilitate data analytical progress, we construct a new large-scale benchmark, MultiHiertt, with QA pairs over Multi Hierarchical Tabular and Textual data. MultiHiertt is built from a wealth of financial reports and has the following unique characteristics: 1) each document contain multiple tables and longer unstructured texts; 2) most of tables contained are hierarchical; 3) the reasoning process required for each question is more complex and challenging than existing benchmarks; and 4) fine-grained annotations of reasoning processes and supporting facts are provided to reveal complex numerical reasoning. We further introduce a novel QA model termed MT2Net, which first applies facts retrieving to extract relevant supporting facts from both tables and text and then uses a reasoning module to perform symbolic reasoning over retrieved facts. We conduct comprehensive experiments on various baselines. The experimental results show that MultiHiertt presents a strong challenge for existing baselines whose results lag far behind the performance of human experts. The dataset and code are publicly available at https://github.com/psunlpgroup/MultiHiertt.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Orthogonal Hierarchical Decomposition for Structure-Aware Table Understanding with Large Language Models

    cs.CL 2026-02 conditional novelty 6.0 of 10

    A framework that splits tables into row-tree and column-tree representations and uses an LLM to arbitrate both achieves large gains on complex-table QA benchmarks, though one backbone underperforms on HiTab.

  2. LaRe: Latent Refocusing for Multimodal Reasoning

    cs.CV 2025-11 reject novelty 6.0 of 10

    LaRe performs iterative visual refocusing in latent space and reports accuracy gains with fewer tokens, but its main experiments compare against baselines trained with less data.

  3. No Universal Prompt: Unifying Reasoning through Adaptive Prompting for Temporal Table Reasoning

    cs.CL 2025-06 reject novelty 5.0 of 10

    An adaptive prompting framework called SEAR is claimed to beat fixed prompting strategies across temporal table QA datasets, but reported scores contradict the headline and rely on an unvalidated metric.

Pith tools