Pith. sign in

REVIEW 6 cited by

SpartQA: : A Textual Question Answering Benchmark for Spatial Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.05832 v1 pith:KHISQ6YZ submitted 2021-04-12 cs.CL cs.AI

classification cs.CLcs.AI
keywords spatialreasoningautomaticallybenchmarklanguagemodelstextwork
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This paper proposes a question-answering (QA) benchmark for spatial reasoning on natural language text which contains more realistic spatial phenomena not covered by prior work and is challenging for state-of-the-art language models (LM). We propose a distant supervision method to improve on this task. Specifically, we design grammar and reasoning rules to automatically generate a spatial description of visual scenes and corresponding QA pairs. Experiments show that further pretraining LMs on these automatically generated data significantly improves LMs' capability on spatial understanding, which in turn helps to better solve two external datasets, bAbI, and boolQ. We hope that this work can foster investigations into more sophisticated models for spatial reasoning over text.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards

    cs.CV 2025-11 conditional novelty 6.0 of 10

    Dense scene-graph-grounded rewards let a 7B multimodal LLM trained on 7K synthetic questions beat SFT and sparse-RL baselines and outscore GPT-4o on average across 12 spatial/real-world benchmarks.

  2. IMoRe: Implicit Program-Guided Reasoning for Human Motion Q&A

    cs.CV 2025-08 conditional novelty 6.0 of 10

    IMoRe couples a MAC-style memory network with program-function embeddings and multi-level ViT motion features to reach state-of-the-art accuracy on Babel-QA and a new HuMMan-QA benchmark.

  3. USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning of LLMs as Urban Agents

    cs.AI 2025-05 conditional novelty 6.0 of 10

    USTBench is the first benchmark that decomposes urban spatiotemporal reasoning into understanding, forecasting, planning, and reflection, and shows LLMs struggle most with planning and reflection.

  4. Exploring Spatial Language Grounding Through Referring Expressions

    cs.CL 2025-02 conditional novelty 6.0 of 10

    On the CopsRef dataset, referring expression comprehension reveals that vision-language models fail at dynamic spatial relations, multi-relation expressions, and negations, while task-specific models handle geometric ...

  5. Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning

    cs.AI 2024-12 conditional novelty 6.0 of 10

    A new synthetic-image benchmark, CDR, shows that leading multimodal models mostly guess on compass direction questions, with chain-of-thought fine-tuning only partially closing the gap.

  6. Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes

    cs.LG 2025-04 conditional novelty 4.0 of 10

    The authors argue that spatial reasoning in multimodal LLMs requires dedicated new recipes in data, architecture, and training objectives, and will not emerge from scaling alone.

Pith tools