REVIEW 6 cited by
SpartQA: : A Textual Question Answering Benchmark for Spatial Reasoning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
This paper proposes a question-answering (QA) benchmark for spatial reasoning on natural language text which contains more realistic spatial phenomena not covered by prior work and is challenging for state-of-the-art language models (LM). We propose a distant supervision method to improve on this task. Specifically, we design grammar and reasoning rules to automatically generate a spatial description of visual scenes and corresponding QA pairs. Experiments show that further pretraining LMs on these automatically generated data significantly improves LMs' capability on spatial understanding, which in turn helps to better solve two external datasets, bAbI, and boolQ. We hope that this work can foster investigations into more sophisticated models for spatial reasoning over text.
Forward citations
Cited by 6 Pith papers
-
SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards
Dense scene-graph-grounded rewards let a 7B multimodal LLM trained on 7K synthetic questions beat SFT and sparse-RL baselines and outscore GPT-4o on average across 12 spatial/real-world benchmarks.
-
IMoRe: Implicit Program-Guided Reasoning for Human Motion Q&A
IMoRe couples a MAC-style memory network with program-function embeddings and multi-level ViT motion features to reach state-of-the-art accuracy on Babel-QA and a new HuMMan-QA benchmark.
-
USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning of LLMs as Urban Agents
USTBench is the first benchmark that decomposes urban spatiotemporal reasoning into understanding, forecasting, planning, and reflection, and shows LLMs struggle most with planning and reflection.
-
Exploring Spatial Language Grounding Through Referring Expressions
On the CopsRef dataset, referring expression comprehension reveals that vision-language models fail at dynamic spatial relations, multi-relation expressions, and negations, while task-specific models handle geometric ...
-
Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning
A new synthetic-image benchmark, CDR, shows that leading multimodal models mostly guess on compass direction questions, with chain-of-thought fine-tuning only partially closing the gap.
-
Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes
The authors argue that spatial reasoning in multimodal LLMs requires dedicated new recipes in data, architecture, and training objectives, and will not emerge from scaling alone.
Discussion (0). Continue with ORCID to comment.