Pith. sign in

REVIEW 10 cited by

Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.00754 v2 pith:GSRFRWAD submitted 2024-08-01 cs.CV cs.LG

classification cs.CVcs.LG
keywords correspondencesreasoningcoarsemllmsspatial-temporalappliedbenchmarkfine-tuning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal language models (MLLMs) are increasingly being applied in real-world environments, necessitating their ability to interpret 3D spaces and comprehend temporal dynamics. Current methods often rely on specialized architectural designs or task-specific fine-tuning to achieve this. We introduce Coarse Correspondences, a simple lightweight method that enhances MLLMs' spatial-temporal reasoning with 2D images as input, without modifying the architecture or requiring task-specific fine-tuning. Our method uses a lightweight tracking model to identify primary object correspondences between frames in a video or across different image viewpoints, and then conveys this information to MLLMs through visual prompting. We demonstrate that this simple training-free approach brings substantial gains to GPT4-V/O consistently on four benchmarks that require spatial-temporal reasoning, including +20.5\% improvement on ScanQA, +9.7\% on OpenEQA's episodic memory subset, +6.0\% on the long-form video benchmark EgoSchema, and +11\% on the R2R navigation benchmark. Additionally, we show that Coarse Correspondences can also enhance open-source MLLMs' spatial reasoning (by +6.9\% on ScanQA) when applied in both training and inference and that the improvement can generalize to unseen datasets such as SQA3D (+3.1\%). Taken together, we show that Coarse Correspondences effectively and efficiently boosts models' performance on downstream tasks requiring spatial-temporal reasoning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames

    cs.CV 2025-05 conditional novelty 7.0 of 10

    A new egocentric benchmark shows vision-language models fail at spatial reasoning across disjoint frames, falling 28 points behind humans and only improving sharply when handed ground-truth 3D coordinates.

  2. GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

    cs.CV 2025-01 conditional novelty 7.0 of 10

    Consistent object IDs across video frames and a reconstructed bird's-eye view let large VLMs, and a fine-tuned 7B model, reach state-of-the-art results on ScanNet 3D question answering, captioning, and grounding.

  3. From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models

    cs.CV 2026-02 conditional novelty 6.0 of 10

    A 3B multimodal LLM trained with patch-level cross-view alignment plus explicit viewpoint-action reasoning outperforms much larger models on two multi-image spatial reasoning benchmarks.

  4. Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A 3B agent trained with supervised tool-use traces and reinforcement learning answers week-long egocentric video questions by dynamically selecting hierarchical retrieval, video-LLM, and VLM tools.

  5. AD^2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions

    cs.CV 2025-06 conditional novelty 6.0 of 10

    AD^2-Bench is a new adverse-weather driving benchmark with hierarchical chain-of-thought annotations and LLM-based quality metrics; 12 MLLMs all scored below 60%.

  6. Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

    cs.CV 2025-05 reject novelty 6.0 of 10

    VeBrain unifies perception, spatial reasoning, and robot control in one MLLM by representing control as keypoint detection plus skill selection, with a robotic adapter for deployment.

  7. Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A new benchmark with clean, corrupted, and text-only driving inputs shows that vision-language models can answer many driving questions without visual information, so standard accuracy metrics overestimate visual grounding.

  8. Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Adding 3D coordinate encodings to sampled RGB-D video frames lets a video LLM beat prior 3D scene understanding models on five benchmarks.

  9. UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning

    cs.CV 2025-05 conditional novelty 5.0 of 10

    UniVG-R1 uses CoT supervised fine-tuning plus GRPO with difficulty-aware reweighting to make Qwen2-VL substantially better at multi-image, reasoning-based visual grounding.

  10. Ponder & Press: Advancing Visual GUI Agent towards General Computer Control

    cs.CV 2024-12 conditional novelty 4.0 of 10

    A screenshot-only GUI agent, built from an instruction-interpreting MLLM plus a LoRA-fine-tuned Qwen2-VL locator, reports state-of-the-art grounding and task success on five GUI benchmarks.

Pith tools