Pith. sign in

Paper Citation Record · LEDGER

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching

As of 6 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2606.03577.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.03577 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T10:48:24.523702Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

71 of 71 outbound references displayed

  • verified exact24
  • verified fuzzy0
  • unresolved44
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6ae4d19d-96f5-4256-b781-81c3a4e02458 · outbound

This paper cites Qwen3-VL Technical Report.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Qwen3-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:36:27.073648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:ad32661cd3818bd73dd3e3f1c644d1e15d5e29f3c8d33187550a3eb9da5bbf74

Observation 1dd6f1b6-e6fe-47e5-9ccf-1113c2729634 · outbound

This paper cites Qwen2.5-vl technical report, 2025.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Qwen2.5-vl technical report, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:8ee2c90d81dc963815e9b1cad14a82b6ac9d1b085c701728185a114fde4f6c00

Observation 89364235-0444-4515-8ee0-27aaf8e1c766 · outbound

This paper cites Hot3d: Hand and object tracking in 3d from egocentric multi-view videos.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Hot3d: Hand and object tracking in 3d from egocentric multi-view videos

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:86e655a04d09384aa54b377dc895f08dfdbe4837526b7180cb9e4faac38372cb

Observation 0813fcfc-354e-457d-b509-a2387b0d01c6 · outbound

This paper cites Graph-cut ransac.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Graph-cut ransac

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:4c479dda12c3bf960cce6d33867154224a5d5ddb5cb4020069a9c85490dd8f48

Observation eee1b83d-cb4c-4b6e-a8e0-51f5e45fa590 · outbound

This paper cites Magsac: marginalizing sample consensus.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Magsac: marginalizing sample consensus

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:40ae11f1fe227304eda6756140df08b56fac4c2cba4a23a3cf86dacbec31abda

Observation 68a30c06-a3af-4a02-ab5a-95bffb1f5cee · outbound

This paper cites Surf: Speeded up robust features.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Surf: Speeded up robust features

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:02458fe08dc6b0434ce5e4eb134f2a168c3d772f0c28f1cfb52eb01d8369b8f6

Observation e44bacd5-5e46-41ba-9e82-6a274be1912b · outbound

This paper cites Depth Pro: Sharp Monocular Metric Depth in Less Than a Second.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Depth Pro: Sharp Monocular Metric Depth in Less Than a Second

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:36:27.111959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:9ab8c0bed5ecda68c458f36cd30907ec38f256a4e8a56b9e9eb8a8c33d6c3d91

Observation a6f9749d-9730-47c1-b5c4-d20f0c48171c · outbound

This paper cites Assignment problems: revised reprint.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Assignment problems: revised reprint

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:077449266e606a520df882d831d64732fa0a8905c3a19943f0a9e11c467a6daa

Observation c1ccb55f-fade-44c5-a941-4f5298ab4061 · outbound

This paper cites SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.141917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:7367b182ad6fbfb255735adb8654c8e562a1a2c9f5f81c4858b56fa36850a849

Observation 705572a1-80de-4b2a-b8c0-5aee124588c2 · outbound

This paper cites Are we on the right way for evaluating large vision-language models?Advances in Neural Informa- tion Processing Systems, 37:27056–27087, 2024.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Are we on the right way for evaluating large vision-language models?Advances in Neural Informa- tion Processing Systems, 37:27056–27087, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:a62204b73a34ac80089c9c3fa0b8bb701120049fc7cfeaead0d1623ce45686af

Observation dfa49c86-523a-47ed-90ed-71d5b80d8d96 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:36:27.084858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:6997f83ce5fe92abf94900767711bfc6d4db1583e0cb2037ff3b734876193949

Observation a01720ff-a0bd-4111-bf13-23da3c5b1999 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:8e45aeb23c833f7e9f7e40ae574b053eb63f916de662e0eef6e5ef4e7cdc521e

Observation 63ef0228-c17b-4df3-8f59-d5de49d22778 · outbound

This paper cites Superpoint: Self-supervised interest point detection and description.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Superpoint: Self-supervised interest point detection and description

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:999f28b28894cf0c98fd33b102f65cad78ad4e7bec0fbf711f88ffd2fdf65089

Observation 77cc1285-2155-41fa-9d5e-fe6602069f07 · outbound

This paper cites Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.097894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:9c2a4ee3730995cfb5cf232c9feb7d6faf042abf3281020c29f52696984e8c26

Observation e7a02385-b46b-42b8-ac01-8abcbc39fb28 · outbound

This paper cites D2-net: A trainable cnn for joint detection and description of local features.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching D2-net: A trainable cnn for joint detection and description of local features

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:cde3082d579f27b38dcc2b530d7fb5811d52fc8a058e16522fed8f93b5ab68f1

Observation 0722f048-ecc0-4eba-840d-a46b94ed7f30 · outbound

This paper cites Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography.Communications of the ACM, 24(6):381–395, 1981.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography.Communications of the ACM, 24(6):381–395, 1981

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:d86dc55bb19c581b65440a5550f3ab18398be6b9c40f2c1e0b6ea069bffb4b85

Observation 5d928d06-39c7-4997-89f6-e414c19740b3 · outbound

This paper cites A graduated assignment algorithm for graph matching.IEEE Transactions on pattern analysis and machine intelligence, 18(4):377–388, 2002.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching A graduated assignment algorithm for graph matching.IEEE Transactions on pattern analysis and machine intelligence, 18(4):377–388, 2002

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:5e3ac80bba20adafd58b0b82b7edbe1c26ecd6a8e67b3ce3a1ca16b052bdffcd

Observation 7757d196-dabc-4a43-a121-62e6d5dbe551 · outbound

This paper cites Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.Nature, 645(8081):633– 638, 2025.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.Nature, 645(8081):633– 638, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:9d11b33490dc00cfc1ca5db2cca32a89d9eceaa41b4cafb6cd07f78be5710ef4

Observation cb58b551-ce95-4bd7-8481-5b8e60218292 · outbound

This paper cites Projective reconstruction and invariants from multiple images.IEEE Transactions on Pattern Analy- sis and Machine Intelligence, 16(10):1036–1041, 1994.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Projective reconstruction and invariants from multiple images.IEEE Transactions on Pattern Analy- sis and Machine Intelligence, 16(10):1036–1041, 1994

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:bdf2eee687c721421b22858e83b3e2d7905da3e045326aeb48ed976bd26b8e47

Observation 401c8c08-c5ba-4db5-afb7-0b9f285f3002 · outbound

This paper cites Metric3d v2: A versatile monocular geomet- ric foundation model for zero-shot metric depth and surface normal estimation.IEEE TPAMI, 2024.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Metric3d v2: A versatile monocular geomet- ric foundation model for zero-shot metric depth and surface normal estimation.IEEE TPAMI, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:98cd49fe6cc764bf668f20e72bdc5bf1881c9d064dd6a6ea5957ab4e1bb589ad

Observation 026976d1-559f-4146-a53f-a64a23478789 · outbound

This paper cites NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:36:27.129942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:44d09551d5c42c698936fbbdf460f4cf2bec7e263c20764002a82a632b8350b8

Observation 1954a7c0-0776-4997-9ddc-a72de2a85f7f · outbound

This paper cites RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.070919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:774c7462423ca7e4a83d55c647b0160d08e05739d54a52fa485288ebb2ea4039

Observation ddfa18ac-5c9e-4185-934d-a56d26e986bc · outbound

This paper cites Omnispatial: Towards comprehensive spatial reasoning benchmark for vision language models.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Omnispatial: Towards comprehensive spatial reasoning benchmark for vision language models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.076549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:18f1eb61a75b09c8090b2035f6837575184115c7544bddeb3d1eaadd56a837d4

Observation 642fa5f6-ee6a-487c-8480-dda05b518f07 · outbound

This paper cites Image Matching across Wide Baselines: From Paper to Practice.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Image Matching across Wide Baselines: From Paper to Practice

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.079352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:e419c25e2292089cf23683aa06a408b8e729222049346721104edbb90fd7365c

Observation a9628a8c-e701-4de6-af9f-848344b9c5b5 · outbound

This paper cites Segment any- thing.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Segment any- thing

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:db1151536207636e9e6c519a20310e26284d7ccafc688c9822ea4b5f1ef1709b

Observation 98f32c98-ebb9-4537-ab21-6110ca5c192a · outbound

This paper cites Building and better understanding vision- language models: Insights and future directions.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Building and better understanding vision- language models: Insights and future directions

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:bed6d5e7aa46fae53de0b022c684fc25f7aed8fa6a0fbafe0ef3af430c5fe26a

Observation 5485ff7a-a715-453b-91e4-b0e02c70a3fe · outbound

This paper cites Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:51be981cb9080663ad0245dbc28d5615100e6cd9fa59137d714a7504eae262d5

Observation 18bf7330-5e3a-4580-80fb-41f7438dcda9 · outbound

This paper cites Improved baselines with visual instruction tuning.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Improved baselines with visual instruction tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:cfc5ac071f01a990f929bfb055365d01b93346eefca6efec34db60b5a494278a

Observation f18e4c8c-ca52-4718-bf4d-3d50251454ed · outbound

This paper cites Zhang, Natalia Neverova, Andrea Vedaldi, Roman Shapovalov, and David Novotny.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Zhang, Natalia Neverova, Andrea Vedaldi, Roman Shapovalov, and David Novotny

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:3cf644a60713cdb8c5c8f4395adde201c0429204bad280f01341e2c00b94054f

Observation 99dd230c-41ca-493d-b93b-3c4b252498d2 · outbound

This paper cites Decoupled Weight Decay Regularization.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Decoupled Weight Decay Regularization

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:36:27.132415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:0bdc06771db88219878591068e9042bedbac56f67c455672f80e27ab4fa9d8bc

Observation bb3d0ad6-5e7a-4760-9381-041ac7f3081d · outbound

This paper cites Distinctive image features from scale- invariant keypoints.International journal of computer vi- sion, 60(2):91–110, 2004.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Distinctive image features from scale- invariant keypoints.International journal of computer vi- sion, 60(2):91–110, 2004

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:aa0769a346117c0c8b6dc7b20b2c2c8074c9856ba1d0e63d678c249297ff5d14

Observation f067a1f9-4463-4d8a-ac2c-ab2324be276d · outbound

This paper cites 3dsrbench: A comprehensive 3d spatial reasoning benchmark.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching 3dsrbench: A comprehensive 3d spatial reasoning benchmark

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:35b3f0adad1c0b63f699f729f1a1737ec552514839d8d06a2d204eb03faf6e1a

Observation 6da8823c-ec06-40dc-ab08-1ad93192093b · outbound

This paper cites Working hard to know your neighbor’s mar- gins: Local descriptor learning loss.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Working hard to know your neighbor’s mar- gins: Local descriptor learning loss

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:1c372963e5769d35ca003d3b76eaf873222c575881de2433c6e76bcc31a01a23

Observation 34f60a06-49f0-4d99-8083-1ca572070ed0 · outbound

This paper cites GPT-4 Technical Report.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching GPT-4 Technical Report

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:36:27.122127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:579c0211a108ce2ea4a3a917760aacb6f7c823a1fb2e5f920beadf4165ad3bba

Observation 14e5a29d-44af-473e-891f-14b87e8ecdef · outbound

This paper cites Hello gpt-4o.OpenAI Blog, 2024.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Hello gpt-4o.OpenAI Blog, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:246de710d291b5d63b0d5d502ff8b837fd805ab11c8c065000a1c3e5cef109d4

Observation c3d48be2-94b0-4323-9ee4-e0545ddd7090 · outbound

This paper cites Manivideo: Generating hand-object manipulation video with dexterous and generalizable grasping.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Manivideo: Generating hand-object manipulation video with dexterous and generalizable grasping

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:8f636b2bf723f2dff1dcee32b110b10302d5ef8d47bd787d53f0df58abd55be3

Observation b0820559-58e6-425a-9fe4-bb31a18f40bd · outbound

This paper cites Wide baseline stereo matching.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Wide baseline stereo matching

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:0dfdf2ad7e3ce95d240afa72b53291c084c5f7640d187f368f35610a0d3becaa

Observation d011b345-3398-475d-b9ab-fe9955d2186b · outbound

This paper cites Sat: Spa- tial aptitude training for multimodal language models.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Sat: Spa- tial aptitude training for multimodal language models

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:36:27.124633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:6bff02879e0646ad6cc8b061a7acdcffe6105ed54ed24f6e74f6c2b0f185ab1d

Observation 496f5e75-c002-42fa-8b79-0ce887b55b4f · outbound

This paper cites Com- mon objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Com- mon objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:f42e1ddbd937d44bd19e837a2e5a1de79ac0bf523490d18c308edcce10a6856e

Observation 8b3afa1c-4231-4cf3-957e-ea354cfcd599 · outbound

This paper cites R2D2: Repeatable and Reliable Detector and Descriptor.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching R2D2: Repeatable and Reliable Detector and Descriptor

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:36:27.112698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:ce8f2df3acb09fdcf8d188c510e1f9e077cec843900a6a50e07469f4cb4fcb77

Observation bd27d2d4-6512-48e6-b00a-63197ea4a73f · outbound

This paper cites Orb: An efficient alternative to sift or surf.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Orb: An efficient alternative to sift or surf

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:15d997e3ac7e3170e83876898af571d86fb1441c598b776522bad7e4408460b0

Observation ab807407-ee27-4dbd-914f-2bb57f1c882b · outbound

This paper cites PhD thesis, INRIA, 1995.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching PhD thesis, INRIA, 1995

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:083d03ebe8cfe3e57bde51feee13a763f281e7c6db95b52a24b2152c416b37fb

Observation f5744ddd-cfb1-4fe1-adad-be546c13c85f · outbound

This paper cites Structure- from-motion revisited.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Structure- from-motion revisited

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:cbe54676348e17051f49a1d4458fb1b416a636a69297f94da18662025419a452

Observation 312c2b02-444a-400b-b41a-3549b832ddb6 · outbound

This paper cites Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:f9b48672012f6cce62823b3613ad6e2e8621916c0ab4785f8a62c75f74537eff

Observation ba3719b2-0304-4446-a01f-f0336982a128 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Gemini: A Family of Highly Capable Multimodal Models

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:36:27.118161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:feefda04ff0f6d78c092c0e6068813a16005ca72deaabd59ab02c1c864f1e6a4

Observation f9cc5f5f-bdb1-4b91-a9c8-5fdde233f377 · outbound

This paper cites Sosnet: Second order similarity reg- ularization for local descriptor learning.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Sosnet: Second order similarity reg- ularization for local descriptor learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:0dda2489ddf9163fb012a15aab6253a1aa69b42579d4b47f6c3a05320807a02a

Observation af5fd0af-da59-447c-8ac7-9698c2996a0b · outbound

This paper cites Vggt: Vi- sual geometry grounded transformer.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Vggt: Vi- sual geometry grounded transformer

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:671f5276487cb8e7df09f202fc91044f7a9e6c91f0237dd69609bd52cb6957f1

Observation 0368829d-6d55-416b-9f5a-7e427061e7e6 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:36:27.120626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:d5768be3ca5b826853443d0315b19fd48f977642cb8ab47bd88337b355f1c7a2

Observation 5654e44d-c9f7-4a85-8f9f-86b073678946 · outbound

This paper cites Moge: Unlocking accurate monocular geometry estimation for open-domain images with optimal training supervision.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Moge: Unlocking accurate monocular geometry estimation for open-domain images with optimal training supervision

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:d03633973f7c8410517aa40d3ff33ccdc35a54493913510e5d514e7fbca6ab57

Observation 5c1a49a8-7ee5-4981-986b-325be80844d1 · outbound

This paper cites Site: towards spatial intelligence thorough evaluation.arXiv preprint arXiv:2505.05456,.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Site: towards spatial intelligence thorough evaluation.arXiv preprint arXiv:2505.05456,

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.127268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:a79368d98d5902d197a54b56a50b6f4fb8a1055668ea22e21fdd619e4916aaab

Observation e4db9203-f400-4198-ba53-7fb855936f39 · outbound

This paper cites Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:36:27.134981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:94569d6525594aa019dee7b257c25dec50a8c1546bbea7e719ec2df705156e0a

Observation cca67599-480f-462f-8ab7-ef016ee9b372 · outbound

This paper cites V?: Guided visual search as a core mechanism in multimodal llms.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching V?: Guided visual search as a core mechanism in multimodal llms

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:5bb7e03a8f3ffb7b56dac9706dceb5b9b232508ec2dabdbacef458eef33e8333

Observation 8eca3dc7-f60e-4af2-8538-1d3eb5a1c9ea · outbound

This paper cites Realworldqa: Real-world visual question answering benchmark.https://x.ai/news/grok-1.5v, 2024.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Realworldqa: Real-world visual question answering benchmark.https://x.ai/news/grok-1.5v, 2024

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:c890b576ef9c035c787f21bc6b78687e033a7e7fcbd0ad3102cf57acc2a0bd5a

Observation 799da79a-0794-4675-b979-ba254ab09e3c · outbound

This paper cites Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:36:27.138259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:ee74083bb081a6e2838a3ff04672c065e74fbf5144e99af9ddb5e9c46c0c5810

Observation f1317aff-d042-47fd-a45f-484f16b20782 · outbound

This paper cites Thinking in space: How mul- timodal large language models see, remember, and recall spaces.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Thinking in space: How mul- timodal large language models see, remember, and recall spaces

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:ac7284411b0d40118010a9da45966fbded8e5cbb110fe72bdf93818ca3bda86f

Observation ca1804b6-e976-403b-b10d-efdc7dd932ac · outbound

This paper cites Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.114650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:671868e782fc57a203eb4b07fa503f3970241fa24e0623b4e824318c2a6699ed

Observation 3abb23f5-eb12-4e73-9339-8c2afe4e338e · outbound

This paper cites Spa- tial mental modeling from limited views.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Spa- tial mental modeling from limited views

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.109497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:da78d1404fafbce57c514cd5a3252c4c8e18ed515970c031b947c5508aceaa3e

Observation cafdc826-d6d9-48dd-ad2e-719f58afb9cb · outbound

This paper cites Lmms- eval: Reality check on the evaluation of large multimodal models, 2024.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Lmms- eval: Reality check on the evaluation of large multimodal models, 2024

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:7cd7b349e882cc0b6cf4bc9c3703f670691e148a68be521273239def0e440c83

Observation b6d39ebf-6efe-4444-ae15-8c6ff73527a1 · outbound

This paper cites Long Context Transfer from Language to Vision.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Long Context Transfer from Language to Vision

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:36:27.103532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:34e8877b2b7456711e1471bf4b717501f484f0781de330741e72cb11a071ebb1

Observation 9ab7348e-9927-4180-ab56-fee3d6ea73df · outbound

This paper cites MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:36:27.094847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:fb70cf448962d9fe2921793493d6a237d88853232c6e19594643d3e76aaa3b92

Observation 59631231-82f1-48ef-91b3-68a6c3106ccb · outbound

This paper cites Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.106597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:8d6ebe75ec8e031a982cdccacd41bace9647358566949876453c922991a4c91f

Observation 7de3e548-9bd1-463c-9c85-c1f7e271e6a2 · outbound

This paper cites RoboRefer: Towards spatial referring with rea- soning in vision-language models for robotics.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching RoboRefer: Towards spatial referring with rea- soning in vision-language models for robotics

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.091415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:b2b1cb376779015f980197ce28fe7a0f164daf339ace7d87c9408ad743fd9f61

Observation b2c6241f-bb69-42d8-bf3e-1dfea16b0a79 · outbound

This paper cites Stereo magnification: learning view synthesis using multiplane images.ACM Transactions on Graphics (TOG), 37(4):1–12, 2018.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Stereo magnification: learning view synthesis using multiplane images.ACM Transactions on Graphics (TOG), 37(4):1–12, 2018

Reference 63

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:6de534f56a52e39888306bc5564a989f881f9dfc0b439ccef85f8bb897d62e83

Observation 5418be54-f9b9-48a6-8f34-7ddef12584e9 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:36:27.090117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:73661a83cd60e27e884110464681d1d29155c806b58a2fb32503438d2b70637f

Observation 44fc81a2-c9dd-4d62-b898-e68c27f24fca · outbound

This paper cites Sega- gent: Exploring pixel understanding capabilities in mllms by imitating human annotator trajectories.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Sega- gent: Exploring pixel understanding capabilities in mllms by imitating human annotator trajectories

Reference 65

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:05d4ce8d9a477a4490b4e904ec47828163e48988a9ae8e5f71f4faf279d77c1d

Observation 16f9fb93-a9d2-48ee-a297-28f0690b62cd · outbound

This paper cites ACTIVE-o3: Empowering MLLMs with Active Perception via Pure Reinforcement Learning.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching ACTIVE-o3: Empowering MLLMs with Active Perception via Pure Reinforcement Learning

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:36:27.100708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:7064d91689fe29254ba7393acb6479f2ebc2a48655b65e202ae68f4acdeacabb

Observation 1750721d-d4f7-405a-8af5-46f9d6fc334d · outbound

This paper cites 7 – Implementation Details:Complete specifi- cations of data generation pipeline, experimental setup, prompt template, and curriculum progression schedules.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching 7 – Implementation Details:Complete specifi- cations of data generation pipeline, experimental setup, prompt template, and curriculum progression schedules

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:562e6cdc27b960d769f1e2de23519ad9a620f56ea15c4f5d588eb4ec03569504

Observation bb318602-0cdf-4692-8b4a-01ddb380f6b1 · outbound

This paper cites an unresolved cited work.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:7c597f264a2523725cf5ce6b245b6dde52cb17e7db30dcb8609bc993aca6e700

Observation 6132f501-f417-4ee5-9729-7154c0949599 · outbound

This paper cites Supervised Fine-tuning.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Supervised Fine-tuning

Reference 69

Resolution
malformed identifier
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:14b3bebbfef0df89ddceb7b7ef61f69d482a51ed99d3e8e3022570863833122f

Observation 8e7103c3-9d8d-46d9-9624-e453095fb1d1 · outbound

This paper cites None” (F5).ReasonMatch-Bench includes regions that legitimately have no correspondence in the other view. F5 measures whether the model uses the “none.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching None” (F5).ReasonMatch-Bench includes regions that legitimately have no correspondence in the other view. F5 measures whether the model uses the “none

Reference 70

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:d8f21ec5ba7bf3d09eedd75ddd5b0d53337d8922355da73002d50cb461e33585

Observation b4e87068-946b-4661-84d3-50b28635c487 · outbound

This paper cites 1": "1",.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching 1": "1",

Reference 71

Resolution
unresolved
no resolver link, observed 2026-06-28T10:48:24.523702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:01a5845489b210fe388b91064a187b9a7074a2ff31315f9740d8b0f02c04f6bc

Pith citing papers

No inbound Pith citation observations are available.