Pith. sign in

Paper Citation Record · LEDGER

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images

As of 23 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2511.19119.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.19119 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:38:58.579422Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved62
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6171e5a5-9779-4182-b1c1-21820f9db851 · outbound

This paper cites an unresolved cited work.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:54.657295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:54.657295Z digest=sha256:1a796cacd9fab91ea10e7fcccb3c7c9395dfb950f90fb394bb208aff8de211ca

Observation 94e10a9f-6b82-451b-a339-8de016a6fb1e · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Scanqa: 3d question answering for spatial scene understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:54.706106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:54.706106Z digest=sha256:0a1ab19dbee97391d963c55db0e1a5f801c03c72e6d72b8c31aa95e1eed13a67

Observation 12207c0e-8757-48fe-8568-fd461fa8b6be · outbound

This paper cites Qwen2.5-VL Technical Report.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:54.781179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:54.781179Z digest=sha256:ff45d4c8988a91e3c2faff413633c70c7c181f4f8983506d63f8e9193bb06de8

Observation 710d92b2-ac15-4abf-bc0b-12b491aa1962 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:54.831472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:54.831472Z digest=sha256:d8521eaca6a1b0461d99e7339c54a4607b43952523eb18fe5104aec7f405ef2e

Observation 0fd9d04f-37b2-42d8-a1e5-62fe5678f162 · outbound

This paper cites Omni3d: A large benchmark and model for 3d object detection in the wild.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Omni3d: A large benchmark and model for 3d object detection in the wild

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:54.864656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:54.864656Z digest=sha256:3016d1ee7f1b819eaae2c9991656b192b8e82d5fe8127eb437d85fcc6ad07b14

Observation 89fc6ba7-bae3-419e-9de8-09a4738a8d4a · outbound

This paper cites SpatialBot: Precise Spatial Understanding with Vision Language Models.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:54.925300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:54.925300Z digest=sha256:9483282b7aa7e407f1fbafc97fe30087128899d21fd42e20a61b659ea963d17c

Observation 1c086a39-011b-49a0-9149-adc2640e85c3 · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:54.954815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:54.954815Z digest=sha256:c1c03436b77e5043e7e2ca55dfb188313f896c5b081e70a86a8424c7a054056a

Observation 6c6b597c-a254-49e1-ab03-f7a18a027b21 · outbound

This paper cites Perception before reasoning: Two-stage reinforce- ment learning for visual reasoning in vision-language mod- els.arXiv preprint arXiv:2509.13031, 2025.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Perception before reasoning: Two-stage reinforce- ment learning for visual reasoning in vision-language mod- els.arXiv preprint arXiv:2509.13031, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:54.974631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:54.974631Z digest=sha256:1fe2b611c301b110501291f4fe1668e58b6a59aa984bd0fb35af4a8034e85935

Observation a8942ee0-7dbf-45b9-ac42-d7d87e28b24f · outbound

This paper cites Spatial- rgpt: Grounded spatial reasoning in vision-language models.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Spatial- rgpt: Grounded spatial reasoning in vision-language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.015330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.015330Z digest=sha256:070c8b9a18b4c28fbb3c2d3569e6de7d5e4480315982c8b3a894e81d1fd957b4

Observation 508972a0-ed4b-495d-ad5b-d2c5e2b04dab · outbound

This paper cites Collins, Ilia Sucholutsky, Umang Bhatt, Kartik Chandra, Lionel Wong, Mina Lee, Cedegao E.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Collins, Ilia Sucholutsky, Umang Bhatt, Kartik Chandra, Lionel Wong, Mina Lee, Cedegao E

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.028226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.028226Z digest=sha256:15111afb04ce2ae4df0e4c4f1d1d92007fd1c8e82c22dffc21986d371e23839f

Observation e73763e9-b8bb-4338-ad19-4d2eadd1c8f7 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.094974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.094974Z digest=sha256:e49e4d7b0c98bee44ee80714c08e8d70c6e4613356463c52b7f8a395e6150dfa

Observation 72fc96a9-1852-4a86-a0bd-ed1dcd958dde · outbound

This paper cites VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.151754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.151754Z digest=sha256:03decf879257293f4bc21e84c56b810a1c5308c1ab4227e4dd706a5a774b6d31

Observation eee8f01c-22c9-4336-b728-07c975bf8000 · outbound

This paper cites Surds: Benchmarking spatial understand- ing and reasoning in driving scenarios with vision language models.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Surds: Benchmarking spatial understand- ing and reasoning in driving scenarios with vision language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.184590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.184590Z digest=sha256:f00e376f99d7898d5c4cfe54705c2936f7309fbb72dc703a0f39d1837efe4a3e

Observation ab6ab7ac-37fd-4510-8642-cfb12d71f050 · outbound

This paper cites 3d-llm: Inject- ing the 3d world into large language models.NeurIPS, 2023.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images 3d-llm: Inject- ing the 3d world into large language models.NeurIPS, 2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.239896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.239896Z digest=sha256:b7cce4c5917eb6c3c3f7b01ff1c09a17d01cde08b6d0f7debcd94da0622b90f6

Observation 909bc398-84c5-4dd4-b423-b1ccc8b10f08 · outbound

This paper cites What’s ”up” with vision-language models? investigating their strug- gle with spatial reasoning.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images What’s ”up” with vision-language models? investigating their strug- gle with spatial reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.287693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.287693Z digest=sha256:0cc62ab5c5bc78c30326d572c8fcd5457ed46c014d9f6a7c166919698ad31413

Observation 346095bd-61b2-48d3-81e5-e9db8400d8db · outbound

This paper cites Lee, Jihyeon Je, Chanho Park, Mikaela Angelina Uy, Leonidas Guibas, and Minhyuk Sung.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Lee, Jihyeon Je, Chanho Park, Mikaela Angelina Uy, Leonidas Guibas, and Minhyuk Sung

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.318595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.318595Z digest=sha256:0eefecff215393d76c84420fb51cc9c7bd486f4b361dc051dc59d8819d35c030

Observation 532c478c-b7da-43c5-85d3-6ad7c014b912 · outbound

This paper cites Seed-bench: Bench- marking multimodal large language models.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Seed-bench: Bench- marking multimodal large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.361460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.361460Z digest=sha256:d00521b6b4535b1bd178b16fc6a8e201f2cce28b49a0ea3737734633c984d84f

Observation cee1622d-630d-4358-9580-72b51b92f0c4 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images LLaVA-OneVision: Easy Visual Task Transfer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.411826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.411826Z digest=sha256:dd250a95b27bdaf98dc7080a71d816b4932ab605a83594a0decc7ce543e30652

Observation 95c1dadc-6fda-4c81-8d38-d9f20f4b0ee3 · outbound

This paper cites Viewspatial-bench: Evaluating multi-perspective spatial localization in vision-language models, 2025.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Viewspatial-bench: Evaluating multi-perspective spatial localization in vision-language models, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.448060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.448060Z digest=sha256:1905b504f00954174f776144022054f61b0ac9617f3df63841f2325be149fac1

Observation cd0f1fc6-39d2-4a64-b2c2-0c78d649dc57 · outbound

This paper cites Visual Instruction Tuning.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Visual Instruction Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.503949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.503949Z digest=sha256:2d034a2d7277b682dce65f808961a358042718415d1ffce0ff869f5d2994b1a0

Observation 8761fb1f-1515-494e-9010-dd67e17c3095 · outbound

This paper cites Spatialladder: Progressive train- ing for spatial reasoning in vision-language models, 2025.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Spatialladder: Progressive train- ing for spatial reasoning in vision-language models, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.545210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.545210Z digest=sha256:db7539ffa96b018a1fb348043ee0a3e9f420c826a37c092945401e931bc369ec

Observation f54e9786-4ca9-4ced-b037-d896d0b231c2 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.596804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.596804Z digest=sha256:1e620584c341b8138b0527240a418bba2cac2bf42a09b88eceb8add060e044f4

Observation a2d028f0-fe31-4c10-8379-6933e4834c96 · outbound

This paper cites an unresolved cited work.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.643951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.643951Z digest=sha256:e5a377aaa662e5cabf9664067213d518629b51779e47c7e6a382dfa0dfb34713

Observation 65d6b3ed-76fc-44a4-bc01-4fb57d4847a8 · outbound

This paper cites A Novel Multi-Agent Deep RL Approach for Traffic Signal Control.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images A Novel Multi-Agent Deep RL Approach for Traffic Signal Control

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.718590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.718590Z digest=sha256:845453f6e3d54047463c5368d8adef629406f5ff5791eea5ce589a599c924510

Observation 7def5476-984c-4f25-ab84-09af786d2853 · outbound

This paper cites Visual spa- tial reasoning.Transactions of the Association for Computa- tional Linguistics, 11:635–651, 2023.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Visual spa- tial reasoning.Transactions of the Association for Computa- tional Linguistics, 11:635–651, 2023

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.770035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.770035Z digest=sha256:0e3ce4fed09cf1650930d0aea221cd7e4c2219edd595d6a7ff9a1d3e865ffeb2

Observation e9a6c7f1-deac-4565-833b-e8f48d1b5405 · outbound

This paper cites Grounding dino: Mar- rying dino with grounded pre-training for open-set object 10 detection.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Grounding dino: Mar- rying dino with grounded pre-training for open-set object 10 detection

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.855026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.855026Z digest=sha256:aa58e47653795862af0afda82c851d50b984c8f681c033929f5fdd846c895d1a

Observation 06e758fc-b272-4fee-b172-b22316866e94 · outbound

This paper cites 3dsrbench: A compre- hensive 3d spatial reasoning benchmark.arXiv preprint arXiv:2412.07825, 2024.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images 3dsrbench: A compre- hensive 3d spatial reasoning benchmark.arXiv preprint arXiv:2412.07825, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.901480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.901480Z digest=sha256:3437bb80b1ed99d8ae1080e01c65cdd6ea2afcc78110478c4c6d868f79385c6c

Observation 809fd4d9-e63b-4352-8c61-e370e45e891c · outbound

This paper cites Sqa3d: Situated question answering in 3d scenes.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Sqa3d: Situated question answering in 3d scenes

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.967821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.967821Z digest=sha256:35b84e14448df1e3cc3a608ce195930eafd78e2572ba2fddbecc2f4ea20896d6

Observation f5c72476-ad3f-498d-b895-80649dd2dc72 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.020436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.020436Z digest=sha256:b9844532c3d170143375e1469ef015881915f540c6854b69334f9649d229e58b

Observation 21436383-2142-4cfc-a7ad-c80fe38dfc3b · outbound

This paper cites GPT-4 Technical Report.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images GPT-4 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.091709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.091709Z digest=sha256:0f22b7eb885b8a8a86195741ca0e642334f6df74a7e6044dea2b9f753616dc26

Observation d48f1012-52dd-4854-80a4-4ad5e0383e39 · outbound

This paper cites Shapellm: Universal 3d object understanding for embodied interaction.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Shapellm: Universal 3d object understanding for embodied interaction

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.095771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.095771Z digest=sha256:feba7e8884619d6bc0a95c21f29b248735c7efd68544e36df0ca3a0ff854e55a

Observation 50e00068-8775-49a4-9a94-2dbaf0723a35 · outbound

This paper cites Learning transferable visual models from natural language supervision.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Learning transferable visual models from natural language supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.102240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.102240Z digest=sha256:7774ef646edbc69ddb6d587526a855bbe0553661b4c68f9415d1971c94783779

Observation 1e11018a-2d3c-4ed3-8394-c1f711d24bc9 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.157395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.157395Z digest=sha256:f9c0163b562f9f525ea0bad7bdc09246dd808f6da8f68b324256c5de0a29c8d5

Observation ec6506ca-7cad-48d3-9bda-080642b9689e · outbound

This paper cites Space3D-Bench: Spatial 3D Question Answering Benchmark.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Space3D-Bench: Spatial 3D Question Answering Benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.248164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.248164Z digest=sha256:3b5045cceb5ee39e80a3258635f15996619dd20f5d7c30e98d4d030780b67443

Observation b37d4bd3-ca58-4b9d-8dc0-095031994996 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Gemini: A Family of Highly Capable Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.274268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.274268Z digest=sha256:20c6dbc5a17256da70d32b2ec7e75780ba670f4d56df6926bf173e18a35e1f49

Observation 0ed634ff-f6e0-4037-8e01-a0acee736260 · outbound

This paper cites NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.380114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.380114Z digest=sha256:f166b57f401dbb517bf0386be77c36e0c62f1884fa6e05864bd4a3e3a61fa310

Observation e059e4b6-49a5-44ef-bc97-727183fc8bdf · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images LLaMA: Open and Efficient Foundation Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.472453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.472453Z digest=sha256:bc15454d50f9ec8320758751599d93e4308e10a6135359354a9403c94116d0cc

Observation df08ed63-aad8-4ade-8dee-f858468399a6 · outbound

This paper cites Cross-modal pro- jection in multimodal llms doesn’t really project visual at- tributes to textual space.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Cross-modal pro- jection in multimodal llms doesn’t really project visual at- tributes to textual space

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.591800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.591800Z digest=sha256:54cc1f5e97ebd52cd15a9605bfec7a3cf5815ff0bae9a6c24bcbdf1cb1b3e93d

Observation 2e583bfd-8301-48ac-8aa6-092681b606bb · outbound

This paper cites Learning 3d semantic scene graphs from 3d indoor reconstructions.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Learning 3d semantic scene graphs from 3d indoor reconstructions

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.673894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.673894Z digest=sha256:5aefc7fbbe91b2585a59ae6f7be91038933805eee5119b1b66ba37e37b47225a

Observation a4a5cb51-d025-44e9-b5f3-a00c59a9361b · outbound

This paper cites Is a picture worth a thou- sand words? delving into spatial reasoning for vision lan- guage models.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Is a picture worth a thou- sand words? delving into spatial reasoning for vision lan- guage models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.735634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.735634Z digest=sha256:56778e674e91e8f4a4457d11fd48ec1e0e2245af1573992d650c35127eadd618

Observation 7acbe468-8842-49d1-99d4-ca4dfe6618ae · outbound

This paper cites Spatial 3d-llm: Exploring spatial awareness in 3d vision-language models,.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Spatial 3d-llm: Exploring spatial awareness in 3d vision-language models,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.777803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.777803Z digest=sha256:95945372eaeeb7767548dc052a5e07c7bb9f577e236216f40f6389397818eb78

Observation 04f4922f-24b2-44f6-bfe3-9389f52a9375 · outbound

This paper cites Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.844553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.844553Z digest=sha256:e24fa708beb5a9d841259061bab9325ff3c7633dc91fdf5edbf4751d733e0af7

Observation 0373d929-38c4-4c28-a8b6-d3554f1e5416 · outbound

This paper cites Pointllm: Empowering large lan- guage models to understand point clouds.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Pointllm: Empowering large lan- guage models to understand point clouds

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.910137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.910137Z digest=sha256:3e2012ee3870112b1e5e57813fcb9041b8156d9c802eeea2a5d557d298ee78a7

Observation 45090bbc-c81e-4083-83e5-efdd44782115 · outbound

This paper cites Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.976381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.976381Z digest=sha256:9e70583560875b756bea3b758c94ccab5818a703b132d77ab9fc45c0de7b34aa

Observation 4ec212f5-83fa-4cbb-89f0-582422b9b7eb · outbound

This paper cites Open-vocabulary object detection using cap- tions.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Open-vocabulary object detection using cap- tions

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.124463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.124463Z digest=sha256:b43619c96469571e9860d6fadf72933636b2ba0c16f88552f45f990897d53994

Observation 09b0e65d-2a0c-49e7-8238-0a0cd844b2c7 · outbound

This paper cites How to enable llm with 3d capacity? a survey of spatial reasoning in llm, 2025.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images How to enable llm with 3d capacity? a survey of spatial reasoning in llm, 2025

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.258520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.258520Z digest=sha256:69d945042c5e20288993cb80719e853aaf49d9a1c336bfd85b840507d6fa6995

Observation 1d04de49-7abf-40cc-ad8c-ccefd429fd3c · outbound

This paper cites From flatland to space: Teaching vision-language models to perceive and reason in 3d.arXiv preprint arXiv:2503.22976, 2025.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images From flatland to space: Teaching vision-language models to perceive and reason in 3d.arXiv preprint arXiv:2503.22976, 2025

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.314772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.314772Z digest=sha256:bd2626bbd730d7a7191f9b5dd547bcc00ab733f89ac28bb65f0e18c494653976

Observation 904bbdd2-6d8e-42bc-9365-6868345a91a9 · outbound

This paper cites Spinbench: Perspective and rotation as a lens on spatial reasoning in vlms, 2025.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Spinbench: Perspective and rotation as a lens on spatial reasoning in vlms, 2025

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.422175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.422175Z digest=sha256:daadc076f19025a88c9d7509c30deac70531d3265111f8c91eebe776ea0f1449

Observation 2a6ba5e5-489d-4e13-b4e2-6678b7a104cf · outbound

This paper cites Video-3d llm: Learning position-aware video representation for 3d scene understanding.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Video-3d llm: Learning position-aware video representation for 3d scene understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.553330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.553330Z digest=sha256:d39e456d84438d0c1a9eac95e33dc6d83bc5628876b5b2d0e5d88d57815f54c9

Observation b94a55ac-cab6-4579-8625-16871717526a · outbound

This paper cites Scanreason: Empowering 3d visual grounding with reasoning capabilities.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Scanreason: Empowering 3d visual grounding with reasoning capabilities

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.596020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.596020Z digest=sha256:c7905d6979c0bb0abdc673f2f6c36fb2c9adeda931d346a11ebbb46c1fb4e93e

Observation f65d7cbe-43bc-4dbc-a2d9-6e2a7126a51f · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.652883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.652883Z digest=sha256:fc6162589cd475df7cacffb944c7aaa3d6993fccc516fcc9955087288117b592

Observation 9d293d6a-d7db-4986-a1e2-6b7bd44fcead · outbound

This paper cites A detailed description of each level and its cor- responding tasks is provided below.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images A detailed description of each level and its cor- responding tasks is provided below

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.761326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.761326Z digest=sha256:cdad3b48120ff18f5196854321fabe264ce230774fa4567370335b4f30df0386

Observation ff698a18-5d01-4deb-9db5-359f450867ae · outbound

This paper cites Specifically, after obtaining high-quality raw data through filtering, we generate image captions and construct scene graphs to serve as our underlying database.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Specifically, after obtaining high-quality raw data through filtering, we generate image captions and construct scene graphs to serve as our underlying database

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.853194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.853194Z digest=sha256:83f9724d78d19d7a39b85d9562d6d2fcf0efcfe560614a6e1bb1cfd47c40626c

Observation f1e64af2-c666-458b-a4b7-1e450b6963cc · outbound

This paper cites - {variation_instruction} - The rewritten question MUST include: (a) A brief motivation clause describing WHY we need this information, consistent with the motivation hint.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images - {variation_instruction} - The rewritten question MUST include: (a) A brief motivation clause describing WHY we need this information, consistent with the motivation hint

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.909081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.909081Z digest=sha256:5527ff2a5b47996fcb748d1f00dae8b15276fbde81c6bb3fb516915045966dd5

Observation 591aa810-2306-491a-b5ac-fbdda1f876b0 · outbound

This paper cites - {answer_constraint} - {task_extra} - You MUST NOT flip yesno.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images - {answer_constraint} - {task_extra} - You MUST NOT flip yesno

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.975775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.975775Z digest=sha256:6a7e48c5eebd3af73e335016ad3bed2b15eb26c38abbceb5fe1838204734c020

Observation 39a5c178-0457-4ead-b97f-b774b381fd20 · outbound

This paper cites thinking.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images thinking

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:58.030558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:58.030558Z digest=sha256:86d15fcca8ea8dfc1a4ff63786d7315ae4579c7835f70d239899515bd5903719

Observation 7ab6194a-8858-4083-bc70-cf1d59d2870e · outbound

This paper cites Scene In- formation adds global context, 2D Visual Prompts improve local grounding, and 3D Bounding Boxes deliver the largest gains through explicit geometric structure.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Scene In- formation adds global context, 2D Visual Prompts improve local grounding, and 3D Bounding Boxes deliver the largest gains through explicit geometric structure

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:58.090253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:58.090253Z digest=sha256:45fd47bb1da713b1742312cd0bf815ddd3c12022a225aa8389025ec9229b5534

Observation b027e5a3-9197-4b17-93b2-27c3c38441ef · outbound

This paper cites For the 3D bounding box information, each object is rep- resented by its center coordinates, spatial dimensions (size), and orientation expressed as a rotation quaternion.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images For the 3D bounding box information, each object is rep- resented by its center coordinates, spatial dimensions (size), and orientation expressed as a rotation quaternion

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:58.171986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:58.171986Z digest=sha256:f442d5ba52c84255335abed350466dee2bf091a2809b1ac359dd24d9e146ea3a

Observation d6e8e8b5-5ad3-4d78-8169-69fca60a2b03 · outbound

This paper cites All inputs are processed with the officialQwen2.5-VLprocessor, which supports dy- namic image resolutions up to 262,144 pixels (512×512).

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images All inputs are processed with the officialQwen2.5-VLprocessor, which supports dy- namic image resolutions up to 262,144 pixels (512×512)

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:58.262429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:58.262429Z digest=sha256:a3e61d6878ca72c0675025546bd5035e76f5642e2ca3e20cd5d70b598a7d34f8

Observation 41a9de49-fc95-4dfc-8e96-9391fe0f4693 · outbound

This paper cites - Use the format <obj>...</obj> to describe your mapping between textual entities and object IDs.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images - Use the format <obj>...</obj> to describe your mapping between textual entities and object IDs

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:58.358583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:58.358583Z digest=sha256:ef870727a11bb031dfd661931d4dbae5b4d1269efc62886a9906cf44f846fce5

Observation bb5c9776-80e7-4524-8b8f-cb623da6ca9e · outbound

This paper cites an unresolved cited work.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:58.429529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:58.429529Z digest=sha256:8200b6e503ab7440a3a04b47bb0d25c02f7a0abd086deb42defa40fa426246b3

Observation 5a1117f2-1d56-48f4-9935-77abe67ff869 · outbound

This paper cites Your output format should strictly follow: <obj> Object mapping: - entity_1 object {id_a} ({caption_a}) - entity_2 object {id_b} </obj> <reasoning>.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Your output format should strictly follow: <obj> Object mapping: - entity_1 object {id_a} ({caption_a}) - entity_2 object {id_b} </obj> <reasoning>

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:58.579422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:58.579422Z digest=sha256:83685d55f75fa37de0824a2f004ee344e1a1a7166214abe530dde3b22f1ec3e8

Pith citing papers

No inbound Pith citation observations are available.