Pith. sign in

Paper Citation Record · LEDGER

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images

As of 8 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2511.19119.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.19119 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:38:58.579422Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved62
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6171e5a5-9779-4182-b1c1-21820f9db851 · outbound

This paper cites an unresolved cited work.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:54.657295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:54.657295Z digest=sha256:9a191a38e97fd2ddb5308d2f31597bbf2d92a0a89f4a5ad7ea1200f246a8fda9

Observation 94e10a9f-6b82-451b-a339-8de016a6fb1e · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Scanqa: 3d question answering for spatial scene understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:54.706106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:54.706106Z digest=sha256:ee1e1de556352fd3105e7f72529e751e95b211f9273e0a14a25e545419f183ad

Observation 12207c0e-8757-48fe-8568-fd461fa8b6be · outbound

This paper cites Qwen2.5-VL Technical Report.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:54.781179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:54.781179Z digest=sha256:b031c79a6a117a094e1cdd16f224a05a6d41693398d6f53f2da79fb3373c2f3e

Observation 710d92b2-ac15-4abf-bc0b-12b491aa1962 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:54.831472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:54.831472Z digest=sha256:55fc4501cd99d8064b476e4f1105621322db295a5517b16510c3090be7809c52

Observation 0fd9d04f-37b2-42d8-a1e5-62fe5678f162 · outbound

This paper cites Omni3d: A large benchmark and model for 3d object detection in the wild.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Omni3d: A large benchmark and model for 3d object detection in the wild

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:54.864656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:54.864656Z digest=sha256:3ed5f9fd926e8aadac6ce43e8ed6692709a4398fb1e17ca86d7f774e3e47db66

Observation 89fc6ba7-bae3-419e-9de8-09a4738a8d4a · outbound

This paper cites SpatialBot: Precise Spatial Understanding with Vision Language Models.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:54.925300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:54.925300Z digest=sha256:bd2e8c5c5b9d7c9d871a5da09b3add7a94c8a7523ec9966310790eef782ee099

Observation 1c086a39-011b-49a0-9149-adc2640e85c3 · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:54.954815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:54.954815Z digest=sha256:6b422cfc6e382f61b20f51a76778866c2f3e5ce222cf87ac01795ce33cab51b3

Observation 6c6b597c-a254-49e1-ab03-f7a18a027b21 · outbound

This paper cites Perception before reasoning: Two-stage reinforce- ment learning for visual reasoning in vision-language mod- els.arXiv preprint arXiv:2509.13031, 2025.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Perception before reasoning: Two-stage reinforce- ment learning for visual reasoning in vision-language mod- els.arXiv preprint arXiv:2509.13031, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:54.974631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:54.974631Z digest=sha256:f64c34cb4f9cdb7e16fe48a4693aa9fdf3287c014ca5992af20fe4e60d9ace10

Observation a8942ee0-7dbf-45b9-ac42-d7d87e28b24f · outbound

This paper cites Spatial- rgpt: Grounded spatial reasoning in vision-language models.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Spatial- rgpt: Grounded spatial reasoning in vision-language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.015330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.015330Z digest=sha256:6bb0047c389cf61d8964220ebc9af3138a36d5c4821df8564186a7037b665748

Observation 508972a0-ed4b-495d-ad5b-d2c5e2b04dab · outbound

This paper cites Collins, Ilia Sucholutsky, Umang Bhatt, Kartik Chandra, Lionel Wong, Mina Lee, Cedegao E.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Collins, Ilia Sucholutsky, Umang Bhatt, Kartik Chandra, Lionel Wong, Mina Lee, Cedegao E

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.028226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.028226Z digest=sha256:2ab9cd1274858ad0575b08536906e3ca13a69bd91f0faa5194e25d71a0014065

Observation e73763e9-b8bb-4338-ad19-4d2eadd1c8f7 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.094974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.094974Z digest=sha256:ed83dd44278e2b3fc56939a90dceb6515e44686588970601a06c8edb8de1717e

Observation 72fc96a9-1852-4a86-a0bd-ed1dcd958dde · outbound

This paper cites VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.151754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.151754Z digest=sha256:e7e2ca733da5e288a0d50291b003f3ce225b84a3c49c3629b385ec2ef37e515a

Observation eee8f01c-22c9-4336-b728-07c975bf8000 · outbound

This paper cites Surds: Benchmarking spatial understand- ing and reasoning in driving scenarios with vision language models.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Surds: Benchmarking spatial understand- ing and reasoning in driving scenarios with vision language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.184590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.184590Z digest=sha256:a5fc82bbc327ecb5acf9036df62f67515d1e4cb264b7e174ca962da15a54418a

Observation ab6ab7ac-37fd-4510-8642-cfb12d71f050 · outbound

This paper cites 3d-llm: Inject- ing the 3d world into large language models.NeurIPS, 2023.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images 3d-llm: Inject- ing the 3d world into large language models.NeurIPS, 2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.239896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.239896Z digest=sha256:80b3a37aa427c70954835643ef7501a81a48ff81c938c4efeae33186451c28e4

Observation 909bc398-84c5-4dd4-b423-b1ccc8b10f08 · outbound

This paper cites What’s ”up” with vision-language models? investigating their strug- gle with spatial reasoning.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images What’s ”up” with vision-language models? investigating their strug- gle with spatial reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.287693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.287693Z digest=sha256:271901b5f36ecca3ef6e892fe82c98695719872e1942e295482b5a96405b7b86

Observation 346095bd-61b2-48d3-81e5-e9db8400d8db · outbound

This paper cites Lee, Jihyeon Je, Chanho Park, Mikaela Angelina Uy, Leonidas Guibas, and Minhyuk Sung.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Lee, Jihyeon Je, Chanho Park, Mikaela Angelina Uy, Leonidas Guibas, and Minhyuk Sung

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.318595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.318595Z digest=sha256:c39faf971a2f446c90164b6c09776cdb21d7fb39780de2d1cd31343b78993bbe

Observation 532c478c-b7da-43c5-85d3-6ad7c014b912 · outbound

This paper cites Seed-bench: Bench- marking multimodal large language models.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Seed-bench: Bench- marking multimodal large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.361460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.361460Z digest=sha256:64170402d7f9104e87c7657446e2ddb0ff1931e9d07fb723bee5d60dfd06abe6

Observation cee1622d-630d-4358-9580-72b51b92f0c4 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images LLaVA-OneVision: Easy Visual Task Transfer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.411826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.411826Z digest=sha256:c2f8a34045d4fb9b3bc216537b8d2a93ec6ddf1c8e8408baa4b3acca41374f5a

Observation 95c1dadc-6fda-4c81-8d38-d9f20f4b0ee3 · outbound

This paper cites Viewspatial-bench: Evaluating multi-perspective spatial localization in vision-language models, 2025.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Viewspatial-bench: Evaluating multi-perspective spatial localization in vision-language models, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.448060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.448060Z digest=sha256:cb1bbce3b2dfa5481d3638797fd1ea02826def52d152e846f6fcd9657aa01cc1

Observation cd0f1fc6-39d2-4a64-b2c2-0c78d649dc57 · outbound

This paper cites Visual Instruction Tuning.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Visual Instruction Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.503949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.503949Z digest=sha256:8eb654eaec190e9cff5fa17cd1eb8877de043ce1ca09ea53d18bc253561004a4

Observation 8761fb1f-1515-494e-9010-dd67e17c3095 · outbound

This paper cites Spatialladder: Progressive train- ing for spatial reasoning in vision-language models, 2025.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Spatialladder: Progressive train- ing for spatial reasoning in vision-language models, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.545210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.545210Z digest=sha256:a3c07a439af23a0d4dd06ddfc6d2a2f39b32a46607f58e8457ca0123c302f694

Observation f54e9786-4ca9-4ced-b037-d896d0b231c2 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.596804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.596804Z digest=sha256:e3279728d797d9ff524bd3c6839593167fe2773a66ca2cafd7161bf6e9fc7375

Observation a2d028f0-fe31-4c10-8379-6933e4834c96 · outbound

This paper cites an unresolved cited work.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.643951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.643951Z digest=sha256:8550420aa00445a7d5d849230f443004f2c4945160bb6e85c30020aa3d778f30

Observation 65d6b3ed-76fc-44a4-bc01-4fb57d4847a8 · outbound

This paper cites A Novel Multi-Agent Deep RL Approach for Traffic Signal Control.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images A Novel Multi-Agent Deep RL Approach for Traffic Signal Control

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.718590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.718590Z digest=sha256:3f87d00d7ead93f86acf7d1597d93d6436f34bdb3649bc1f6f8c190002567079

Observation 7def5476-984c-4f25-ab84-09af786d2853 · outbound

This paper cites Visual spa- tial reasoning.Transactions of the Association for Computa- tional Linguistics, 11:635–651, 2023.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Visual spa- tial reasoning.Transactions of the Association for Computa- tional Linguistics, 11:635–651, 2023

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.770035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.770035Z digest=sha256:84ec595daa90dcf049c0e231bcf60ec4f71cb6aa86898aafa4fefb5b5dfb72c7

Observation e9a6c7f1-deac-4565-833b-e8f48d1b5405 · outbound

This paper cites Grounding dino: Mar- rying dino with grounded pre-training for open-set object 10 detection.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Grounding dino: Mar- rying dino with grounded pre-training for open-set object 10 detection

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.855026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.855026Z digest=sha256:5ea292d70f837aa97665943fceb7aa21db2584a205100f48183048dd6567274d

Observation 06e758fc-b272-4fee-b172-b22316866e94 · outbound

This paper cites 3dsrbench: A compre- hensive 3d spatial reasoning benchmark.arXiv preprint arXiv:2412.07825, 2024.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images 3dsrbench: A compre- hensive 3d spatial reasoning benchmark.arXiv preprint arXiv:2412.07825, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.901480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.901480Z digest=sha256:cda01c4a91f561de333fdc7cf0f09d3c9168356f985d5154190afb6bd184a345

Observation 809fd4d9-e63b-4352-8c61-e370e45e891c · outbound

This paper cites Sqa3d: Situated question answering in 3d scenes.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Sqa3d: Situated question answering in 3d scenes

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.967821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.967821Z digest=sha256:18dcb011782c2108f878c7c15a0e55ef284099d3be252ff44066e95f918623ea

Observation f5c72476-ad3f-498d-b895-80649dd2dc72 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.020436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.020436Z digest=sha256:1d8a86b422a8bfa9d1ce5a66581557b60d19a05d89b2635b04952ac86c303f64

Observation 21436383-2142-4cfc-a7ad-c80fe38dfc3b · outbound

This paper cites GPT-4 Technical Report.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images GPT-4 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.091709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.091709Z digest=sha256:bd2e22f7409f04fbb1e7dd2d1915b790b14a36c909ece27908026199f5fed61e

Observation d48f1012-52dd-4854-80a4-4ad5e0383e39 · outbound

This paper cites Shapellm: Universal 3d object understanding for embodied interaction.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Shapellm: Universal 3d object understanding for embodied interaction

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.095771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.095771Z digest=sha256:9c05d7a94c11957de9d18ee66ec7c9ac432d059b05737dd65e5472d77de94eda

Observation 50e00068-8775-49a4-9a94-2dbaf0723a35 · outbound

This paper cites Learning transferable visual models from natural language supervision.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Learning transferable visual models from natural language supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.102240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.102240Z digest=sha256:d9f531959d4418792616c8cbcffbb7816f1600eb8fb5fbfc04b858dec3c27c77

Observation 1e11018a-2d3c-4ed3-8394-c1f711d24bc9 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.157395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.157395Z digest=sha256:801d7c7bd8f4af6c6b7b4fccfee1cc4c5abcd1a6db8cfc8274a09aa9b42b0458

Observation ec6506ca-7cad-48d3-9bda-080642b9689e · outbound

This paper cites Space3D-Bench: Spatial 3D Question Answering Benchmark.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Space3D-Bench: Spatial 3D Question Answering Benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.248164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.248164Z digest=sha256:2b270a47e43351663984b74db7e0deab8b0d177021de4e33c254c6f635fff7c8

Observation b37d4bd3-ca58-4b9d-8dc0-095031994996 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Gemini: A Family of Highly Capable Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.274268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.274268Z digest=sha256:7ca95451e8e3ee25835b0ef7d824e33efa84c00afb6a32f4102dc75c850c1408

Observation 0ed634ff-f6e0-4037-8e01-a0acee736260 · outbound

This paper cites NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.380114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.380114Z digest=sha256:e9325f0dbd01daf12f5457a563c42844853e73e700cc032e51e8242cf94cb034

Observation e059e4b6-49a5-44ef-bc97-727183fc8bdf · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images LLaMA: Open and Efficient Foundation Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.472453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.472453Z digest=sha256:60f0d2b07bc9766352e50f9e612d28286e2f851a819410c49bc79d0e7bf83e39

Observation df08ed63-aad8-4ade-8dee-f858468399a6 · outbound

This paper cites Cross-modal pro- jection in multimodal llms doesn’t really project visual at- tributes to textual space.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Cross-modal pro- jection in multimodal llms doesn’t really project visual at- tributes to textual space

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.591800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.591800Z digest=sha256:b6bd05cd79e75465b455bdbfe684caff8532601d4cec5096d78cf619161094f4

Observation 2e583bfd-8301-48ac-8aa6-092681b606bb · outbound

This paper cites Learning 3d semantic scene graphs from 3d indoor reconstructions.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Learning 3d semantic scene graphs from 3d indoor reconstructions

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.673894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.673894Z digest=sha256:564e1b93128d0d5f5d95733d4196bc8f8e7a5ab41e8d85319513b55148eabf99

Observation a4a5cb51-d025-44e9-b5f3-a00c59a9361b · outbound

This paper cites Is a picture worth a thou- sand words? delving into spatial reasoning for vision lan- guage models.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Is a picture worth a thou- sand words? delving into spatial reasoning for vision lan- guage models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.735634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.735634Z digest=sha256:d35e10cf9d0dedf019163b88d55083126467876d2d2a0af61954b6a39ecf56d5

Observation 7acbe468-8842-49d1-99d4-ca4dfe6618ae · outbound

This paper cites Spatial 3d-llm: Exploring spatial awareness in 3d vision-language models,.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Spatial 3d-llm: Exploring spatial awareness in 3d vision-language models,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.777803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.777803Z digest=sha256:2e094d684770dbc669f9d6b521113249127c15512a15f779440e9b369ab80fc4

Observation 04f4922f-24b2-44f6-bfe3-9389f52a9375 · outbound

This paper cites Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.844553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.844553Z digest=sha256:e1476c74f5319f36189963389a9ffc6dea0544a6adec1174e911546e8341e6ae

Observation 0373d929-38c4-4c28-a8b6-d3554f1e5416 · outbound

This paper cites Pointllm: Empowering large lan- guage models to understand point clouds.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Pointllm: Empowering large lan- guage models to understand point clouds

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.910137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.910137Z digest=sha256:ae82144abc058d6617c5e3cfd34681f25cb1ef3cfc5c011d4fce76e4cdb7aa79

Observation 45090bbc-c81e-4083-83e5-efdd44782115 · outbound

This paper cites Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.976381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.976381Z digest=sha256:645c5e0c49463782bfa5e3aad1d78c8d6c23fc261cfd8d285be73d3aa1497fda

Observation 4ec212f5-83fa-4cbb-89f0-582422b9b7eb · outbound

This paper cites Open-vocabulary object detection using cap- tions.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Open-vocabulary object detection using cap- tions

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.124463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.124463Z digest=sha256:a0d5cb1fa651644118012b65bcc7e4808e9c55ce07fd5d2065ab7e6d8b2893d6

Observation 09b0e65d-2a0c-49e7-8238-0a0cd844b2c7 · outbound

This paper cites How to enable llm with 3d capacity? a survey of spatial reasoning in llm, 2025.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images How to enable llm with 3d capacity? a survey of spatial reasoning in llm, 2025

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.258520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.258520Z digest=sha256:dfaa9831055cff36f737660945c1df5401e230c737dda2d4666e099d56584044

Observation 1d04de49-7abf-40cc-ad8c-ccefd429fd3c · outbound

This paper cites From flatland to space: Teaching vision-language models to perceive and reason in 3d.arXiv preprint arXiv:2503.22976, 2025.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images From flatland to space: Teaching vision-language models to perceive and reason in 3d.arXiv preprint arXiv:2503.22976, 2025

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.314772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.314772Z digest=sha256:719799b019216b8fb98c674a11cea3a767163ffbb7b6f67e995b05a957f1bab3

Observation 904bbdd2-6d8e-42bc-9365-6868345a91a9 · outbound

This paper cites Spinbench: Perspective and rotation as a lens on spatial reasoning in vlms, 2025.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Spinbench: Perspective and rotation as a lens on spatial reasoning in vlms, 2025

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.422175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.422175Z digest=sha256:611b4551b4b75b103532faf9afd63c58dc3c9777beb07fb4fef59bd71ccd2f4f

Observation 2a6ba5e5-489d-4e13-b4e2-6678b7a104cf · outbound

This paper cites Video-3d llm: Learning position-aware video representation for 3d scene understanding.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Video-3d llm: Learning position-aware video representation for 3d scene understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.553330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.553330Z digest=sha256:d457de35222797784ead7af23afcca6985dc923da424e99ab2edf7d0655ffe08

Observation b94a55ac-cab6-4579-8625-16871717526a · outbound

This paper cites Scanreason: Empowering 3d visual grounding with reasoning capabilities.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Scanreason: Empowering 3d visual grounding with reasoning capabilities

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.596020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.596020Z digest=sha256:65bc1a6af62165d02edb3a5d472e6389165467b40b84940cf64baacd4756bf31

Observation f65d7cbe-43bc-4dbc-a2d9-6e2a7126a51f · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.652883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.652883Z digest=sha256:c95454dd06cf003e5e9cb5fc6093a11e1526cd934b6596c8dbb3db6b3cbcd383

Observation 9d293d6a-d7db-4986-a1e2-6b7bd44fcead · outbound

This paper cites A detailed description of each level and its cor- responding tasks is provided below.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images A detailed description of each level and its cor- responding tasks is provided below

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.761326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.761326Z digest=sha256:ece7877b2160afac46b5ec4d2f48e06610409ef19e3debf832f0323a5c3946a6

Observation ff698a18-5d01-4deb-9db5-359f450867ae · outbound

This paper cites Specifically, after obtaining high-quality raw data through filtering, we generate image captions and construct scene graphs to serve as our underlying database.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Specifically, after obtaining high-quality raw data through filtering, we generate image captions and construct scene graphs to serve as our underlying database

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.853194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.853194Z digest=sha256:d1a1bd448e6fdd67ddd02d6fc26b0135ab05bdac6af75396f3a00ecc39866f72

Observation f1e64af2-c666-458b-a4b7-1e450b6963cc · outbound

This paper cites - {variation_instruction} - The rewritten question MUST include: (a) A brief motivation clause describing WHY we need this information, consistent with the motivation hint.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images - {variation_instruction} - The rewritten question MUST include: (a) A brief motivation clause describing WHY we need this information, consistent with the motivation hint

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.909081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.909081Z digest=sha256:4a25dee462f902ca0509938878f50969d91da594a23ead2c0a2bcc50c7e519ed

Observation 591aa810-2306-491a-b5ac-fbdda1f876b0 · outbound

This paper cites - {answer_constraint} - {task_extra} - You MUST NOT flip yesno.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images - {answer_constraint} - {task_extra} - You MUST NOT flip yesno

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.975775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.975775Z digest=sha256:f78e68e3b70a958a3a9429583f14f981fd26800a1f0b2c0d5acabeabff2d685f

Observation 39a5c178-0457-4ead-b97f-b774b381fd20 · outbound

This paper cites thinking.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images thinking

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:58.030558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:58.030558Z digest=sha256:b145edc0ce28329d6bbbdd89e79df9c857c5e3d6c6ae1171d5e6247b0d041d8f

Observation 7ab6194a-8858-4083-bc70-cf1d59d2870e · outbound

This paper cites Scene In- formation adds global context, 2D Visual Prompts improve local grounding, and 3D Bounding Boxes deliver the largest gains through explicit geometric structure.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Scene In- formation adds global context, 2D Visual Prompts improve local grounding, and 3D Bounding Boxes deliver the largest gains through explicit geometric structure

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:58.090253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:58.090253Z digest=sha256:5ca095b9f74137103a9cf40f2d9d59bfb203c79d4734d63e1f94b837198708d1

Observation b027e5a3-9197-4b17-93b2-27c3c38441ef · outbound

This paper cites For the 3D bounding box information, each object is rep- resented by its center coordinates, spatial dimensions (size), and orientation expressed as a rotation quaternion.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images For the 3D bounding box information, each object is rep- resented by its center coordinates, spatial dimensions (size), and orientation expressed as a rotation quaternion

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:58.171986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:58.171986Z digest=sha256:7ba9d3746d1eab498aea829c25c854dcc889c2500acada66e47ac2cba349a27a

Observation d6e8e8b5-5ad3-4d78-8169-69fca60a2b03 · outbound

This paper cites All inputs are processed with the officialQwen2.5-VLprocessor, which supports dy- namic image resolutions up to 262,144 pixels (512×512).

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images All inputs are processed with the officialQwen2.5-VLprocessor, which supports dy- namic image resolutions up to 262,144 pixels (512×512)

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:58.262429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:58.262429Z digest=sha256:64ce6b28a92a7f58cb9f45ce0bb8a7eef09c385c5edd53b62f81c1ee5d5399bc

Observation 41a9de49-fc95-4dfc-8e96-9391fe0f4693 · outbound

This paper cites - Use the format <obj>...</obj> to describe your mapping between textual entities and object IDs.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images - Use the format <obj>...</obj> to describe your mapping between textual entities and object IDs

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:58.358583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:58.358583Z digest=sha256:40ac8b7bb74b04bf06a88c68c492ed6e1c9cfae667290fcfa686df05802db978

Observation bb5c9776-80e7-4524-8b8f-cb623da6ca9e · outbound

This paper cites an unresolved cited work.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:58.429529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:58.429529Z digest=sha256:cfb682fefd2dba348ed500213dfa38ad4a6f1889cb4dfc6f482deb427c0f81f5

Observation 5a1117f2-1d56-48f4-9935-77abe67ff869 · outbound

This paper cites Your output format should strictly follow: <obj> Object mapping: - entity_1 object {id_a} ({caption_a}) - entity_2 object {id_b} </obj> <reasoning>.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Your output format should strictly follow: <obj> Object mapping: - entity_1 object {id_a} ({caption_a}) - entity_2 object {id_b} </obj> <reasoning>

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:58.579422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:58.579422Z digest=sha256:4a9518f89e3a8a54b33d6596ea8bf14b5ff4b2be43bad597ce9b7795a34267af

Pith citing papers

No inbound Pith citation observations are available.