Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T20:38:58.579422Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2511.19119.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T20:38:58.579422Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6171e5a5-9779-4182-b1c1-21820f9db851 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94e10a9f-6b82-451b-a339-8de016a6fb1e · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Scanqa: 3d question answering for spatial scene understanding
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12207c0e-8757-48fe-8568-fd461fa8b6be · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Qwen2.5-VL Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 710d92b2-ac15-4abf-bc0b-12b491aa1962 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fd9d04f-37b2-42d8-a1e5-62fe5678f162 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Omni3d: A large benchmark and model for 3d object detection in the wild
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89fc6ba7-bae3-419e-9de8-09a4738a8d4a · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images SpatialBot: Precise Spatial Understanding with Vision Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c086a39-011b-49a0-9149-adc2640e85c3 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c6b597c-a254-49e1-ab03-f7a18a027b21 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Perception before reasoning: Two-stage reinforce- ment learning for visual reasoning in vision-language mod- els.arXiv preprint arXiv:2509.13031, 2025
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8942ee0-7dbf-45b9-ac42-d7d87e28b24f · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Spatial- rgpt: Grounded spatial reasoning in vision-language models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 508972a0-ed4b-495d-ad5b-d2c5e2b04dab · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Collins, Ilia Sucholutsky, Umang Bhatt, Kartik Chandra, Lionel Wong, Mina Lee, Cedegao E
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e73763e9-b8bb-4338-ad19-4d2eadd1c8f7 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72fc96a9-1852-4a86-a0bd-ed1dcd958dde · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eee8f01c-22c9-4336-b728-07c975bf8000 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Surds: Benchmarking spatial understand- ing and reasoning in driving scenarios with vision language models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab6ab7ac-37fd-4510-8642-cfb12d71f050 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images 3d-llm: Inject- ing the 3d world into large language models.NeurIPS, 2023
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 909bc398-84c5-4dd4-b423-b1ccc8b10f08 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images What’s ”up” with vision-language models? investigating their strug- gle with spatial reasoning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 346095bd-61b2-48d3-81e5-e9db8400d8db · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Lee, Jihyeon Je, Chanho Park, Mikaela Angelina Uy, Leonidas Guibas, and Minhyuk Sung
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 532c478c-b7da-43c5-85d3-6ad7c014b912 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Seed-bench: Bench- marking multimodal large language models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cee1622d-630d-4358-9580-72b51b92f0c4 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images LLaVA-OneVision: Easy Visual Task Transfer
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95c1dadc-6fda-4c81-8d38-d9f20f4b0ee3 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Viewspatial-bench: Evaluating multi-perspective spatial localization in vision-language models, 2025
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd0f1fc6-39d2-4a64-b2c2-0c78d649dc57 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Visual Instruction Tuning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8761fb1f-1515-494e-9010-dd67e17c3095 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Spatialladder: Progressive train- ing for spatial reasoning in vision-language models, 2025
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f54e9786-4ca9-4ced-b037-d896d0b231c2 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2d028f0-fe31-4c10-8379-6933e4834c96 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65d6b3ed-76fc-44a4-bc01-4fb57d4847a8 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images A Novel Multi-Agent Deep RL Approach for Traffic Signal Control
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7def5476-984c-4f25-ab84-09af786d2853 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Visual spa- tial reasoning.Transactions of the Association for Computa- tional Linguistics, 11:635–651, 2023
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9a6c7f1-deac-4565-833b-e8f48d1b5405 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Grounding dino: Mar- rying dino with grounded pre-training for open-set object 10 detection
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06e758fc-b272-4fee-b172-b22316866e94 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images 3dsrbench: A compre- hensive 3d spatial reasoning benchmark.arXiv preprint arXiv:2412.07825, 2024
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 809fd4d9-e63b-4352-8c61-e370e45e891c · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Sqa3d: Situated question answering in 3d scenes
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5c72476-ad3f-498d-b895-80649dd2dc72 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21436383-2142-4cfc-a7ad-c80fe38dfc3b · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images GPT-4 Technical Report
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d48f1012-52dd-4854-80a4-4ad5e0383e39 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Shapellm: Universal 3d object understanding for embodied interaction
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50e00068-8775-49a4-9a94-2dbaf0723a35 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Learning transferable visual models from natural language supervision
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e11018a-2d3c-4ed3-8394-c1f711d24bc9 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec6506ca-7cad-48d3-9bda-080642b9689e · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Space3D-Bench: Spatial 3D Question Answering Benchmark
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b37d4bd3-ca58-4b9d-8dc0-095031994996 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Gemini: A Family of Highly Capable Multimodal Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ed634ff-f6e0-4037-8e01-a0acee736260 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e059e4b6-49a5-44ef-bc97-727183fc8bdf · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images LLaMA: Open and Efficient Foundation Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df08ed63-aad8-4ade-8dee-f858468399a6 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Cross-modal pro- jection in multimodal llms doesn’t really project visual at- tributes to textual space
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e583bfd-8301-48ac-8aa6-092681b606bb · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Learning 3d semantic scene graphs from 3d indoor reconstructions
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4a5cb51-d025-44e9-b5f3-a00c59a9361b · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Is a picture worth a thou- sand words? delving into spatial reasoning for vision lan- guage models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7acbe468-8842-49d1-99d4-ca4dfe6618ae · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Spatial 3d-llm: Exploring spatial awareness in 3d vision-language models,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04f4922f-24b2-44f6-bfe3-9389f52a9375 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0373d929-38c4-4c28-a8b6-d3554f1e5416 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Pointllm: Empowering large lan- guage models to understand point clouds
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45090bbc-c81e-4083-83e5-efdd44782115 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ec212f5-83fa-4cbb-89f0-582422b9b7eb · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Open-vocabulary object detection using cap- tions
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09b0e65d-2a0c-49e7-8238-0a0cd844b2c7 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images How to enable llm with 3d capacity? a survey of spatial reasoning in llm, 2025
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d04de49-7abf-40cc-ad8c-ccefd429fd3c · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images From flatland to space: Teaching vision-language models to perceive and reason in 3d.arXiv preprint arXiv:2503.22976, 2025
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 904bbdd2-6d8e-42bc-9365-6868345a91a9 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Spinbench: Perspective and rotation as a lens on spatial reasoning in vlms, 2025
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a6ba5e5-489d-4e13-b4e2-6678b7a104cf · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Video-3d llm: Learning position-aware video representation for 3d scene understanding
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b94a55ac-cab6-4579-8625-16871717526a · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Scanreason: Empowering 3d visual grounding with reasoning capabilities
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f65d7cbe-43bc-4dbc-a2d9-6e2a7126a51f · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d293d6a-d7db-4986-a1e2-6b7bd44fcead · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images A detailed description of each level and its cor- responding tasks is provided below
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff698a18-5d01-4deb-9db5-359f450867ae · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Specifically, after obtaining high-quality raw data through filtering, we generate image captions and construct scene graphs to serve as our underlying database
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1e64af2-c666-458b-a4b7-1e450b6963cc · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images - {variation_instruction} - The rewritten question MUST include: (a) A brief motivation clause describing WHY we need this information, consistent with the motivation hint
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 591aa810-2306-491a-b5ac-fbdda1f876b0 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images - {answer_constraint} - {task_extra} - You MUST NOT flip yesno
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39a5c178-0457-4ead-b97f-b774b381fd20 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images thinking
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ab6194a-8858-4083-bc70-cf1d59d2870e · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Scene In- formation adds global context, 2D Visual Prompts improve local grounding, and 3D Bounding Boxes deliver the largest gains through explicit geometric structure
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b027e5a3-9197-4b17-93b2-27c3c38441ef · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images For the 3D bounding box information, each object is rep- resented by its center coordinates, spatial dimensions (size), and orientation expressed as a rotation quaternion
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6e8e8b5-5ad3-4d78-8169-69fca60a2b03 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images All inputs are processed with the officialQwen2.5-VLprocessor, which supports dy- namic image resolutions up to 262,144 pixels (512×512)
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41a9de49-fc95-4dfc-8e96-9391fe0f4693 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images - Use the format <obj>...</obj> to describe your mapping between textual entities and object IDs
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb5c9776-80e7-4524-8b8f-cb623da6ca9e · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Unresolved cited work
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a1117f2-1d56-48f4-9935-77abe67ff869 · outbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Your output format should strictly follow: <obj> Object mapping: - entity_1 object {id_a} ({caption_a}) - entity_2 object {id_b} </obj> <reasoning>
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.