Pith. sign in

Paper Citation Record · LEDGER

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?

As of 21 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 14 inbound Pith citation observations for arXiv:2505.11907.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11907 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:51:31.758682Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:32:28.221452Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T14:44:59.676019Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy35
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4214ac27-6a40-462f-b73d-e87ebcebe72e · outbound

This paper cites DiMeR: Disentangled Mesh Reconstruction Model.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? DiMeR: Disentangled Mesh Reconstruction Model

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:31.466858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:31.466858Z digest=sha256:91e6bcbf35faed8b0c09b12cc54d2c315901465922f596a91a6e70a4ba188ee8

Observation 13c3f5ed-6459-4174-b27a-c63e3a76e60a · outbound

This paper cites Benchmarking Multi-modal Semantic Segmentation under Sensor Failures: Missing and Noisy Modality Robustness.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Benchmarking Multi-modal Semantic Segmentation under Sensor Failures: Missing and Noisy Modality Robustness

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:31.472307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:31.472307Z digest=sha256:2b7fc5f74452e21b55b3f03d8c05d8f67259ec23fe71eb5f59eefc678b6d1b86

Observation 8fad7963-54e1-4a98-9828-eb9d623d3ff6 · outbound

This paper cites Retrieval Augmented Generation and Understanding in Vision: A Survey and New Outlook.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Retrieval Augmented Generation and Understanding in Vision: A Survey and New Outlook

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:31.477574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:31.477574Z digest=sha256:bf4a40b89132712314c2ed48fe7939be6df4714554787ff848f8371152b7d98e

Observation 70ab66af-2035-4dca-8452-8ead588d9fb0 · outbound

This paper cites an unresolved cited work.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:51:32.832702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.483190Z digest=sha256:c99b23dc9cf7eeeac05fca07af38f765499a9eafae0cb9f219e92072034fabd3

Observation 0784b6be-8b51-4dff-a70b-90249b7e4bfa · outbound

This paper cites Realrag: Retrieval-augmented realistic image generation via self-reflective contrastive learning.arXiv preprint arXiv:2502.00848, 2025.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Realrag: Retrieval-augmented realistic image generation via self-reflective contrastive learning.arXiv preprint arXiv:2502.00848, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:31.488912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:31.488912Z digest=sha256:eaa382e15702dec67c53d21e7b6b1c6ada7a0926462359a0369b77ffe14f2b17

Observation 028b6d53-39c1-4865-99e6-28ca29b58ce1 · outbound

This paper cites Distilling efficient vision transformers from cnns for semantic segmentation.Pattern Recognition, 158:111029, 2025.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Distilling efficient vision transformers from cnns for semantic segmentation.Pattern Recognition, 158:111029, 2025

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.817168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.494478Z digest=sha256:a8faead6aece168aaaf3b39309d4e87ed53cc5d9ab23f33cab50706bb60b29de

Observation 121fc705-5832-4cee-af3c-84de5d094459 · outbound

This paper cites Open panoramic segmentation.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Open panoramic segmentation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.802445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.500793Z digest=sha256:e6799eff2d71e426f3f3c65af8ecc7e94e4b92180e67a927ba6eef20610f7fa9

Observation beb1a02b-cdc7-43f1-990e-0cd6a24b67a6 · outbound

This paper cites an unresolved cited work.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:51:32.787547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.505514Z digest=sha256:90f95b9caf31107a1aebdf1202a1fb712622cde62001272d20f39ea4474d7e15

Observation 7c636476-17af-4d46-ab90-d74adba73f63 · outbound

This paper cites Bending reality: Distortion-aware transformers for adapting to panoramic semantic segmentation.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Bending reality: Distortion-aware transformers for adapting to panoramic semantic segmentation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.773275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.510445Z digest=sha256:49824a2263e80bdf3370281f0eb064ebb7fa98f52e9751b40b3a2aa6ca076f2d

Observation d8d1ac6c-5b6f-47d1-a115-7afb6215fd73 · outbound

This paper cites Both style and distortion matter: Dual-path unsupervised domain adaptation for panoramic semantic segmentation.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Both style and distortion matter: Dual-path unsupervised domain adaptation for panoramic semantic segmentation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.760108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.515073Z digest=sha256:2f170e7a8cfc465fa86c4d277ff303f1b33e30c7084b7cbc823b583a77c72316

Observation 9843ecef-aee0-4fd2-8d84-185b9ae87b95 · outbound

This paper cites Look at the neighbor: Distortion-aware unsupervised domain adaptation for panoramic semantic segmentation.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Look at the neighbor: Distortion-aware unsupervised domain adaptation for panoramic semantic segmentation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.744876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.520281Z digest=sha256:ccab27bbe400cad5d83bcd33f1b92cf7045774fbb7d5c68ec64028432d1862f6

Observation 0e76e5c1-375f-4509-bb81-41f4456f96f2 · outbound

This paper cites Semantics distortion and style matter: Towards source-free UDA for panoramic segmentation.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Semantics distortion and style matter: Towards source-free UDA for panoramic segmentation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.728543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.525080Z digest=sha256:25f5a7f7137361a5e75e0059243106fcc1330434a586aeded049fde710158d57

Observation 626b6e06-a49b-471e-8644-f9d1e75a95a0 · outbound

This paper cites 360SFUDA++: Towards source-free UDA for panoramic segmentation by learning reliable category prototypes.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? 360SFUDA++: Towards source-free UDA for panoramic segmentation by learning reliable category prototypes.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.709614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.529410Z digest=sha256:750555e23c1c42a717fd99389bae7d227c6981cff14e66d136f89bd6ffe9afc8

Observation 398d28f0-daa8-46a1-a793-53f9a5265b38 · outbound

This paper cites GoodSAM: Bridging domain and capacity gaps via segment anything model for distortion-aware panoramic semantic segmentation.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? GoodSAM: Bridging domain and capacity gaps via segment anything model for distortion-aware panoramic semantic segmentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.693015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.534045Z digest=sha256:d17f043089b979fb5dcd62423c5023a5c7a765ccaeeb0db0b3bd39d86900ff1f

Observation 3142aa25-10d0-486a-ae87-5048ad68b0c4 · outbound

This paper cites GoodSAM++: Bridging Domain and Capacity Gaps via Segment Anything Model for Panoramic Semantic Segmentation.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? GoodSAM++: Bridging Domain and Capacity Gaps via Segment Anything Model for Panoramic Semantic Segmentation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:31.540206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:31.540206Z digest=sha256:828d29f0145378c9f151a184e7633d8e7859937d3b9b6895185f01bf3d8ed925

Observation eb6dbcc6-b488-4061-98af-5429ad2a95e6 · outbound

This paper cites Interact360: Interactive identity-driven text to 360° panorama generation.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Interact360: Interactive identity-driven text to 360° panorama generation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.679012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.544942Z digest=sha256:0b4a745cf49722f90c46c440f7c0123c2fa933746f9c1d4b57728018bd6419b4

Observation 570585d6-6026-4e15-8ee7-a09043b878ad · outbound

This paper cites OmniSAM: Omnidirectional segment anything model for UDA in panoramic semantic segmentation.arXiv preprint arXiv:2503.07098, 2025.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? OmniSAM: Omnidirectional segment anything model for UDA in panoramic semantic segmentation.arXiv preprint arXiv:2503.07098, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:31.549290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:31.549290Z digest=sha256:f64bae173e2f9712d4703da327d8e89d7f4e57484030fe992e81a9931d95ba5e

Observation 6df75775-2aec-45fd-b1fb-a298f94f6a80 · outbound

This paper cites A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:31.553533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:31.553533Z digest=sha256:810e91d76154df31e02306a227b0ee1a37e0fa842b520b53450265976ec31de9

Observation 08e7849a-553b-40aa-8ccb-4b3ae5c41d3a · outbound

This paper cites MMUnlearner: Reformulating Multimodal Machine Unlearning in the Era of Multimodal Large Language Models.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? MMUnlearner: Reformulating Multimodal Machine Unlearning in the Era of Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:31.558531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:31.558531Z digest=sha256:43c5418c8e1bf709183327549ed2b06d61939ed1dc9e7119a0ae8b944c22159d

Observation eac0d469-bb35-4c0f-a042-2660fef78b6b · outbound

This paper cites Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:31.563568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:31.563568Z digest=sha256:7335726d6d9603da4df192a0a2d1393f83786d145c5d0dd773252f327de7b75b

Observation d5fca540-2f9e-4915-9589-d34fb13ffacb · outbound

This paper cites ManipLLM: Embodied multimodal large language model for object-centric robotic manipulation.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? ManipLLM: Embodied multimodal large language model for object-centric robotic manipulation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.664657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.568515Z digest=sha256:5d29445b57221911af2edcf4de406986045ea5c214159d9b4cc81793e338fe84

Observation 07c0994c-d3dc-4572-96c4-6ff697c6052b · outbound

This paper cites PanoContext-Former: Panoramic total scene understanding with a transformer.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? PanoContext-Former: Panoramic total scene understanding with a transformer

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.649071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.572268Z digest=sha256:b2f1245d3997ce9149fa43df79eeac1a0c7fe14bd028e178585299f9f9fc6bc8

Observation 0635ce79-230d-4b04-99de-1e391dc4052c · outbound

This paper cites DeepPanoContext: Panoramic 3D scene understanding with holistic scene context graph and relation-based optimization.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? DeepPanoContext: Panoramic 3D scene understanding with holistic scene context graph and relation-based optimization

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.633053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.576309Z digest=sha256:0a8fc760bba0c8eeb9b340b121376bdf80fb723614820b656f41fdcf5eba0374

Observation 7c3e25e5-62da-46bd-92ad-97c3d1f54668 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Evaluating object hallucination in large vision-language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:31.580309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:31.580309Z digest=sha256:4001eaaf04d67a34c5ef9fc869ded4680ff74a0e761c2dde7654d1b6581b8706

Observation a02d84da-aec8-44a2-8a7c-4edf2972f277 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:31.584843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:31.584843Z digest=sha256:36648316b5f8359ee7b88f73bad164f8d8cd5862dd47e9b62279916e965ad123

Observation 9cf1cdc1-6641-4683-9450-9f27a3adffd1 · outbound

This paper cites Improved baselines with visual instruction tuning.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Improved baselines with visual instruction tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:31.588941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:31.588941Z digest=sha256:9104f5a910d4f1e533aa8e4b534327c2114a3fa0a0063390f9dfe92268e74af7

Observation 9aae7c2e-e0c8-4a49-83e9-8c91683df178 · outbound

This paper cites The Llama 3 Herd of Models.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? The Llama 3 Herd of Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:31.594178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:31.594178Z digest=sha256:429e88fd0546eeea93dd2d03161193b1a6def8e36b0040e297fa6cde70babc2c

Observation 6e42aa9d-577f-474c-88c4-75a07bb12a9e · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:31.598531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:31.598531Z digest=sha256:facb379cba438702f4d5b1c59d7164eceb4277795bb21988ee804e5626dff2b9

Observation 1cf03ce4-eff8-4aaf-81f4-04ba940ade45 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:31.603940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:31.603940Z digest=sha256:1f2a1add4b13ea2c0ff3729d74609ab6f636e4666e122a221160479a25bb7404

Observation 92036eef-d850-4dbf-a0a6-b39be2cca922 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:31.608920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:31.608920Z digest=sha256:7272bcf75ba16de134b68ec20a3b3e19ac54ddcc1c97a53a23dd63de9c330fe4

Observation 75f581e6-b00b-420f-b677-da9e5e3027ae · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in visual question answering.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Making the V in VQA matter: Elevating the role of image understanding in visual question answering

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:31.613470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:31.613470Z digest=sha256:0e335fb82acd687f2283723573f1249320691bb63e9d53ef83f6f46f037ec72c

Observation 7f905c13-fc30-4379-9e92-0a22c32257e8 · outbound

This paper cites Hudson and Christopher D.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Hudson and Christopher D

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.590353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.618972Z digest=sha256:f3c5bbc31963917b1a6f63b6370dcf45ef2ee0e1a81420ef046185614ff6236f

Observation c9b02bf9-8ced-45f6-940a-e782356d66e1 · outbound

This paper cites Towards VQA models that can read.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Towards VQA models that can read

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.575694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.623910Z digest=sha256:8cf32b6be6f95273b41bc0ea7b34c9d921d277ca1ebfd235ae8329074bfd3814

Observation 7581812c-b3dc-4a77-8c4b-41c72ae610a2 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:31.628793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:31.628793Z digest=sha256:8ad5f33073288656008617298188026d5b52993d1a5c5294f09604d2fff7dfab

Observation c3fbdf7c-4920-4435-bb66-f403d66fc6a9 · outbound

This paper cites MMBench: Is your multi-modal model an all-around player? InECCV, 2024.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? MMBench: Is your multi-modal model an all-around player? InECCV, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.560204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.633746Z digest=sha256:e762702067b3af11fd26c3699f426e4a5ad9d44f927d2de79a2ece78bf572ba9

Observation 2670ec3b-f06e-4451-91ea-fb3680d9d5d8 · outbound

This paper cites MM-Vet: Evaluating large multimodal models for integrated capabilities.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? MM-Vet: Evaluating large multimodal models for integrated capabilities

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.545683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.638767Z digest=sha256:eeeb71d6b5210b187e7306bb8bf089db269220f897e0e226ee96cac5a3f06308

Observation 89e0d854-ff64-4357-b169-3230df3c043a · outbound

This paper cites ConvBench: A multi-turn conversation evaluation benchmark with hierarchical ablation capability for large vision-language models.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? ConvBench: A multi-turn conversation evaluation benchmark with hierarchical ablation capability for large vision-language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.532148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.643896Z digest=sha256:4ffaa67ec8fad5685e99d062c183c2308407d5f77c6a943f04267e7994ac03b7

Observation 4d151204-0300-428d-9c56-7d688f13ce34 · outbound

This paper cites MMDU: A multi-turn multi-image dialog understanding benchmark and instruction-tuning dataset for LVLMs.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? MMDU: A multi-turn multi-image dialog understanding benchmark and instruction-tuning dataset for LVLMs

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.517026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.648750Z digest=sha256:590a3c86972754947e5d20d06bb54a21449c2fed35bbbe33eb96795285080006

Observation 7cb02c7a-f0ab-472b-a916-b689cc2dd027 · outbound

This paper cites Negative object presence evaluation (NOPE) to measure object hallucination in vision-language models.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Negative object presence evaluation (NOPE) to measure object hallucination in vision-language models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.499498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.653583Z digest=sha256:c692ab8ab6bf0304509c81d8ca7d9e5a03c369f39698bcd54f2028ab54c33b77

Observation 3edc5dcc-c52c-4eb8-a66a-47e2e2d5bc74 · outbound

This paper cites Is a picture worth a thousand words? Delving into spatial reasoning for vision language models.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Is a picture worth a thousand words? Delving into spatial reasoning for vision language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.485937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.658304Z digest=sha256:39eeebfffa4c10b7290ba55fcdd688b387059173aa7cc0f91c66c88297814e46

Observation 600d5c69-87e5-4f7b-9d4a-e8c202713e70 · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:31.662128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:31.662128Z digest=sha256:a8932cae81d1eaa83aaa357397edfac3ab56dfbb8563093e5e73f1b79b7448d1

Observation ae96d098-aa7a-4242-a8f6-01602e25d62e · outbound

This paper cites PanoContext: A whole-room 3D context model for panoramic scene understanding.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? PanoContext: A whole-room 3D context model for panoramic scene understanding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.471976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.666353Z digest=sha256:6a645e54b7c820e16282c3488eee3ff35b4035f28663b7dd2db3e9d2d79fa7a7

Observation adb0e3b3-a19c-4190-9a70-c7d35c92f7ac · outbound

This paper cites PANDORA: A panoramic detection dataset for object with orientation.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? PANDORA: A panoramic detection dataset for object with orientation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.456784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.670313Z digest=sha256:5eaa5589b6c3c88992608eab6b0920eb55ca27efab05c10f50c2deb884aacdaf

Observation af5873e6-6efc-4e53-93dc-a8116a359538 · outbound

This paper cites 360+x: A panoptic multi-modal scene understanding dataset.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? 360+x: A panoptic multi-modal scene understanding dataset

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.441868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.673907Z digest=sha256:ba2413b7c50449be50a8b9db8def22d09eb01eea96027d35c45d1a466227ba8f

Observation 2077dc25-a16d-427b-a09c-a2f8a7629596 · outbound

This paper cites an unresolved cited work.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:51:32.427967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.677636Z digest=sha256:998b2aa611d76302f18cb11a3ff21713f236d74dcaa31158e6b3bb22176dff0a

Observation 219a45de-cd58-4170-a147-611527598695 · outbound

This paper cites KITTI-360: A novel dataset and benchmarks for urban scene understanding in 2D and 3D.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? KITTI-360: A novel dataset and benchmarks for urban scene understanding in 2D and 3D.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.413359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.681513Z digest=sha256:d5b302dcfe7ca7ed5d8aa8824fc6574da3830edfabbf91bba340fd5a79e20350

Observation 30fd200c-f404-4ea8-93a9-37c453595474 · outbound

This paper cites DensePASS: Dense panoramic semantic segmentation via unsupervised domain adaptation with attention-augmented context exchange.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? DensePASS: Dense panoramic semantic segmentation via unsupervised domain adaptation with attention-augmented context exchange

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.399002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.685963Z digest=sha256:93e3237b41c0d6f0bd9149427e75dd60c6cf8981bb152d2bbdada09e4d80fd78

Observation c8a77633-2c0d-4a9c-ad96-c3e1df8835a0 · outbound

This paper cites Waymo open dataset: Panoramic video panoptic segmentation.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Waymo open dataset: Panoramic video panoptic segmentation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.385373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.690488Z digest=sha256:ea0339dc5f1d64eff8cbe11527a11e00f8061edba536a98c0ee638cd6b90a5c0

Observation fa630c73-4294-4d64-9245-51bebdc91ddc · outbound

This paper cites Matterport3D: Learning from RGB-D data in indoor environments.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Matterport3D: Learning from RGB-D data in indoor environments

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.372905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.694941Z digest=sha256:f59d4d46a0e136770a50512f387b3b8e9e5183648c497a9657d8bfd253fd5b67

Observation e2761595-c591-484d-bc73-7c1f1be9feef · outbound

This paper cites Joint 2D-3D-Semantic Data for Indoor Scene Understanding.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Joint 2D-3D-Semantic Data for Indoor Scene Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:31.700107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:31.700107Z digest=sha256:2839ffaecf8a8e32e49da2b77594b3cbdb63b1ee3a0581295e8c40dd6776d0a3

Observation aa6ac5ed-d5e8-413f-a500-b3c9ffaac082 · outbound

This paper cites Pano-A VQA: Grounded audio-visual question answering on 360° videos.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Pano-A VQA: Grounded audio-visual question answering on 360° videos

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.357532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.704956Z digest=sha256:1bee05a01e8cb9055a11eef922d32e6220d4f4b7da1de731958cf5d2ceb46519

Observation d807fb79-af2c-4d0c-ab4f-f156358a3020 · outbound

This paper cites Visual question answering on 360° images.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Visual question answering on 360° images

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.341341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.709690Z digest=sha256:16eb16d76c3c386be511efe9a0f94caac8c5261568fc85201b56d9d9e4e7f5cd

Observation d5e6b545-85d8-4e2c-b233-f8d7dd2c8a1d · outbound

This paper cites Karen Liu, Hyowon Gweon, Jiajun Wu, Li Fei-Fei, and Silvio Savarese.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Karen Liu, Hyowon Gweon, Jiajun Wu, Li Fei-Fei, and Silvio Savarese

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.326278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.714215Z digest=sha256:83da229bbd804c7852167492bd4ec9d8d0c259051fbffa92cdd2db4c6ad760be

Observation d80ecb80-6266-4bbb-ac13-1ca8b84e11e9 · outbound

This paper cites Video question answering for people with visual impairments using an egocentric 360-degree camera.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Video question answering for people with visual impairments using an egocentric 360-degree camera

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.311488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.718565Z digest=sha256:d1a06d66d9102be98f57702d7c4902e1a01de15cb5eb07cd0ba4719a1fd070d2

Observation 3bb88956-953d-4f05-a67f-bcd29659e3ce · outbound

This paper cites Xing, Hao Zhang, Joseph E.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Xing, Hao Zhang, Joseph E

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.296586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.722680Z digest=sha256:8913aa9174f9b61f8d1ee41bc822250182b8c92167e0c3ea618a8c1023eb7c5d

Observation 217aec4a-e930-457d-a9e2-17dca4a54a01 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:31.728159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:31.728159Z digest=sha256:43cad3b72581efec7ef278c3fd2dcf647178823bc28a4897d988100a4c5e1bbe

Observation de5ea926-adf9-46ab-8fad-1903fde67e1c · outbound

This paper cites GPT-4 Technical Report.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? GPT-4 Technical Report

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:31.732763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:31.732763Z digest=sha256:4c60cd37e994c4fe3dfc69639aa2e3d5da90d75104f1c9d70b7d06ac8579fa71

Observation 87ad5648-def1-49e3-b02a-7cd1f3e3a691 · outbound

This paper cites Focus ONLY on these categories.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Focus ONLY on these categories

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.282222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.737771Z digest=sha256:e6a364f4c66ac127c5569d702a691891a15b1ea03544c71fdaadb7119073c476

Observation e9785a36-596e-456e-94ee-c5a27f1faf84 · outbound

This paper cites an unresolved cited work.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:31.742672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:31.742672Z digest=sha256:a44d166d45621f640b060b23a4353646656de660abd2d67eb1c13791a471053f

Observation 9723daed-ab6e-42c5-8fdc-42d47941d9cb · outbound

This paper cites an unresolved cited work.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:31.746930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:31.746930Z digest=sha256:9f391aef9fe9903c024974212b49987567bc63e711c4dc1693a8bbc6b68f0799

Observation 5285f894-7fb9-4a2d-aad6-9f9daeba6d50 · outbound

This paper cites an unresolved cited work.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:51:32.247369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.750846Z digest=sha256:c2ab67d78792b0aff36f721fdc2494aa56143034eeb368a35c5bd71d7baf7f45

Observation 38b4334b-acc8-497f-a704-7a11033745d0 · outbound

This paper cites Closer objects should be placed near the center of the grid, while distant objects should be placed toward the edges.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? Closer objects should be placed near the center of the grid, while distant objects should be placed toward the edges

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.232321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.754752Z digest=sha256:d2384a0945a3ac4b2a868ad88c2b7903fb09097e4cdcec8f5bbdd9b71b1e3baa

Observation 529e2f2f-8085-41b3-85d7-3439d0be5ec7 · outbound

This paper cites cate- gory name.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? cate- gory name

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:51:32.214292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:51:31.758682Z digest=sha256:531f3c73ae8d4a91cf4d40cd51ebad56d0d03e1a556d7fbc17abd752faacd1af

Pith citing papers

Observation 5abb6290-bf85-4906-895c-34d5a20f2458 · inbound

Omnidirectional Spatial Modeling from Correlated Panoramas cites this paper.

Omnidirectional Spatial Modeling from Correlated Panoramas Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T11:56:37.717534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:56:37.717534Z digest=sha256:75ef1a6aaa36ca99eaba6f68b3680d4e2c8f1af81c77c327810f11644ec3df7b

Observation 43e4b6e8-1baf-4d16-a8c0-8bb6bfbead22 · inbound

One Flight Over the Gap: A Survey from Perspective to Panoramic Vision cites this paper.

One Flight Over the Gap: A Survey from Perspective to Panoramic Vision Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?

Reference 272

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:28.221452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:28.221452Z digest=sha256:de606356687541a5bc3a3bf5d32b0c8337f2c2dc6d0a15a4a3b46c84c8abd021

Observation cc70dddd-d1b8-4cf1-9a8b-80502f10d403 · inbound

Beyond Thinking: Imagining in 360$^\circ$ for Humanoid Visual Search cites this paper.

Beyond Thinking: Imagining in 360$^\circ$ for Humanoid Visual Search Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:18.289687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-12T02:58:46.728868Z digest=sha256:44533515c0a7793baa7d5e803b9428d8005a0de4e0081312041b9a0666519add

Observation 2c2bf1a0-1fb2-42d4-ac68-d17823b0c0fa · inbound

Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images cites this paper.

Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:57:27.391208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T06:56:37.792662Z digest=sha256:ac272d72c179e3c0ae687059f2f97c81ac6d67b8ec94192121156c054fd9c3ea

Observation 2f122681-9e17-437f-833e-f78c75ec0125 · inbound

Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images cites this paper.

Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:39:29.060872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T21:38:09.573431Z digest=sha256:ff56894cafa2f2badb1713e4655546f1e36eaf5b5d1c4a681b00a39967e1ca5e

Observation 2a4ae096-c122-466e-8e03-116dbae3ed40 · inbound

Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images cites this paper.

Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:29:09.857085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T22:24:38.232462Z digest=sha256:a3ebf7f59d205786bd8e0aca88ebf9b7ed41e7f83b07ffe8e8afac3ffd00935a

Observation a2a1c787-b80a-49f8-9dae-5e07b9cc7707 · inbound

PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World cites this paper.

PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:42:57.553435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:40:59.877854Z digest=sha256:c2d8376b5b89191cd25f358e2545445ff91c3b8e8ea1c134c195f731fe736cdf

Observation bf4e4261-fd0e-4fd9-b8fd-ddea000b385f · inbound

PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World cites this paper.

PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.120350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T16:57:03.172340Z digest=sha256:9159102c9c09e140458bd1851091f0b111f120e6e8ffe97573b63455f2f11588

Observation 77cc1285-2155-41fa-9d5e-fe6602069f07 · inbound

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching cites this paper.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.097894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:e61f5602edec38315d5fa03ff1b055f96afef2117f625345cb0a6b239da374f2

Observation 1162cc6d-6e32-4727-b1c5-185155f95162 · inbound

Panoramic Scene Understanding: A Survey from Distortion-Aware Engineering to Sphere-Native Modeling cites this paper.

Panoramic Scene Understanding: A Survey from Distortion-Aware Engineering to Sphere-Native Modeling Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-29T20:03:56.622701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T04:35:58.372801Z digest=sha256:5acc438b10d53d475e0a34e19e8ac47e0c47f76909246431ed2501f88327e738

Observation d6159810-5c67-4fa8-9da0-f4b16dd4cc83 · inbound

Panoramic Scene Understanding: A Survey from Distortion-Aware Engineering to Sphere-Native Modeling cites this paper.

Panoramic Scene Understanding: A Survey from Distortion-Aware Engineering to Sphere-Native Modeling Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T09:58:47.727146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:58:47.727146Z digest=sha256:813826f5c81a8217347aa2cc849e128638aefbaf80504936ad70a684854a230b

Observation fc623953-8777-4802-9eaf-ef25c84ff2d4 · inbound

OmniCoT: A Benchmark for Global and Multi-Step Panoramic Reasoning cites this paper.

OmniCoT: A Benchmark for Global and Multi-Step Panoramic Reasoning Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:14:19.082359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T06:11:09.693576Z digest=sha256:7f50881b750d1f2c832ab12486d33a0f934c5dac16c839c9224c3216152c7bb9

Observation 3f274e56-1d19-4a06-a910-12c2785f3bcb · inbound

Seek to Segment: Active Perception for Panoramic Referring Segmentation cites this paper.

Seek to Segment: Active Perception for Panoramic Referring Segmentation Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:48:32.618936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T14:39:22.617747Z digest=sha256:63b657346bdeb6ac19bd2c34594a1c431b7f993461ea659a673e89d4138991a0

Observation 86bf2789-2a3d-4959-8d59-7f1a160e6572 · inbound

EAGOR: Embodied Reasoning in Omni-direction cites this paper.

EAGOR: Embodied Reasoning in Omni-direction Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-08T14:44:59.678569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T14:44:32.242554Z digest=sha256:9f15de8a77b926cc68ed5745851c3f5212b115d8f9ac2b2c9b035de5f944596a