Pith. sign in

Paper Citation Record · LEDGER

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation

As of 7 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 1 inbound Pith citation observation for arXiv:2506.23120.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23120 v1

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:53:04.127385Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T22:13:52.383043Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

77 of 77 outbound references displayed

  • verified exact2
  • verified fuzzy41
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 325774cd-e2b0-4437-9db1-1771425d731c · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:57.793154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:57.793154Z digest=sha256:bdf5946628f998b86cef8d49aa1f2e08bd87d0bab4bcd1f61bdaf8677c00eca5

Observation 6a7764da-d893-49aa-80d6-879fc4821c98 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Scanqa: 3d question answering for spatial scene understanding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:14.729965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:57.848548Z digest=sha256:b25dfbfa9093f531f9261f0652057c4beff28751c2ce9660df410d268b3fb8ad

Observation ec84b564-9542-4b2a-b53d-af68d50282d6 · outbound

This paper cites Semantic scene com- pletion via integrating instances and scene in-the-loop.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Semantic scene com- pletion via integrating instances and scene in-the-loop

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:14.576806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:57.941641Z digest=sha256:5dc6ecc900f3be78848edfdff29e76cd4caf7693953135b030f3c37e2911781c

Observation eb38820f-994a-4424-a259-5980f3e22e0d · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:14.406809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:58.068312Z digest=sha256:554778a1df7f17b02df0cd6ce4a30a578c1725916cd15516ec205928611e7678

Observation 7525dbd6-533c-4a95-babf-e4d55a965115 · outbound

This paper cites Language conditioned spatial relation reasoning for 3d object grounding.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Language conditioned spatial relation reasoning for 3d object grounding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:14.238738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:58.206308Z digest=sha256:336ce9ff92e8c0dee83bd218fe046612a9a078026508aa53c1fb394c70f8f630

Observation 5b1a3463-d775-4faf-b2e6-b2f6a048ebe2 · outbound

This paper cites Ll3da: Visual interactive instruction tuning for omni-3d understand- ing reasoning and planning.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Ll3da: Visual interactive instruction tuning for omni-3d understand- ing reasoning and planning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:14.061588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:58.293363Z digest=sha256:32a2d32daaa931d2987f01ac76ee00ff11139ed5b0ad91ef858ba5185b41ae88

Observation becd4c34-3a8b-4b1e-9027-ac8dbaaa3d1b · outbound

This paper cites Mppnet: Multi-frame feature intertwining with proxy points for 3d temporal ob- ject detection.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Mppnet: Multi-frame feature intertwining with proxy points for 3d temporal ob- ject detection

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:13.902145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:58.381758Z digest=sha256:7f8ba37d2f2e1f8c261661d5023caea5be516ca95bb1ded0ee4d66a015a774fa

Observation fd2f6363-7d1b-458d-a749-37362235b4e2 · outbound

This paper cites Trajectoryformer: 3d object tracking transformer with predictive trajectory hypotheses.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Trajectoryformer: 3d object tracking transformer with predictive trajectory hypotheses

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:13.687509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:58.480676Z digest=sha256:8a2625c25b025628072befba5cc3426cae8429213e42545902a756e9560e2463

Observation 1c299235-7d13-492f-b05f-0ee47acad1ef · outbound

This paper cites Grounded 3D-LLM with Referent Tokens.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Grounded 3D-LLM with Referent Tokens

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:58.607061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:58.607061Z digest=sha256:c48fa623af40194c8d28a408647f2de2a4903ad94d79678a5135c9a689465245

Observation a1c83751-31cb-426a-ac9f-b69ab4ad7c73 · outbound

This paper cites Reslt: Residual learning for long-tailed recog- nition.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Reslt: Residual learning for long-tailed recog- nition

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:13.513529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:58.701391Z digest=sha256:638444ece8dba7e8a81f04289f5addac85f83a7beae6b26bee27489dbf49147a

Observation 458a1831-5f2a-48fc-803e-39cdd608875a · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:58.811072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:58.811072Z digest=sha256:7ead5dd299d261d467165b4e2f9b18d8beb7399302550b9eb6e54c9ac4442253

Observation 94ac9a92-49d9-42d3-8d0f-78cc8cc7c625 · outbound

This paper cites MTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text Alignment.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation MTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text Alignment

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:53:04.929386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:58.917530Z digest=sha256:2efb7f65d202bf99c49eab54345608977860e53694928a95b8fd06af93837717

Observation e112cc66-d919-4cc3-b7ea-c1c26b72a5df · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.052695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.052695Z digest=sha256:6a5e53a3daaf79fdcc377d301e88ca49777af9c461ac9f96a264c5b70ac59c8c

Observation 90af0a97-5bb2-413d-9280-9877aa54cf5e · outbound

This paper cites Pla: Language-driven open- vocabulary 3d scene understanding.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Pla: Language-driven open- vocabulary 3d scene understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.155141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.155141Z digest=sha256:9d0803f063594af508a6e6d40ef4f42e7e046229bf070dc794cf89d7cd33b4e1

Observation de2e4887-707f-4e28-96b1-bdf9f352fea9 · outbound

This paper cites The llama 3 herd of models, 2024.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation The llama 3 herd of models, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:13.360728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:59.243701Z digest=sha256:ba58a2f86397600776589736e9011b013cc248aff387c2ecf42f000a282e4067

Observation 44a72786-e440-4241-b3d0-a9436f35572d · outbound

This paper cites Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.369243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.369243Z digest=sha256:a468237f1fc19af2672833d195dc0915416d6b1c3d48d132955f50b9910abe09

Observation ad57e631-7596-4529-8f48-cc1ce6ee5a4d · outbound

This paper cites ImageBind-LLM: Multi-modality Instruction Tuning.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation ImageBind-LLM: Multi-modality Instruction Tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.473536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.473536Z digest=sha256:3c87c45763bd08beee1b69f50f32fa801f5e0c131a6ce08005bb0a4a46995f21

Observation 97ce80bc-d7d0-4338-a09c-467d0f2c3547 · outbound

This paper cites 3d-llm: Inject- ing the 3d world into large language models.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation 3d-llm: Inject- ing the 3d world into large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:13.146921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:59.578599Z digest=sha256:7147902a7f4d7340ea9afb116112efc721393bf76b9d428fcad0e7b17854e7dc

Observation 0f9edb34-9b72-4b97-8b5d-91ae6ec18f30 · outbound

This paper cites Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.681855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.681855Z digest=sha256:44ee12346f97d59a42b4e111028a82403f24eb7cece2881e4d441de78888e8c7

Observation 34d95fbc-e937-4715-9210-663522ad7705 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation An Embodied Generalist Agent in 3D World

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.769267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.769267Z digest=sha256:bfaaf47e7b1927a9fc6e315dc88bf27ca022be68de07fe83b2bb91230c8519ae

Observation 4b4618bd-cdea-4d2c-828d-1afa9f1c645b · outbound

This paper cites Multi- view transformer for 3d visual grounding.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Multi- view transformer for 3d visual grounding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:12.946706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:59.867853Z digest=sha256:f2d1a4f758e754917b5ac315ef1640aa1b4a8c745c5fea4e50dd53c5c6e7c422

Observation 62e5aba4-dde3-4665-b802-8ea3177c3189 · outbound

This paper cites SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:53:04.545494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:59.994335Z digest=sha256:cd791346827cb00f2f912fa7f158fc8a9595544b604961a99a17af615275f643

Observation ef4f5f50-f464-48c0-82e9-bf42cd0c62ed · outbound

This paper cites Guided point contrastive learn- ing for semi-supervised point cloud semantic segmentation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Guided point contrastive learn- ing for semi-supervised point cloud semantic segmentation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:12.761403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:00.083655Z digest=sha256:d34320b844d303085bb267ae132449fe050decaa6a08752d9e322b1c8a6928ef

Observation 831e018b-d7b8-4946-89c5-d2f9cd460c8e · outbound

This paper cites Segment Anything.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Segment Anything

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:00.204418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:00.204418Z digest=sha256:7eb18dedb017cf8291b4759af4128464b04b7f2226d08464390334de59c208b2

Observation 7188b26a-7abb-4603-90da-ee7f35813e5a · outbound

This paper cites Oneformer3d: One transformer for unified point cloud segmentation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Oneformer3d: One transformer for unified point cloud segmentation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:12.583338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:00.306050Z digest=sha256:d9af341ceb90efa52ccf72ca48c8f17d7ca6e0b3a2f0fcfd717065c5da65dd5d

Observation 76ace732-993f-481d-af74-7ab8261c1a13 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Lisa: Reasoning segmentation via large language model

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:12.403341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:00.437200Z digest=sha256:8851f8281754566b43cf893b46252ad013e037a9debcd4d116302bd75c06f133

Observation d9653dec-3380-4c78-8316-8b21d5210f57 · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:00.558586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:00.558586Z digest=sha256:c62ee599cbfa699c75b957cfa43b29bfc5883c745c03ed0f9bd74fcd7b9d5efb

Observation b4becde9-e5fd-46d0-80c8-0fe3f452cab6 · outbound

This paper cites Large-scale point cloud semantic segmentation with superpoint graphs.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Large-scale point cloud semantic segmentation with superpoint graphs

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:12.158422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:00.644188Z digest=sha256:40770d2358e3f9195098bec05e6b4b697372e41fe4ec459279734f8ea9af2b2b

Observation dc14eb7e-771a-4c7f-b789-4daefd08ace7 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.967553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:00.717881Z digest=sha256:317a277d7de4999c86c49bdba4dc20ecc9fb064029974676550687e1e88aed6f

Observation 3ea31733-91bf-48ee-8dce-ef838f65b9e2 · outbound

This paper cites Uni3D-LLM: Unifying Point Cloud Perception, Generation and Editing with Large Language Models.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Uni3D-LLM: Unifying Point Cloud Perception, Generation and Editing with Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:00.799961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:00.799961Z digest=sha256:e8c06ee9a72ecd9264eaebaeb25961a359d7d807bcc7d27b79f3ca6c702f6673

Observation c4d0f4aa-fc8b-469b-bf92-4d7ef8153b75 · outbound

This paper cites Decoupled Weight Decay Regularization.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Decoupled Weight Decay Regularization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:00.901821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:00.901821Z digest=sha256:4ff3763f537656c5e4e3b67a02b3e95b6cf20bd4500402ca7bbc7d530de54c00

Observation 27d46f3b-d777-4076-b839-5e64e9a65f93 · outbound

This paper cites SQA3D: Situated Question Answering in 3D Scenes.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation SQA3D: Situated Question Answering in 3D Scenes

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:00.971816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:00.971816Z digest=sha256:b572bba4374e693987fa3c4539002e474a632c91efdb8a5da97294cf2f64807d

Observation 146b51d3-c45f-413a-90dd-24a29d730431 · outbound

This paper cites An end-to-end transformer model for 3d object detection.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation An end-to-end transformer model for 3d object detection

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.753791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:01.046044Z digest=sha256:fbfa1624c119d9e32dc4601c2b713ea496fd43a5ffb443bedf87109d8d6e8020

Observation dd071416-a327-43c7-a1db-a6996a849c5f · outbound

This paper cites Boosting few-shot 3d point cloud segmentation via query-guided enhancement.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Boosting few-shot 3d point cloud segmentation via query-guided enhancement

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.575644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:01.149234Z digest=sha256:ef0218cbb3b8509173ec690e6ea1b36a8cd36428eea18fb4093fc560f2d24ee1

Observation 50ced8bf-31e5-4b03-8fc6-e2c65b4f4d2d · outbound

This paper cites Hierarchi- cal dense correlation distillation for few-shot segmentation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Hierarchi- cal dense correlation distillation for few-shot segmentation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.421226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:01.255571Z digest=sha256:67958e06f6e5556ab663a6bf4b348595ee49e282cdc4b1935b29fc59c9338d87

Observation f7cae70b-0dc2-4b0c-8eb7-ce7ca5f1d669 · outbound

This paper cites Scalable Language Model with Generalized Continual Learning.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Scalable Language Model with Generalized Continual Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:01.348484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:01.348484Z digest=sha256:528c72f7f80dfbce5e5d48cbee3d11b3a2c8f25411d33dbcd3f9f6d64178f732

Observation 584d8855-8324-4fd3-aae2-bdc8b758c3ba · outbound

This paper cites Oa-cnns: Omni- adaptive sparse cnns for 3d semantic segmentation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Oa-cnns: Omni- adaptive sparse cnns for 3d semantic segmentation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.173646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:01.438300Z digest=sha256:8cd450b769c00b4abef57c1ebd1de1bc802d397a7540c38542b1cfcf7f709d18

Observation 0d582eb0-37e5-4246-a451-0eb88695eccc · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Openscene: 3d scene understanding with open vocabularies

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:01.538579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:01.538579Z digest=sha256:4a388b97b3fcb1dc34e202a8c4a130d8b74008f24e42ec7cd6666c88bda0e5b9

Observation 56dc775b-f6a2-446e-a8c0-172202b6e50c · outbound

This paper cites Explore the potential of clip for training-free open vocab- ulary semantic segmentation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Explore the potential of clip for training-free open vocab- ulary semantic segmentation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.983916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:01.634815Z digest=sha256:106cb11a5fa23478cfd2d6bedf4dc6b2ae1d0962f8c9ac151b88f6df022e39ba

Observation 61f5a869-bb7a-4f88-b675-271d3b0fe13b · outbound

This paper cites Pointr- cnn: 3d object proposal generation and detection from point cloud.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Pointr- cnn: 3d object proposal generation and detection from point cloud

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.816844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:01.702626Z digest=sha256:af5a9535bc4b6a5c5307f457cc4874fcf36b14151c66c0551ecc0689eec1b144

Observation 1096ace6-d48f-4595-b0ac-f2ab5e1a36b7 · outbound

This paper cites OpenMask3D: Open-Vocabulary 3D Instance Segmentation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation OpenMask3D: Open-Vocabulary 3D Instance Segmentation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:01.780508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:01.780508Z digest=sha256:ce09cdb359fd72269f84f5119324b235ac1eb8c8590d8c8e730529e22a07e6c5

Observation 78a01afb-911b-4a4b-b494-590f26f7df4c · outbound

This paper cites Mind the interference: Retaining pre-trained knowledge in parameter efficient continual learning of vision-language models.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Mind the interference: Retaining pre-trained knowledge in parameter efficient continual learning of vision-language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.628744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:01.855126Z digest=sha256:d14ddea20de3995f48f7d9598f259dbc1cb96153be8a67d12607ecd6eda35f74

Observation 934c49e3-811d-49b1-b6ca-78cb5612802d · outbound

This paper cites Learning shape-aware embedding for scene text detection.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Learning shape-aware embedding for scene text detection

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.428096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:01.946814Z digest=sha256:878b7b279ceb921ff594f830baa29905cabe8815eed43770488ebc6387383f33

Observation 0c26d983-be1a-410b-ae62-43726fbc61fc · outbound

This paper cites Prior guided feature enrich- ment network for few-shot segmentation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Prior guided feature enrich- ment network for few-shot segmentation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.189030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:02.042318Z digest=sha256:2f2b752282a1afe2926c537681d5dd8d62ed928fb46a0d8bfc2d7ab4461e9daa

Observation 8e314eef-7b3c-45b6-b717-432117460287 · outbound

This paper cites Adaptive perspective distillation for semantic segmen- tation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Adaptive perspective distillation for semantic segmen- tation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.980659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:02.106679Z digest=sha256:ae167d853a921310d40e2441cb1123d3b17b0edac108b962654fc3bff9669d8d

Observation 17d04d7b-54fd-4bcc-8dec-367477883c72 · outbound

This paper cites Generalized few-shot semantic segmentation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Generalized few-shot semantic segmentation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.786638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:02.173893Z digest=sha256:0de9830ca43462acaec30bb74d0f1e90c24bf5ccf7e003ba906563947857d156

Observation 8cc5ac61-fe44-428f-8d39-d6f24a46680b · outbound

This paper cites Learning context-aware classifier for semantic segmentation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Learning context-aware classifier for semantic segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.577862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:02.264287Z digest=sha256:76b9406b809cc9fe245619a5f1c4ef0c7c75f7d01a0f6991b734e27d810f9672

Observation 04974178-964d-4fc1-abcf-d397d4456184 · outbound

This paper cites Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:02.361077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:02.361077Z digest=sha256:43c4cfa64f9cf8d0639aca72b1a64a9c98d582308d70d8b3ca38cd2f6a7c5138

Observation 19030596-2323-44f8-bf0a-ff83d7969880 · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:02.436049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:02.436049Z digest=sha256:52f234ebb701968c4c8e996961a2783088164eb99acd712fe440997ab98bbd9b

Observation 8e24a87e-fa6e-47ec-80aa-350393141259 · outbound

This paper cites PointLLM: Empowering Large Language Models to Understand Point Clouds.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation PointLLM: Empowering Large Language Models to Understand Point Clouds

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:02.513697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:02.513697Z digest=sha256:13a17491db9cfc8a50ce453bf629c63be6f76178d3e38e71f82cc6ee2a7e7e8f

Observation cd0acc59-f07b-445c-a6af-dd175fe92d31 · outbound

This paper cites Llm- grounder: Open-vocabulary 3d visual grounding with large language model as an agent.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Llm- grounder: Open-vocabulary 3d visual grounding with large language model as an agent

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.433475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:02.621278Z digest=sha256:65e75faa561c94787df16c565c9713dbc70a92311723e99b101da9a1821cf887

Observation 222944e9-9eef-4452-891d-bcb190a9e218 · outbound

This paper cites LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:02.670233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:02.670233Z digest=sha256:dd50638a0eced4d0305f0b1162954205b270ab4da64b7b29e6d23e8881e0f15d

Observation 691d61e1-8cbe-46c6-bf68-9bd49f55c352 · outbound

This paper cites Unified language-driven zero-shot domain adaptation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Unified language-driven zero-shot domain adaptation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.286190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:02.785477Z digest=sha256:5551e02b311322c51cbc29acfc657e609d7c3a26990b5af9a1adbeffac959961

Observation e364c948-8f20-4ccd-9c55-f72f539d2bca · outbound

This paper cites Visionzip: Longer is better but not necessary in vision language models.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Visionzip: Longer is better but not necessary in vision language models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.155714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:02.857736Z digest=sha256:40faa5880d2057c2a94697e2369b60bdc556f391d87248200943662fe5be45a8

Observation 59b157f9-8fa1-4d26-9496-c327aeed8fc7 · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:02.923647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:02.923647Z digest=sha256:30a6eef85257dbfacfcefad54b6238ae6cc396f66c0986299cf50ce63a27f795

Observation 5ab31b1a-6074-49f5-a2cf-f4ec5ac7f4da · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:02.977100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:02.977100Z digest=sha256:501e613f5c013b931de171d7a85e138fac87240eae04502628c9cc53063a9c3b

Observation 944d160a-1c9a-4f48-a492-c43445372d40 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation OPT: Open Pre-trained Transformer Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:03.034745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:03.034745Z digest=sha256:4a969645cd133a1b336e471901ef19dfc3016b1c79e9ee8415deda85ab03dae7

Observation 1ea536c1-6748-486c-aed5-8131ee72c7df · outbound

This paper cites Multi3drefer: Grounding text description to multiple 3d objects.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Multi3drefer: Grounding text description to multiple 3d objects

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.978931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:03.093626Z digest=sha256:8aa68e884b2296be3faf92563b84daa4b1575322bbbe1bb996770eb5ff986fb8

Observation ce0c7d61-939f-4525-b2a1-c6a915a0a248 · outbound

This paper cites 3dvg- transformer: Relation modeling for visual grounding on point clouds.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation 3dvg- transformer: Relation modeling for visual grounding on point clouds

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.839030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:03.159558Z digest=sha256:cd42a6d6f33ac5b432307b01ad30bf8a9380be40755b48fb3f39e5a47a5d4abd

Observation 5fd12392-a455-441a-8630-a817b0b433fe · outbound

This paper cites ScanReason: Empowering 3D Visual Grounding with Reasoning Capabilities.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation ScanReason: Empowering 3D Visual Grounding with Reasoning Capabilities

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:03.225641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:03.225641Z digest=sha256:2ec0a0497dd08dcb16bebb85b5cfe1cfdabc09a3d4b72561291045c96e0ba91a

Observation 8820c6be-a73e-4ae8-a85b-38ae4e588bf0 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:03.285974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:03.285974Z digest=sha256:528b67dfd3b020eb36076d5a5b849223c57c98cf65227d96f42d0e0f6867b264

Observation c264643a-7b76-4de0-9b77-f692ca3febba · outbound

This paper cites in the right of.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation in the right of

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.641534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:03.364743Z digest=sha256:3c62de0a099b602afc5f884e0846047399c5aece1a860b44131f12f4d31f4a12

Observation 79c14c01-dd42-438c-8312-73dd76869e8b · outbound

This paper cites Focus on detecting and distinguishing objects.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Focus on detecting and distinguishing objects

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.491329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:03.415149Z digest=sha256:d450c0da059c80f5fd4cafee69e54483499e16987de2ae13bc8c7f92c1566dd1

Observation 325f6a96-a76f-4951-8e96-bad0ba2a8387 · outbound

This paper cites an unresolved cited work.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:53:08.344472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:03.466751Z digest=sha256:3bb123e42b99db8ac4baa3d948b9eea3170942b74bace943da0eb5d2a96983e3

Observation 05db9656-4a26-484c-add9-192c02bf0113 · outbound

This paper cites an unresolved cited work.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:53:08.151275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:03.522108Z digest=sha256:2e8999c7fd6a4b243cd11b447def6d7187184cc0ab961ead563115110124c780

Observation 560ad566-91ec-441d-af03-e42af7682409 · outbound

This paper cites an unresolved cited work.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:53:08.004162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:03.581218Z digest=sha256:6a8d0d85c5e2612a8a6f8697cc93f0992758934c6708549bed3418ed51e60db9

Observation 61cf7d28-9358-4867-85b6-8c1c07519fab · outbound

This paper cites an unresolved cited work.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:53:07.821554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:03.628077Z digest=sha256:82791ac0a4872b17f1f0ad44e19cd4455fc737a0c84cbf1d25577523996fe222

Observation ac8e90e1-e941-4283-8a6f-0a8799e076db · outbound

This paper cites an unresolved cited work.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:53:07.673168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:03.675226Z digest=sha256:ffd3d5beb3000e67195748668856f1efc9aaf48f5435fe67b9e01ca371df4623

Observation ba3b2e56-4854-4d61-ad78-6809cc68d482 · outbound

This paper cites For multiple objects of the same category, use [<ins_id>][<ins_id>][<ins_id>], etc.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation For multiple objects of the same category, use [<ins_id>][<ins_id>][<ins_id>], etc

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:07.463313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:03.739329Z digest=sha256:0badac6e21800f765e647733ad9babbb2edc45aa66bc278f1de97f9db8654bda

Observation 90115f2d-f5d1-4533-8061-275982f0466d · outbound

This paper cites an unresolved cited work.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:53:07.134037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:03.785678Z digest=sha256:0887bfdfda96c8b9afc707d51931c76f2cd5a795eeb666ffed50452d06696000

Observation 83c98ba2-e58d-42af-8cde-eca333407c73 · outbound

This paper cites an unresolved cited work.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:53:06.791230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:03.839144Z digest=sha256:afab9d785df1116d4388344cb1002d8787ac1699f7078c6741930343d7c68ec8

Observation 684938be-0073-4770-8639-a5573a11fc08 · outbound

This paper cites Don't output anything else.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Don't output anything else

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:06.518179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:03.898635Z digest=sha256:cb0c504a0239d2caf826379a0c26b871c362b7aa0c4572dafb87966dbf795f00

Observation dfe122b5-8d40-4404-9fa1-d59bf2f82354 · outbound

This paper cites Analyze the properties and spatial relationship of objects in the scene.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Analyze the properties and spatial relationship of objects in the scene

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:06.182178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:03.961199Z digest=sha256:3a3c4e1e426c9353e639f30f06e0a01f7116a712174ca1fccd471caf00d16ec1

Observation 555da127-ea1a-47e9-93b8-43a22e8f0a3c · outbound

This paper cites Don't output anything else.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Don't output anything else

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:05.893817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:04.025687Z digest=sha256:edf0fe08b84548fc333651d298ef2aae194d55e3e859cc51ae5db94c65d186a5

Observation 1dc3a3b7-e30e-4d0a-a3b0-e4bf029bbf62 · outbound

This paper cites an unresolved cited work.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:53:05.705709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:04.088203Z digest=sha256:eb275fa2645c42503fe8a43dfb0b4b5f39a7bebc27c184f181fe541d03d6008a

Observation 0c966068-2eff-4233-b2eb-aea3fb3858e4 · outbound

This paper cites Analyze the description and identify description-related objects.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Analyze the description and identify description-related objects

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:05.526378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:04.123579Z digest=sha256:945a8747f5c8a24e69ab826f4ab3312a5d71ccad95611429ab8e593dc4333a8d

Observation 4ab6c0e7-5250-42a7-a73c-9fe22efa05d4 · outbound

This paper cites Don't output any other things.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Don't output any other things

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:05.267420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:53:04.127385Z digest=sha256:3ce0ad78cdaf544bf1fe096383f27f3540b3fcdac37d17b20386e0398b03c435

Pith citing papers

Observation ba066641-729c-4b3d-b6c6-854f6ebf7940 · inbound

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation cites this paper.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:52.383043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:52.383043Z digest=sha256:ab5e80b154415d045e37ef9f6329f55ea55c91cf0d75ad2c4ec48ceed787881b