Pith. sign in

Paper Citation Record · LEDGER

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models

As of 17 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 1 inbound Pith citation observation for arXiv:2505.00788.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00788 v3

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:40:27.818667Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T18:42:12.968388Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 406a59fe-1e26-4614-8284-1220506dbdf7 · outbound

This paper cites GPT-4 Technical Report.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.528694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.528694Z digest=sha256:cad5de7e5bbde4fb6cf07e8ca67e80e71c0b5c8740e527ba0cb4d106c26161fb

Observation 3899654e-51cf-4047-8147-8fc65620184d · outbound

This paper cites SpaceLLaV A.https://huggingface.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models SpaceLLaV A.https://huggingface

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.715270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.533093Z digest=sha256:fa883dcc80457b9e821a779ffe12fa3a4d422429c31a43d14d3a8255d0584d22

Observation ef159a4d-37d4-4506-bfd3-d6e4a9577b20 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.537060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.537060Z digest=sha256:6eedb0494d25abbe6fd50c4df3af79103d329573658c7a108e3e381b18fc0ac7

Observation 54101c4c-4803-4e16-8d9f-cd5527539fa6 · outbound

This paper cites Claude 3.5 Sonnet.https : / / www.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Claude 3.5 Sonnet.https : / / www

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.691472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.540724Z digest=sha256:fe6e66cbdd93713ed6583a42f2df779d9e4bc67ac4832c27cbba72da40e7f2a6

Observation d08d78bf-58cd-48f3-a510-6f425e17e097 · outbound

This paper cites Apollo syntheic dataset, 2019.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Apollo syntheic dataset, 2019

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.676342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.544404Z digest=sha256:9d5d15b6eeeee129908dad66337d2f44a96c0ad91db5c6649e4c61446780a15b

Observation 1c09667c-7bfd-433e-beed-de79b75fe683 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Scanqa: 3d question answering for spatial scene understanding

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.662827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.548786Z digest=sha256:53397ebc6d975789a42915c716eae04e15c46fa2056b5e31f076fcdc01422f86

Observation b680ca9b-60bc-408e-b398-2ded5eba2dc2 · outbound

This paper cites Qwen Technical Report.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Qwen Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.553190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.553190Z digest=sha256:a079e2dfab07a6f2b9766684dfc804c47e2f70d66d8681ae453dd41fa23a5510

Observation 210bf3bf-22e8-47c5-928b-c18b4892d11b · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.557935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.557935Z digest=sha256:3674b3c4cef771c411cd7db2908b515b45d3878e3a11865d26780c84ee7ac1ec

Observation 014d3939-6660-4473-93c9-d6dac68e29ea · outbound

This paper cites ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.563585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.563585Z digest=sha256:574679768b04696eb0fafe3360eea419e0a6151cc674af92348896bc44eba15f

Observation 2379b6d3-b0e9-4702-8586-5485bee21225 · outbound

This paper cites ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.568854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.568854Z digest=sha256:68bd9fd0547a816909d50d34df7f1f3bc01140a91f9cee54da5d600e0c6d5d90

Observation bdcc0a7c-ab19-42be-8fca-2fcc0b0824dc · outbound

This paper cites Omni3D: A large benchmark and model for 3D object detection in the wild.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Omni3D: A large benchmark and model for 3D object detection in the wild

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.649771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.573588Z digest=sha256:96c4ef3743fe3bd963a2dc6a7a037d407af35c81a5940142ab553ec1d4e33273

Observation f7518ae4-2fea-4820-9626-77596264aef2 · outbound

This paper cites nuscenes: A multi- modal dataset for autonomous driving.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models nuscenes: A multi- modal dataset for autonomous driving

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.634820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.577671Z digest=sha256:3de8d68dbc58a993cb02a3f629b41ea2495418969fc28f7048f1ba1852b06e58

Observation b3b2ee18-f9ae-4442-b275-15616a4374b7 · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Honeybee: Locality-enhanced projector for multimodal llm

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.621016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.581845Z digest=sha256:cde2e96b3da8e4652c5210afaed02f0b86813203aa926cf10b0793a14d2e6043

Observation 6613005b-fd73-4555-b321-9c286ba2a0c4 · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.586160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.586160Z digest=sha256:efad014cd2b4166c41c917a9ad2df41a6ff29911d21a365352eec6f55ede3780

Observation bb75f2f4-66a5-4a0d-b871-a3b9d668f78a · outbound

This paper cites Vitamin: Designing scalable vision models in the vision-language era.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Vitamin: Designing scalable vision models in the vision-language era

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.598273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.590245Z digest=sha256:bf28a5f41d142412493e0ba6692026bdb6ae2d78929727c0a58a2f900d35f719

Observation ee237961-8d51-49cb-87d3-00b205a249fb · outbound

This paper cites SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.594324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.594324Z digest=sha256:0720e5490d9fc0993f6955ef796ad00c2d3463b3301ce9c26bf8ca4b0ebf3a41

Observation c7428130-25d9-401f-b0f8-ff71c63dfd4e · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Imagenet: A large-scale hierarchical image database

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.584277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.598798Z digest=sha256:413a84a73379e0b2c7d6c3028f9690834a16db375a810ed6a0ccfd893d587b0c

Observation e2fe4363-1272-494e-bdbc-866bceb42454 · outbound

This paper cites The Llama 3 Herd of Models.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.602876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.602876Z digest=sha256:be078574c0ac72a480121a17ca3b2de05a5617d72eaefbe50e034c615d9e3421

Observation 8f4a9e3c-3739-4cf3-b85d-a0bcc6ded7c3 · outbound

This paper cites Prob- ing the 3d awareness of visual foundation models.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Prob- ing the 3d awareness of visual foundation models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.571158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.607343Z digest=sha256:553b7ea9866fc825da3a1084c1f2dac423a3ee91aebfdf718801cb06b4f89fa1

Observation 01858a9a-0c1a-4c10-9b5b-b5316d65ceee · outbound

This paper cites DataComp: In search of the next generation of multimodal datasets.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models DataComp: In search of the next generation of multimodal datasets

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.611469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.611469Z digest=sha256:e24a9fe1ca39dc777f147920d39a0725306bf3bfc0a5691e6c270605705a6020

Observation bf252cec-806e-40cb-bddd-14c085680d2c · outbound

This paper cites Are we ready for autonomous driving? the kitti vision benchmark suite.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Are we ready for autonomous driving? the kitti vision benchmark suite

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.558228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.615737Z digest=sha256:5fdd2b1126b3a2a34642c85d30b003b0f2eb719290754fd59dc12d031d4d813d

Observation 963710cf-9f38-43ac-9697-a4faa7d1eb9a · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.619809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.619809Z digest=sha256:3df123c5e302f5e8d4ca2aa246afb219247f78691fb7bcd2fad5f547ec0f9e79

Observation 5853c801-4a09-4244-a610-be3722600d4d · outbound

This paper cites Masked autoencoders are scalable vision learners.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Masked autoencoders are scalable vision learners

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.535650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.624025Z digest=sha256:eecffbb465daf34a67a2f4c8d24ca9b09559506680722b20d1438da541f29706

Observation 0bb241a8-9b5a-49e6-a74c-b4964909fd04 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.628147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.628147Z digest=sha256:160fd4f944d78aa7c65018c8c20e70f9a180ea095c1d55b76d4c44334daf5c3c

Observation ab3e944d-9ffc-41be-8881-0afa42753c6a · outbound

This paper cites Open- clip, 2021.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Open- clip, 2021

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.632345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.632345Z digest=sha256:0d2c26c283a2c97f8f5886386bf28834c50a9b398c350f9b564bb0c559a1f378

Observation efad3a13-2355-4f59-857a-bcf21ed9e00d · outbound

This paper cites Novum: Neural object volumes for robust object classification.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Novum: Neural object volumes for robust object classification

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.636714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.636714Z digest=sha256:ca9103a4ef54a0725bdf7e9a35502cd3db45a285fa0599d050ba66fc244a8f59

Observation 0a18a1e3-b8c5-4f12-919f-d9372cfeb9b4 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Scaling up visual and vision-language representation learning with noisy text supervision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.641122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.641122Z digest=sha256:8c831d3ddf524d5b43d076cffcf235170f8a9aa51869f10c5a228a804636a564

Observation 8b70a587-e6b4-442f-a52b-039d3219ebb4 · outbound

This paper cites Mistral 7B.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Mistral 7B

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.645314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.645314Z digest=sha256:9ced4baf4d7ff69a20b21e81d6e2df484e670514a4a488d55af18e599a2d20e0

Observation 23def1e4-6e16-456d-a272-19374143cbe0 · outbound

This paper cites Perspective fields for single image cam- era calibration.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Perspective fields for single image cam- era calibration

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.496761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.650512Z digest=sha256:07468954550707d5957933064500126361d331938506c63b3e67fd60cbc78de6

Observation 0f29b4b2-cd60-4d8a-bccf-fe8980c20c1c · outbound

This paper cites Segment Anything.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Segment Anything

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.654030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.654030Z digest=sha256:47d474a29402249ad29842f4a098b9c512e23fae84f22b0e46d691869ae820cf

Observation d6ad7280-ab4e-4084-8836-82d3c4087e82 · outbound

This paper cites Segment any- thing.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Segment any- thing

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.482808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.657739Z digest=sha256:a3042afbc05548605f810ff0c44d8f77eb495ecc1314c9e1e4c49ee397fa7557

Observation e8503dba-ab6d-41a7-ba90-27a7802de402 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.468752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.661340Z digest=sha256:2abf49af34155f3af62f031f30c74fdd5010ea905a5ca852ec079471a1d68b7b

Observation 6133f4d5-e8dc-43b8-b982-5d7bc1ba9234 · outbound

This paper cites an unresolved cited work.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:40:28.455596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.664986Z digest=sha256:e83c8c3e44ef2b93a962bb388f84968791ccdf472408baa6446d23b81a484d80

Observation ec3d869b-35a5-4bf8-a173-9680cbc1cfd0 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.668620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.668620Z digest=sha256:576893813337a024a294fa642d7066fd8dfa87b5ed401e737a04c15bf76f5e58

Observation da067013-52b3-4923-b835-b331d5dff1cf · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.672593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.672593Z digest=sha256:a9f5232d2430ad7c273c73fe7f4082bf756a50c86540ef6bd5d07034ac8b7f3c

Observation 6a9eeaed-3940-4f5e-bf21-6cac2194feff · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.676116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.676116Z digest=sha256:a0890271db1a6bda54046b0c93565d4ac6f3775b0efa61a86da548eeeaa5541a

Observation e191db48-ad61-4223-9dab-3e802c457a8f · outbound

This paper cites What If We Recaption Billions of Web Images with LLaMA-3?.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models What If We Recaption Billions of Web Images with LLaMA-3?

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.679528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.679528Z digest=sha256:f3927bdc00911db7462a90c6961bd40998e943d3c043e9af44b56c08688f1c7e

Observation 2d8537d9-59b3-4597-bf88-eb1310a983ab · outbound

This paper cites Learning customized visual models with retrieval-augmented knowledge.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Learning customized visual models with retrieval-augmented knowledge

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.424103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.683992Z digest=sha256:76d54ee26324246788990c7ece18b82183c948d0ec57866d007d07b9ea5b5acd

Observation b5370f09-d4d9-4e65-823c-dd477f4ebdd9 · outbound

This paper cites Improved baselines with visual instruction tuning.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Improved baselines with visual instruction tuning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.411147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.688146Z digest=sha256:3b4813e2c69ee79f394a82ee7273c317f65fdd4354f35445195b8378f62482a0

Observation 8563a318-bdce-4f59-94e2-8c66228b8f8b · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.692491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.692491Z digest=sha256:242abadc3bdf9969788683794b11af4886bf1f85f79b0ed55faa7ad61621a8c4

Observation 9078385e-9047-431d-88a5-8539b76bdac5 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.389689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.696559Z digest=sha256:1187c77eb797e8fa83ca9dca0bf18f1172fe90522541dcc09fa94e7a16a831bd

Observation 2d40ee3a-6590-4210-971b-7cb5e430e012 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.701110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.701110Z digest=sha256:736d732dc7bc8db39c2eb70c415d3fa80513351f49217fbf6f5fce4b88492305

Observation 1b4f0ab1-d640-46cc-b582-d973fb3af87f · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.706213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.706213Z digest=sha256:6e5b83edb9fe693853c28090f115858c52dc1105cfd4a07374d2482f2f2e74e6

Observation 9d41879e-6ea7-45c5-9dd6-5fb92975967a · outbound

This paper cites Robust category-level 6d pose estimation with coarse-to-fine rendering of neural features.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Robust category-level 6d pose estimation with coarse-to-fine rendering of neural features

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.710906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.710906Z digest=sha256:67e7cbfbc85d924fa9607f78d076f8fa84a3a1818f7fbce8d01cfd18ae741a5a

Observation cb3e3658-1c82-496c-8a94-595420f171ad · outbound

This paper cites Imagenet3d: Towards general-purpose object-level 3d understanding.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Imagenet3d: Towards general-purpose object-level 3d understanding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.367631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.715210Z digest=sha256:fcba48eecaf1d445a46bc8ee662835fb05941a3dfa583c19304e624b86502465

Observation ee39a5d8-b409-4f20-afb8-66877bafd94d · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.719431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.719431Z digest=sha256:a829a6076da892007940b38366a35e88eb85980f75fea600badb01c142dd9bcc

Observation a0f212ba-c183-4ba0-ac22-e03927a53aff · outbound

This paper cites Dinov2: Learning robust visual features without supervision.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Dinov2: Learning robust visual features without supervision

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.353857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.723958Z digest=sha256:8eeb1cd9e1e8ac64af2eab28b60b30f516e6f2e6fb6a5f3274dbedc417670072

Observation 46c97bf3-ed26-462e-8532-8f29b56a2c81 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Learn- ing transferable visual models from natural language super- vision

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.340434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.728221Z digest=sha256:fe2a2ba809e839b9e734481af85e07cc723677dc713188e8949f3a273580d9a5

Observation ab98a02a-21ff-4a25-aa69-358024560e93 · outbound

This paper cites GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.732388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.732388Z digest=sha256:1eefc56221c771c38ed3a077d01b68dec8261020d832e0337e5b0da11c6c0d5c

Observation 5ce633a1-7150-4fe4-b5d2-0c05685ba58f · outbound

This paper cites Susskind.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Susskind

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.737271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.737271Z digest=sha256:f86688c73bb4f267b16812670d7397bdb1ca62748346f35cbb51f95e3b839d7d

Observation 0480b152-9d4f-456f-8a1f-4cf1911f0bd2 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models High-resolution image synthesis with latent diffusion models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.742087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.742087Z digest=sha256:f864e145e9899e1ecd26d18655992da39d303793224315e4639639bbe1fa3f73

Observation 006893b9-ae58-480a-858a-34eae99b0e66 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in Neural In- formation Processing Systems, 35:25278–25294, 2022.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in Neural In- formation Processing Systems, 35:25278–25294, 2022

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.746271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.746271Z digest=sha256:818fe4e4013bf29ba655869d4e66bbdddda129d380f8554ce1ae3faacb6dffd1

Observation d9eed4e5-c451-414b-9f1c-d369abc0d3cf · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.750251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.750251Z digest=sha256:58d1a892c09dd77f3ddf51d06e9c9c250b120698ce61e42e93f697116a78b125

Observation 38d7f83c-8884-48a2-a43a-5028083c44e9 · outbound

This paper cites Sun rgb-d: A rgb-d scene understanding benchmark suite.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Sun rgb-d: A rgb-d scene understanding benchmark suite

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.292913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.754431Z digest=sha256:e2cd1693023b36353de7a43f1dcea7bbc89d6eda9df4c4894159a957042308c9

Observation 5079e0ca-85ec-40d6-8dd8-fc7220e45669 · outbound

This paper cites Core knowl- edge.Developmental science, 10(1):89–96, 2007.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Core knowl- edge.Developmental science, 10(1):89–96, 2007

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.279345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.758501Z digest=sha256:88cf314fd9be082e62e54bb363028cecfbbd9e5f2da795a0dcc499590b1c3c38

Observation 2d34e57a-31d7-4600-b65c-438238908d63 · outbound

This paper cites Revisiting unreasonable effectiveness of data in deep learning era.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Revisiting unreasonable effectiveness of data in deep learning era

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.266026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.762558Z digest=sha256:ba8be2628f23fd1257ec4883241a58fec124ebcd5d8a0c133058fb0bdc082b03

Observation 136a0093-736c-4ee3-81e0-c392bb18cacf · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Gemini: A Family of Highly Capable Multimodal Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.766525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.766525Z digest=sha256:b76492f9ca8f82d9b8cce7897e656cd0aa14254434b5cb53b2359f3e86072357

Observation 7432001d-b856-4980-8200-07a6a7a57b7b · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.770707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.770707Z digest=sha256:59e29ae33b9cbde1467802fdb7f8075ce7a27d9bb4bd34e75942c75a732beba4

Observation 58112f68-ab1b-4fd0-9de8-9a674ab0493b · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.775308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.775308Z digest=sha256:0057953419ac43ff65e4853c22fb0f7bcf1b83415b496f5263ab71dea91bc482

Observation 1cb36f74-af69-48b2-b64b-007267717ddc · outbound

This paper cites 3d-aware visual question answering about parts, poses and occlusions.Advances in Neural Information Processing Systems, 36, 2024.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models 3d-aware visual question answering about parts, poses and occlusions.Advances in Neural Information Processing Systems, 36, 2024

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.252560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.779103Z digest=sha256:6c51d111147eaafbc3afdd4209c198f81cd3f12320fd5e7611f5416a493d07ac

Observation 87ad28bf-6065-4e22-a3aa-3ad93f02c52b · outbound

This paper cites Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.782481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.782481Z digest=sha256:a5481edd20f2d527fdce0409c70e0e00e6525595848241972de2df87c324e251

Observation 9413468e-6207-41cb-a66c-0ba8a4f1ed2b · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Depth anything: Unleashing the power of large-scale unlabeled data

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.786283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.786283Z digest=sha256:51e02cc26e1c289266a6f46dcf31ef4e19e68b87ccc6f747fcedacd0e121babf

Observation 5bcdd412-64c8-4b6c-bae0-b479654111c3 · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Depth anything: Unleashing the power of large-scale unlabeled data

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.789631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.789631Z digest=sha256:8b705b028a7727d69d6f0f106595e11a9611d1b4709624d9688e21bcb022ccac

Observation 641aa501-4832-4c54-a095-8c145200b324 · outbound

This paper cites 3D Question Answering.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models 3D Question Answering

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.793012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.793012Z digest=sha256:16c6ddad55197b9f047b2ca1aaf234afddd9d61a9f068bed5c08257376eefbe1

Observation 0a28dadf-2403-4a41-b966-8c54e4654bec · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.796715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.796715Z digest=sha256:027bbc99e12bd3627177e32ccc2eab9d36489774c6083dc0f977fcb98776cfeb

Observation c0976f13-42b8-4a8e-b0c9-79e2abede03f · outbound

This paper cites Recognize Anything: A Strong Image Tagging Model.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Recognize Anything: A Strong Image Tagging Model

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.800978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.800978Z digest=sha256:64990a539ca72844e0df14179224a05b4f7eb412956fd595681c6bf0798fa96e

Observation 05828b61-7e53-4419-969c-22fdbc6382e6 · outbound

This paper cites Recognize anything: A strong image tagging model.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Recognize anything: A strong image tagging model

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.222985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.805680Z digest=sha256:fcb5b1216a1116d7e20b1f764d8b3b0ba60146607c96d3cb9aeaac9d3d5598d7

Observation c6ea85e2-6324-4c84-89cb-3e5cc5a09a3b · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.210551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.809925Z digest=sha256:97a4faca1a657f1f2e28144c96ad0e1d89b5f0772af5c99b006b63047838f25b

Observation 3a2932da-f0a4-413f-a27c-6f50635ac041 · outbound

This paper cites iBOT: Image BERT Pre-Training with Online Tokenizer.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models iBOT: Image BERT Pre-Training with Online Tokenizer

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.814069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.814069Z digest=sha256:fd3f22c1bbb9f6316fb5c0686775ce5020391dcc9161573095ac32eb98f3b095

Observation 40f19090-a735-4d0d-a78f-2b57a9580da8 · outbound

This paper cites yes” as the answer and 120 questions have “no.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models yes” as the answer and 120 questions have “no

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.195914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.818667Z digest=sha256:247c6891e2e8264de6652b2e652f5ee6dbbd12c9790b8929d6d35a79a357bb84

Pith citing papers

Observation 2eeb54c1-f742-4702-aacc-667cb7965203 · inbound

SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL cites this paper.

SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T18:42:12.968388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:42:12.968388Z digest=sha256:8a43c303d146974578713fe812b5077b738fd31078653cf5815c95ca2b654ca1