Pith. sign in

Paper Citation Record · LEDGER

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models

As of 17 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 1 inbound Pith citation observation for arXiv:2505.00788.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00788 v3

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:40:27.818667Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T18:42:12.968388Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 406a59fe-1e26-4614-8284-1220506dbdf7 · outbound

This paper cites GPT-4 Technical Report.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.528694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.528694Z digest=sha256:76528234bc094e326612f54da88ae806a2da814db10b1c975394084c5231398d

Observation 3899654e-51cf-4047-8147-8fc65620184d · outbound

This paper cites SpaceLLaV A.https://huggingface.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models SpaceLLaV A.https://huggingface

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.715270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.533093Z digest=sha256:f49fa15eb6e3387afcd6a60a0f2e65956688cf1fcd9199c60c3a6ad8afa2523e

Observation ef159a4d-37d4-4506-bfd3-d6e4a9577b20 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.537060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.537060Z digest=sha256:ccd6bfe13039d7fbf8e17b4a7bbfca040fe81d055c62da344882d39d9c47b186

Observation 54101c4c-4803-4e16-8d9f-cd5527539fa6 · outbound

This paper cites Claude 3.5 Sonnet.https : / / www.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Claude 3.5 Sonnet.https : / / www

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.691472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.540724Z digest=sha256:20c2c8cbd86472d952b980358a070b336ebf785b82e1e7c7f81e4e38917abe5d

Observation d08d78bf-58cd-48f3-a510-6f425e17e097 · outbound

This paper cites Apollo syntheic dataset, 2019.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Apollo syntheic dataset, 2019

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.676342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.544404Z digest=sha256:75398f3274e2f1935aa87a6de2547210d585b9dbcc3420b94a5faa7b5832c714

Observation 1c09667c-7bfd-433e-beed-de79b75fe683 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Scanqa: 3d question answering for spatial scene understanding

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.662827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.548786Z digest=sha256:21daa0d168f6fa5c66f90b3f654abb315c5a5593a3111ac780c0c0e3bcea071f

Observation b680ca9b-60bc-408e-b398-2ded5eba2dc2 · outbound

This paper cites Qwen Technical Report.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Qwen Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.553190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.553190Z digest=sha256:2b4e498bd25c3719f217fcf6b372869a02be7edffa81a8b3d9336a34bfb6a4c2

Observation 210bf3bf-22e8-47c5-928b-c18b4892d11b · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.557935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.557935Z digest=sha256:d8bb79eafc8022f7857ff41411d684f35b9a3804709f40b4edf384545a95bad4

Observation 014d3939-6660-4473-93c9-d6dac68e29ea · outbound

This paper cites ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.563585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.563585Z digest=sha256:7b5075af98adaa5c4db2315c39737da6dad1d417a3d56315db0672959be47ca3

Observation 2379b6d3-b0e9-4702-8586-5485bee21225 · outbound

This paper cites ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.568854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.568854Z digest=sha256:2c42cd5b1a34326ff6c3acaa356f8cec72e5c53b248dbf8d43878c64ab0182a8

Observation bdcc0a7c-ab19-42be-8fca-2fcc0b0824dc · outbound

This paper cites Omni3D: A large benchmark and model for 3D object detection in the wild.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Omni3D: A large benchmark and model for 3D object detection in the wild

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.649771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.573588Z digest=sha256:e60b3c908d00d2171586c69b56e8cdab37dc09896ffee4be42854c11a728448e

Observation f7518ae4-2fea-4820-9626-77596264aef2 · outbound

This paper cites nuscenes: A multi- modal dataset for autonomous driving.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models nuscenes: A multi- modal dataset for autonomous driving

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.634820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.577671Z digest=sha256:3f45a27857b9f8f692fbbee7e08a5cae08a723812f80cf5f8cef3cb374091a95

Observation b3b2ee18-f9ae-4442-b275-15616a4374b7 · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Honeybee: Locality-enhanced projector for multimodal llm

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.621016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.581845Z digest=sha256:b26dce8527f5fe995f01f913eb9e350204c47127c0153ee4b6aa2d2c55e89f24

Observation 6613005b-fd73-4555-b321-9c286ba2a0c4 · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.586160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.586160Z digest=sha256:9bf698fd7357c5e7f4c7a1d8198909ef4b77513ed885876bae303c1646ee00f5

Observation bb75f2f4-66a5-4a0d-b871-a3b9d668f78a · outbound

This paper cites Vitamin: Designing scalable vision models in the vision-language era.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Vitamin: Designing scalable vision models in the vision-language era

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.598273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.590245Z digest=sha256:e7ebb4e2ae20856b5b24f885d86986ab33e8796f9390b43f66d18b2b0cbee420

Observation ee237961-8d51-49cb-87d3-00b205a249fb · outbound

This paper cites SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.594324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.594324Z digest=sha256:4afbac1ab5625a877803a9c371bdb03b3d274ddeae8efaf1334cffb71b28bee1

Observation c7428130-25d9-401f-b0f8-ff71c63dfd4e · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Imagenet: A large-scale hierarchical image database

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.584277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.598798Z digest=sha256:63731034329bfdbdc78cef9452903898a28323c08be6868eed1060ad00650dbb

Observation e2fe4363-1272-494e-bdbc-866bceb42454 · outbound

This paper cites The Llama 3 Herd of Models.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.602876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.602876Z digest=sha256:d186ba999acad0b4df001176d0a1149e49acf10a7b2438ac879d661d0b95fa7e

Observation 8f4a9e3c-3739-4cf3-b85d-a0bcc6ded7c3 · outbound

This paper cites Prob- ing the 3d awareness of visual foundation models.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Prob- ing the 3d awareness of visual foundation models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.571158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.607343Z digest=sha256:d695b584750ba45e685f3f0206a33a4ec7b3ec8d993e9d3969bd314e8dc7dbab

Observation 01858a9a-0c1a-4c10-9b5b-b5316d65ceee · outbound

This paper cites DataComp: In search of the next generation of multimodal datasets.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models DataComp: In search of the next generation of multimodal datasets

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.611469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.611469Z digest=sha256:44e64b10efea7c73b5d987ba16bc1a2b882fa58781c7a4777a8478156b95c996

Observation bf252cec-806e-40cb-bddd-14c085680d2c · outbound

This paper cites Are we ready for autonomous driving? the kitti vision benchmark suite.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Are we ready for autonomous driving? the kitti vision benchmark suite

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.558228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.615737Z digest=sha256:abde6a3a663d500d851d8cfcca2e4b9c6e086a17f298d515a072f6179e8049e3

Observation 963710cf-9f38-43ac-9697-a4faa7d1eb9a · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.619809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.619809Z digest=sha256:c5718b91804bc8b0939a640c9dcca0a1b6c10df945e04879f162d12b9911b270

Observation 5853c801-4a09-4244-a610-be3722600d4d · outbound

This paper cites Masked autoencoders are scalable vision learners.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Masked autoencoders are scalable vision learners

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.535650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.624025Z digest=sha256:106da55386f8448eecbd27ec3c83ee3de5f77b60664f9e0c7429bf3721e1ceb0

Observation 0bb241a8-9b5a-49e6-a74c-b4964909fd04 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.628147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.628147Z digest=sha256:7d0159e7b97c8b4a85cf730000dd386eecc0357a88df104ceb1e69786d8edd45

Observation ab3e944d-9ffc-41be-8881-0afa42753c6a · outbound

This paper cites Open- clip, 2021.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Open- clip, 2021

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.632345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.632345Z digest=sha256:e25aff65b49b69d779a48a3c2780b14f3fe247d407b1797a1f8c87f9b8651304

Observation efad3a13-2355-4f59-857a-bcf21ed9e00d · outbound

This paper cites Novum: Neural object volumes for robust object classification.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Novum: Neural object volumes for robust object classification

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.636714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.636714Z digest=sha256:7b55bc766f72679ac40f15ac462105426d6f837f6dc60ac748c487e6fcb67756

Observation 0a18a1e3-b8c5-4f12-919f-d9372cfeb9b4 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Scaling up visual and vision-language representation learning with noisy text supervision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.641122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.641122Z digest=sha256:972296b6c72f73c51dc72ddf9f4df01a55843ca79fd0116606e4b00fa258eff5

Observation 8b70a587-e6b4-442f-a52b-039d3219ebb4 · outbound

This paper cites Mistral 7B.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Mistral 7B

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.645314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.645314Z digest=sha256:86d0d05f49ad987d9b4c11cacde25f8f19d1706fc67d259610c1b957b56578a2

Observation 23def1e4-6e16-456d-a272-19374143cbe0 · outbound

This paper cites Perspective fields for single image cam- era calibration.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Perspective fields for single image cam- era calibration

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.496761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.650512Z digest=sha256:fb4c3507e320e0d08c727a1ce2f0487d07fbdbb8277baf2338adffda68c2ee55

Observation 0f29b4b2-cd60-4d8a-bccf-fe8980c20c1c · outbound

This paper cites Segment Anything.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Segment Anything

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.654030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.654030Z digest=sha256:30bbe624bedd385318eee75e5187fae15ae60426514937b20d80654ce4b7819d

Observation d6ad7280-ab4e-4084-8836-82d3c4087e82 · outbound

This paper cites Segment any- thing.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Segment any- thing

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.482808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.657739Z digest=sha256:cdd8272d015fef6fd07df443ebf71e07b1d1c8340832e29fade8808da69f576d

Observation e8503dba-ab6d-41a7-ba90-27a7802de402 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.468752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.661340Z digest=sha256:620c05e407294365087920f6b218d4e0064cf17565e1f8bc52a792e441c3324a

Observation 6133f4d5-e8dc-43b8-b982-5d7bc1ba9234 · outbound

This paper cites an unresolved cited work.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:40:28.455596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.664986Z digest=sha256:6a7de45b160b9404b6073a9417c27e8b7a37c7f7317fcc4327645191d807b9a8

Observation ec3d869b-35a5-4bf8-a173-9680cbc1cfd0 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.668620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.668620Z digest=sha256:aa752c972f54a340ae229a3b1f835deea02a172e2b34292d39a41bcc6278c4e6

Observation da067013-52b3-4923-b835-b331d5dff1cf · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.672593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.672593Z digest=sha256:999628bb0b0b9a7360ba17c28068301b34a76a7c3bdda5eb48d10de4975e22f7

Observation 6a9eeaed-3940-4f5e-bf21-6cac2194feff · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.676116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.676116Z digest=sha256:573975d4fc34babc55a75489771f567e8bfaba10c44bf06934fb2158014f4a5e

Observation e191db48-ad61-4223-9dab-3e802c457a8f · outbound

This paper cites What If We Recaption Billions of Web Images with LLaMA-3?.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models What If We Recaption Billions of Web Images with LLaMA-3?

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.679528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.679528Z digest=sha256:15ca8fae9ec36abc52d9d762a35317267d617161b955dea401e21a748a395276

Observation 2d8537d9-59b3-4597-bf88-eb1310a983ab · outbound

This paper cites Learning customized visual models with retrieval-augmented knowledge.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Learning customized visual models with retrieval-augmented knowledge

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.424103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.683992Z digest=sha256:5d054409c593567f144ed059df4ca9ce806002bc773d4f7bee4acb575b3d5b5c

Observation b5370f09-d4d9-4e65-823c-dd477f4ebdd9 · outbound

This paper cites Improved baselines with visual instruction tuning.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Improved baselines with visual instruction tuning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.411147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.688146Z digest=sha256:5b510f6fcdfd416ab7186ea2410f86ffbd23a4cd1e18aa022ca48c4d084fcd99

Observation 8563a318-bdce-4f59-94e2-8c66228b8f8b · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.692491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.692491Z digest=sha256:38b44c791a693a918e282e7766a275b177841a075cdc104d0b6e3a7aee1f5c46

Observation 9078385e-9047-431d-88a5-8539b76bdac5 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.389689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.696559Z digest=sha256:07ddf7b83482c628a691a2cec57809ca6e198f417f89f41c19ff90a9a73e777c

Observation 2d40ee3a-6590-4210-971b-7cb5e430e012 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.701110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.701110Z digest=sha256:3bf06d31a7816f1e43190c28484b7459a6c70643ec30ac337efd50966749efc1

Observation 1b4f0ab1-d640-46cc-b582-d973fb3af87f · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.706213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.706213Z digest=sha256:99405dfbf03563bf069865ed7bc269860c64ce9959abe6c70586efbc6e523455

Observation 9d41879e-6ea7-45c5-9dd6-5fb92975967a · outbound

This paper cites Robust category-level 6d pose estimation with coarse-to-fine rendering of neural features.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Robust category-level 6d pose estimation with coarse-to-fine rendering of neural features

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.710906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.710906Z digest=sha256:7750a4c1369fddba1b88cadfc922fe2bf19aa43beb4135ee84b7c50f8952cebc

Observation cb3e3658-1c82-496c-8a94-595420f171ad · outbound

This paper cites Imagenet3d: Towards general-purpose object-level 3d understanding.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Imagenet3d: Towards general-purpose object-level 3d understanding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.367631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.715210Z digest=sha256:45c152153f1769c26c02bdb957b37377352d79f31425564f01b2bc686a9ddf2a

Observation ee39a5d8-b409-4f20-afb8-66877bafd94d · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.719431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.719431Z digest=sha256:a8ce9828ac4b1aa2433db069bb29352a325d3c607dd2f8297108fa0ab9661c3e

Observation a0f212ba-c183-4ba0-ac22-e03927a53aff · outbound

This paper cites Dinov2: Learning robust visual features without supervision.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Dinov2: Learning robust visual features without supervision

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.353857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.723958Z digest=sha256:bba1f46da3e314e5e2cfe733f4dd3f11cd6fc67cadc5d0fd3120c649c8a9144d

Observation 46c97bf3-ed26-462e-8532-8f29b56a2c81 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Learn- ing transferable visual models from natural language super- vision

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.340434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.728221Z digest=sha256:feed95ca26c53ce8e98a3e36f9497237e666d3fd31c52539db9959b36e76f811

Observation ab98a02a-21ff-4a25-aa69-358024560e93 · outbound

This paper cites GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.732388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.732388Z digest=sha256:6f7a6d1181b817064a2aa22601f0ecbf0b9ec1e39c266b90aacbf32cd49edef9

Observation 5ce633a1-7150-4fe4-b5d2-0c05685ba58f · outbound

This paper cites Susskind.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Susskind

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.737271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.737271Z digest=sha256:5c83a9c689c7f837cd6b039dd3f6eb3cc6c58f4c8777a9f05383587046ec1267

Observation 0480b152-9d4f-456f-8a1f-4cf1911f0bd2 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models High-resolution image synthesis with latent diffusion models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.742087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.742087Z digest=sha256:d543a6974e78efd417ce85e11a703e1ce03d02935bfaebd679de6fbbb4c37481

Observation 006893b9-ae58-480a-858a-34eae99b0e66 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in Neural In- formation Processing Systems, 35:25278–25294, 2022.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in Neural In- formation Processing Systems, 35:25278–25294, 2022

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.746271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.746271Z digest=sha256:0bdb15f2305a77de3d0983c5b3e0ffa45ac56813d75cd98cf107f995ccfe3c66

Observation d9eed4e5-c451-414b-9f1c-d369abc0d3cf · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.750251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.750251Z digest=sha256:4aa297c29c808464e3cbb5fda5c43a73bf613d2f07794ba66ed8c1ddb31f61b6

Observation 38d7f83c-8884-48a2-a43a-5028083c44e9 · outbound

This paper cites Sun rgb-d: A rgb-d scene understanding benchmark suite.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Sun rgb-d: A rgb-d scene understanding benchmark suite

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.292913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.754431Z digest=sha256:f207523883217f18ec11f584971a422250e0ce06526429d0259d759c8acfe6e3

Observation 5079e0ca-85ec-40d6-8dd8-fc7220e45669 · outbound

This paper cites Core knowl- edge.Developmental science, 10(1):89–96, 2007.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Core knowl- edge.Developmental science, 10(1):89–96, 2007

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.279345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.758501Z digest=sha256:21b14e337349911e72ee54466bae2d436e78c51812ee115d65c98d30c5b7737d

Observation 2d34e57a-31d7-4600-b65c-438238908d63 · outbound

This paper cites Revisiting unreasonable effectiveness of data in deep learning era.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Revisiting unreasonable effectiveness of data in deep learning era

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.266026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.762558Z digest=sha256:e8a97e7d88516adca4c265d4340f20449041fd7ac1ba344d66f2c267f3a59ac2

Observation 136a0093-736c-4ee3-81e0-c392bb18cacf · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Gemini: A Family of Highly Capable Multimodal Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.766525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.766525Z digest=sha256:89d32ef81845a78dd2145e7c8285bdaf1fcd03575c9ba2e15d99f6f308c64d21

Observation 7432001d-b856-4980-8200-07a6a7a57b7b · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.770707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.770707Z digest=sha256:83afff03422f1d0d72ba9c366b9618a08e13cd6c4672dcde3c21709c8075c777

Observation 58112f68-ab1b-4fd0-9de8-9a674ab0493b · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.775308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.775308Z digest=sha256:364d18696b3960d92c7c8e128f536a1eea429c198f1bd9ea70a6999ece63aa11

Observation 1cb36f74-af69-48b2-b64b-007267717ddc · outbound

This paper cites 3d-aware visual question answering about parts, poses and occlusions.Advances in Neural Information Processing Systems, 36, 2024.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models 3d-aware visual question answering about parts, poses and occlusions.Advances in Neural Information Processing Systems, 36, 2024

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.252560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.779103Z digest=sha256:5db08e4ecf7af771867abeab79f77408c8804d273c05d1dc42491bda83f0a0af

Observation 87ad28bf-6065-4e22-a3aa-3ad93f02c52b · outbound

This paper cites Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.782481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.782481Z digest=sha256:4b75401198466537893ca90b32e76da31fb3992b1e4f0ad085f2213bcc9bccd7

Observation 9413468e-6207-41cb-a66c-0ba8a4f1ed2b · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Depth anything: Unleashing the power of large-scale unlabeled data

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.786283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.786283Z digest=sha256:4757fd8059dd841c8d8b69278b1d1de13a541b91942bb448a651524d93d217e4

Observation 5bcdd412-64c8-4b6c-bae0-b479654111c3 · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Depth anything: Unleashing the power of large-scale unlabeled data

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.789631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.789631Z digest=sha256:9758e5216f0c9862628dcef118eb2561862913662bd079fc3d665751a4949155

Observation 641aa501-4832-4c54-a095-8c145200b324 · outbound

This paper cites 3D Question Answering.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models 3D Question Answering

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.793012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.793012Z digest=sha256:2cd0a22e5bcac74585529b34314fab605e973abaecb43cc62c12de43e4f8c4a0

Observation 0a28dadf-2403-4a41-b966-8c54e4654bec · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.796715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.796715Z digest=sha256:269431f02b3ac551021e5bcdbd210cb3fae59100bc9c75951e826d85f5129c2a

Observation c0976f13-42b8-4a8e-b0c9-79e2abede03f · outbound

This paper cites Recognize Anything: A Strong Image Tagging Model.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Recognize Anything: A Strong Image Tagging Model

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.800978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.800978Z digest=sha256:0c9762830c959d091564fde69b234e8088cfabd1161c991d83248f8a4ce83f44

Observation 05828b61-7e53-4419-969c-22fdbc6382e6 · outbound

This paper cites Recognize anything: A strong image tagging model.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Recognize anything: A strong image tagging model

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.222985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.805680Z digest=sha256:9223002c7ba9d706237f6e2ab3d6ab30315476b9dbf093d762c416c12a3c5ea8

Observation c6ea85e2-6324-4c84-89cb-3e5cc5a09a3b · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.210551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.809925Z digest=sha256:c8502af312a4ca7827fbe25b5eac5bd93e189848eba2a3604362895f84d75ed2

Observation 3a2932da-f0a4-413f-a27c-6f50635ac041 · outbound

This paper cites iBOT: Image BERT Pre-Training with Online Tokenizer.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models iBOT: Image BERT Pre-Training with Online Tokenizer

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.814069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.814069Z digest=sha256:2ec87f88b27196e8b1798d970cf006f64f07b61e9843ba165cb3035525a3670c

Observation 40f19090-a735-4d0d-a78f-2b57a9580da8 · outbound

This paper cites yes” as the answer and 120 questions have “no.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models yes” as the answer and 120 questions have “no

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:40:28.195914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:40:27.818667Z digest=sha256:d3c7f5f8c753115f843d669b171d273ed040bc66f1bbee8b32264cfd1a7f4871

Pith citing papers

Observation 2eeb54c1-f742-4702-aacc-667cb7965203 · inbound

SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL cites this paper.

SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T18:42:12.968388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:42:12.968388Z digest=sha256:e3cc14e1ba165f1573cef3191cd4d3a235b2845a0ef3057ac174a13cfa77f47b