Pith. sign in

Paper Citation Record · LEDGER

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models

As of 22 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 1 inbound Pith citation observation for arXiv:2603.16461.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.16461 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T18:06:09.998507Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-07T13:45:53.346402Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T08:46:26.794017Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved61
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1fcca2d3-6baa-47b7-8752-d99503cc7058 · outbound

This paper cites In: European conference on computer vision.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: European conference on computer vision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.807506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.807506Z digest=sha256:82246441594153aa7eaeca3487f41a3f4b6eddfcc9403e53a9998114cf96be1f

Observation 588d605b-4154-431c-985a-d2060adfd6fe · outbound

This paper cites Qwen3-VL Technical Report.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.811884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.811884Z digest=sha256:9bef036321901b1e14abb0a73ec01b947754fdf75d93bd75feb60c70524a31d8

Observation f15bb0ec-165c-4d4c-a020-de0815af47f5 · outbound

This paper cites Qwen2.5-VL Technical Report.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.815651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.815651Z digest=sha256:5a793d1c11a092939b3752ff4099af1d1451430422596d4d554e2f5704cdb675

Observation 25a8376b-b6bb-462f-8fe2-911d0a8d9ffd · outbound

This paper cites arXiv preprint arXiv:2509.25413 (2025).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2509.25413 (2025)

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.819064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.819064Z digest=sha256:3c121525622472e35cdb5fa6f499afb906df4073aaf8cc1a6b237a3e9a81c791

Observation 47a236d9-50b0-4486-9aa9-2dd88e24b991 · outbound

This paper cites arXiv preprint arXiv:2410.01647 (2024).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2410.01647 (2024)

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.822371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.822371Z digest=sha256:0dbd3eee8d5e9979173f41c619fcd70762947fe47df143750a138aafe8c9f1ca

Observation d91e9bb2-c90b-47e9-995b-1e675a7f17e5 · outbound

This paper cites arXiv preprint arXiv:2603.00912 (2026).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2603.00912 (2026)

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.825904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.825904Z digest=sha256:adc64c1efa60cd2c4632c06f0ec67e2841ac7bb722cb0f8fa5958f3178b26c8b

Observation c9d2f36b-d8fd-44fa-af04-709617b99728 · outbound

This paper cites Advances in Neu- ral Information Processing Systems36, 71862–71873 (2023).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Advances in Neu- ral Information Processing Systems36, 71862–71873 (2023)

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.829387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.829387Z digest=sha256:4a16419fbecd50d0b153199f1d26fa032b1233fafb7e1d979e6ef997db0594d4

Observation e1e9b0cd-778b-47c5-bfb4-bb6f69abcef3 · outbound

This paper cites IEEE Transactions on Pattern Analysis and Machine Intelligence (2025) 16 J.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models IEEE Transactions on Pattern Analysis and Machine Intelligence (2025) 16 J

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.832361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.832361Z digest=sha256:88241f9188b9873d37ef42d5cd0ec0832f177461014a223e2fc2119e531cb157

Observation 54160019-3dae-432b-aaec-93892e3d85f5 · outbound

This paper cites In: ECCV (2020).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: ECCV (2020)

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.835431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.835431Z digest=sha256:d39ae26d481f64d1e0ea9e227f3a1b8a22826d7245d82e97d8b9388d006d9177

Observation f8eb9a3c-6c97-408b-94bb-96d489c796a1 · outbound

This paper cites an unresolved cited work.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.838718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.838718Z digest=sha256:8ff3de769db82bebe097a5ccc3c43a56f746ff4c09ce5a4beea2923fb8255e07

Observation 865af6f7-82c8-426b-b558-9455f2c59ac4 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.841836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.841836Z digest=sha256:0a9f9acc6ca8d5a5006c7c52016dc9000fbb31e071b9fd3faafa47096b1fbcf2

Observation bed21cab-d60b-4396-8cd9-23f130b683c3 · outbound

This paper cites Grounded 3D-LLM with Referent Tokens.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Grounded 3D-LLM with Referent Tokens

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.844829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.844829Z digest=sha256:357b23c8f6f5537938f87c8c7049e65eb9d21148ba1873497c5370107314b6a1

Observation 1f64a274-e475-47fe-b58a-bd512e9662a3 · outbound

This paper cites arXiv preprint arXiv:2510.13800 (2025).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2510.13800 (2025)

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.848516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.848516Z digest=sha256:a91d4ff34a10ccd241c440f1fbc119b750a6df0ae2c821315575291baee6071f

Observation 65e5a41c-1917-4984-9027-4369647c380a · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.851409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.851409Z digest=sha256:3421643b532235a3f4f798870c1e30d3c715a7d1165ac249cdfa6aac64edb39d

Observation 1652069e-2d10-4ccb-b348-c4396a15c58b · outbound

This paper cites In: Proceedings of the IEEE/CVF international conference on computer vision.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the IEEE/CVF international conference on computer vision

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.854350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.854350Z digest=sha256:32009178ed9c0f0ebb0b320445cd527a7e98f75e668acc59e8583f1f7aaecb45

Observation e6760b3f-42a1-4d6a-b51e-d5f753e5cbc7 · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.857411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.857411Z digest=sha256:74b4851af33c1cec84aff6a228929d650d6397ee80dd7fedee65e5087729993d

Observation 3dc177dd-407d-4269-87e4-f8868b2ef40c · outbound

This paper cites Generating Context-Aware Natural Answers for Questions in 3D Scenes.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Generating Context-Aware Natural Answers for Questions in 3D Scenes

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.860375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.860375Z digest=sha256:a5410108c88b73745357f441f421378367abed09a7a6a719e20fe1557f44ddd9

Observation 2ecc57ab-2221-4b00-8e81-684fbf08f0ed · outbound

This paper cites VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.863681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.863681Z digest=sha256:a4a75bc0e7a6322fe44d881df3a0dce1487fe1c9eaec87b55b5d1682bb3d2ab7

Observation 3bde5895-7a6b-4d60-ab7c-f20a72793119 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.866812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.866812Z digest=sha256:ef6a4a3c7e7a868dad614c6d0add52bfb3000494103b0124a0d6ab2d73bc0b0e

Observation 2a134d4a-f2a0-4caa-bd2a-39920d8eadf3 · outbound

This paper cites arXiv preprint arXiv:2601.11442 (2026).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2601.11442 (2026)

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.870167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.870167Z digest=sha256:39d5543d6e2c6833061b2024c25c49c741465468e1ef7fd7a266874fe7412032

Observation 3efb21c4-a27f-4919-9719-3a056ab52254 · outbound

This paper cites Cambridge university press (2003).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Cambridge university press (2003)

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.872927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.872927Z digest=sha256:7ee715a37000b2f876d2b1971b9d06f62d9704fd04e52372633ab70ae2961f83

Observation 198bf949-7719-4d36-8145-8662a659a0c8 · outbound

This paper cites Advances in Neural Information Processing Systems36, 20482–20494 (2023).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Advances in Neural Information Processing Systems36, 20482–20494 (2023)

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.875787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.875787Z digest=sha256:cb1b067d56d603e6ba9d4ae847fa584fb944b2420ee93e6c92ce504a9f12e477

Observation d4faea79-27b4-400f-8d07-ef075bd40540 · outbound

This paper cites Advances in Neural Information Processing Systems 37, 113991–114017 (2024).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Advances in Neural Information Processing Systems 37, 113991–114017 (2024)

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.878556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.878556Z digest=sha256:e10e2587f42d599afe889472da6668202df87f70560bb8bb918c24c5dd29336b

Observation 98843dd9-ff48-49ea-855b-63b32843dc1f · outbound

This paper cites An Embodied Generalist Agent in 3D World.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models An Embodied Generalist Agent in 3D World

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.881685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.881685Z digest=sha256:1b9f3a868b3ab7c6cb320949898bed9c17e7bd7c0dc637f58428288a685ac11a

Observation 900369b8-9543-4269-95d4-364584a09cd7 · outbound

This paper cites GPT-4o System Card.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models GPT-4o System Card

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.884802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.884802Z digest=sha256:1acd11e637d6cd30de5c62f81f8a76700d9802697b96ed9cd0f6d6a750413cab

Observation 23ae736e-4d64-4c03-82b2-9b1a1fc80846 · outbound

This paper cites MapAnything: Universal Feed-Forward Metric 3D Reconstruction.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models MapAnything: Universal Feed-Forward Metric 3D Reconstruction

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.887988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.887988Z digest=sha256:f06d61f6a8a722f22db6a34c1a7d3fab4a572b57e4350cc732a489ea9622ebc5

Observation 66d9e8f8-8ab9-47af-bb2e-9b7deed37f30 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.891673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.891673Z digest=sha256:6862bcf17edb9a6eb6a3c79d2e45ead2a8d1b6d12f05daeac2fc3d487ff71db2

Observation 7c711c68-0cd5-4ef3-98d1-cbe8c254770f · outbound

This paper cites Unified Semantic Transformer for 3D Scene Understanding.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Unified Semantic Transformer for 3D Scene Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.895072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.895072Z digest=sha256:0e51298bd5b7a1c7b24656ec7c5181d7f16ad985f730bf8719b7bc1ce0c13bc8

Observation 6b970326-5b68-49a3-9c94-3448979b1a73 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.898270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.898270Z digest=sha256:3f6a4c96778c0505f6993d5acd60d645bc3600a92f17647948034dde633ffada

Observation 57d1d639-d6e9-4f89-9fd7-6ca76690e7d7 · outbound

This paper cites arXiv preprint arXiv:2510.22706 (2025).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2510.22706 (2025)

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.901428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.901428Z digest=sha256:e2a0d436890f1e86f52c6f88ecf32e525c0c38f12054e90565230680641937d5

Observation 5ddf83bc-5c11-4ff4-8d26-92fcb9e5b908 · outbound

This paper cites In: International conference on machine learning.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: International conference on machine learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.904718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.904718Z digest=sha256:d95d8ee4c530b7a36fe57001b046751786cf42d63170573171a9c6f1e1109ca4

Observation 1c2048f2-77fc-4502-be11-51f8fd6b0096 · outbound

This paper cites Depth Anything 3: Recovering the Visual Space from Any Views.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Depth Anything 3: Recovering the Visual Space from Any Views

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.907725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.907725Z digest=sha256:5529a1f6b389bc30a2598713ab0ada1860455a03cd54b95a9f2fe1efb35f22f9

Observation d1939b48-2b0a-46e4-ab41-826f6277cc82 · outbound

This paper cites arXiv preprint arXiv:2405.10255 (2024).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2405.10255 (2024)

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.910861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.910861Z digest=sha256:34ad0446a8469d199cf2d3c3405b06bc316d00513fc37cfb4bc6f359e4fdde07

Observation 0db83715-bbc9-488e-a1a5-06f8bbdab8dc · outbound

This paper cites In: Ad- vances in Neural Information Processing Systems (2025).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Ad- vances in Neural Information Processing Systems (2025)

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.914100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.914100Z digest=sha256:dbca5fc6d569fbba8d0b924383220187f3c7906b843beacfa77c3f3ded69dfd5

Observation f23dd2db-0b40-466f-9558-fc66b03e51eb · outbound

This paper cites Ad- vances in Neural Information Processing Systems37, 23464–23487 (2024).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Ad- vances in Neural Information Processing Systems37, 23464–23487 (2024)

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.917331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.917331Z digest=sha256:cae44112368513c9ab802744b3d68a2d9ac242db5431d175911fab1bcb26fa26

Observation 9b22861a-8f8a-46fd-b74a-89800da0ddc2 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.920391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.920391Z digest=sha256:b758ad4a33ccdc57b71eb32512f7b07198048cd6aaa2801c88e06e77487584bf

Observation db357836-d2ad-4316-87a1-1dac645ee97f · outbound

This paper cites In: proceedings of the IEEE/CVF International Conference on Computer Vision.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.923398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.923398Z digest=sha256:d0790bad0696be30a5baf28dd3cba946ffff22cfb9894e8a59b4dfa30a7cefba

Observation 0ee15ea0-69cc-44c5-9b2c-8fcee3ae62f6 · outbound

This paper cites GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.926358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.926358Z digest=sha256:94b211959116844cf7b183d50a5e2a942f30fb58e001374952b12bcd853e579c

Observation f15d34ea-ab75-40ff-903a-ad20c3dda213 · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.930155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.930155Z digest=sha256:b1ca635d8836d59c2e9f03a0724d9859a58eb6e9e51271ef0ebb280c263adc6b

Observation eacb1e50-027d-460a-8a57-898cd3310017 · outbound

This paper cites In: International conference on machine learning.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: International conference on machine learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.933431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.933431Z digest=sha256:908ce613251e00b2d19dfd322e83309f01bed5e0b3279c613159894c7ed0cab3

Observation 326b211d-d163-4176-b133-6e05c5826e14 · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.936478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.936478Z digest=sha256:acb63e8ee6f5738c0e32092d88f531b3288f5883e6088504014815210e4adfa7

Observation fd1a76a0-11b3-4788-8627-18c2e0d58dcd · outbound

This paper cites FastVGGT: Training-Free Acceleration of Visual Geometry Transformer.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models FastVGGT: Training-Free Acceleration of Visual Geometry Transformer

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.939650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.939650Z digest=sha256:3d03ca17bf5d00e84e4ea42f9a30c05692198f2da223bb2365b8ab09b6cda952

Observation 9d43fbef-8407-46c1-a296-a98901fcaeef · outbound

This paper cites arXiv preprint arXiv:2505.23044 (2025).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2505.23044 (2025)

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.943321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.943321Z digest=sha256:36a9fb6cd403e3b195e8436b80e0800e58fea636a8d14c56eb5dfcf622a0f31f

Observation 2f6bd746-e000-4fc3-b695-c9b154cf2774 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.946206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.946206Z digest=sha256:3272aa6d4d5676620e2f32cd3043e18eb54096619ce8b5831fd54eaf16d8177f

Observation e8811f6d-b96f-4be4-908d-793fcdd8b17a · outbound

This paper cites arXiv preprint arXiv:2511.18416 (2025).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2511.18416 (2025)

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.949484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.949484Z digest=sha256:68224d583497e0ac7ca96a7e62b8453c8c8e5e594e31a7ea605940607c8a17f0

Observation 0ad40dd2-a235-4352-8aa0-e4fa2126fb4d · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the Computer Vision and Pattern Recognition Conference

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.952498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.952498Z digest=sha256:0c42b1d7ffc40ab811df1a15fd7d164417b0190e52a2e293e358e84e487008c9

Observation f2c85b51-1ce0-4a92-aff2-fab8181a1f78 · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the Computer Vision and Pattern Recognition Conference

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.955607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.955607Z digest=sha256:67108ced9611ed85485cd397c26afe2581bc7af92f69553562e44aaf09bc7a15

Observation 0833805c-8384-4d0d-8a36-13af845644f8 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.958696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.958696Z digest=sha256:03923900dbc1bf49949b89d14e7a6b42b1a9aaee83a5a08ef9faf04b5eaa73e6

Observation 0a629d8d-6cb2-4e7c-86c7-512679084b05 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.961597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.961597Z digest=sha256:1fd242067bfbb19b2a5b11ac6d5199f40e838ea3596120a1009c851f308938c7

Observation 51656720-c56f-4e2d-9d4c-ac62b7aa7b21 · outbound

This paper cites $\pi^3$: Permutation-Equivariant Visual Geometry Learning.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models $\pi^3$: Permutation-Equivariant Visual Geometry Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.964640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.964640Z digest=sha256:0d499ca43b7b5bc093796df1e2962c74c86aa300f997cde02f58cd5bb8c826ee

Observation 68d5ae86-8032-4c38-8bb0-745e36e66c59 · outbound

This paper cites Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.967760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.967760Z digest=sha256:c9d0a3085cd89a956684a86d9327c516cd25680a95a383e633e3b9aed76d0dda

Observation 0a8323a6-c07f-4d7a-b729-77691eb92f5a · outbound

This paper cites arXiv preprint arXiv:2508.11952 (2025).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2508.11952 (2025)

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.970802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.970802Z digest=sha256:fa4ec1d22e66df801095233611d0df5363de9fdf1c7c03045f8570e510b697e8

Observation d383857f-9b51-417f-b0ad-bd974ab490f3 · outbound

This paper cites In: Pro- ceedings of the Computer Vision and Pattern Recognition Conference.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Pro- ceedings of the Computer Vision and Pattern Recognition Conference

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.973850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.973850Z digest=sha256:b4b4d59278f55c0d007411f5cdacbdd702434be842457f11935513e3f4df81f1

Observation 79cf545c-addb-4ace-9c0f-ae7939298575 · outbound

This paper cites arXiv preprint arXiv:2511.05491 (2025).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2511.05491 (2025)

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.976837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.976837Z digest=sha256:7405cd7d2642ad2fafea348292224ea34df6a4ed198b9d9c698e2dcbcadf36b3

Observation 54d77854-2572-4d74-804c-64e05ccdb025 · outbound

This paper cites arXiv preprint arXiv:2601.02281 (2026).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2601.02281 (2026)

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.979716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.979716Z digest=sha256:4374bfdae33fed8372a7238efa3d0edfb9debbb913825ad9a86b4c7eb91e636a

Observation 154fce46-323e-4dc9-9ba1-038e1a5b9367 · outbound

This paper cites arXiv preprint arXiv:2503.22976 (2025).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2503.22976 (2025)

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.982621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.982621Z digest=sha256:316e90c1c5d696b06009b87630fa69ad987f842b611921a875e90e0f9d15b017

Observation f0e2179d-74d7-4342-9ece-3a69ed43b3fd · outbound

This paper cites arXiv preprint arXiv:2505.24625 (2025).

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2505.24625 (2025)

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.985621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.985621Z digest=sha256:fc57d2a8342ad77729e6ce818a71508da66c06775ee3c48f4518d56c70c6e736

Observation f6c37f8d-50d4-47e7-84f0-4a726b9ed11a · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models In: Proceedings of the Computer Vision and Pattern Recognition Conference

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.989175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.989175Z digest=sha256:4b39564f8891bbdcf054575968748860930a431699059209b5d72225b9658e94

Observation 046d1603-6434-4f23-b3ef-35111921709d · outbound

This paper cites arXiv preprint arXiv:2510.25760 (2025) GAP-MLLM 19.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models arXiv preprint arXiv:2510.25760 (2025) GAP-MLLM 19

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.992189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.992189Z digest=sha256:ae70f2c01d8396fadbd9ff19402f3da17140959e87a0e8778302967172b5c7fe

Observation 78606104-cf6a-4120-8cd3-97f27074cf27 · outbound

This paper cites LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.995051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.995051Z digest=sha256:d2fe91ef3649ada820e53b4f25dedcaa1d45ed6fcd62e784fa6c7b44de01750d

Observation 7e7adcd3-e2ef-4b33-af1b-ab9fc07ef5a2 · outbound

This paper cites label"and the point’s 3D coordinate in.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models label"and the point’s 3D coordinate in

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.998507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.998507Z digest=sha256:f17448be6d617045d74f41efad9480d3567ba3bb99ae76837bfcb1f7fda49d52

Pith citing papers

Observation d237b7b1-64a0-4574-a2ca-0b54252f643b · inbound

SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness cites this paper.

SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-21T02:21:33.200827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T13:45:53.346402Z digest=sha256:758e478102b689fed2b7465d1f138ae767eb814089c5d15288e8ef0819ccebb1