Pith. sign in

Paper Citation Record · LEDGER

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT

As of 5 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2605.09719.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.09719 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T02:42:09.922300Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact9
  • verified fuzzy26
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e92edeac-4e94-4f76-87c6-b4391a0fadbe · outbound

This paper cites Learning transferable visual models from natural language supervision.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT Learning transferable visual models from natural language supervision

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T22:01:53.549751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:f3d00fff3c82b727c6627feeb3af49e27f8765f249c9a937e005ca41963653b6

Observation ddb2c839-ccc8-48d3-a68d-d30bdf3bd1c2 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T22:01:53.588332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:79257728cb2927df82d852880d8b55b48c903fda429c4f2f9deb4293cf196983

Observation 525f2561-b90f-43f9-97c3-16719d2ae51b · outbound

This paper cites Superlora: Parameter-efficient unified adaptation for large vision models.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT Superlora: Parameter-efficient unified adaptation for large vision models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T22:01:53.625470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:e0d52f12a9c7054db5235534d6ac5780c6766d71ab671ee16d5b5efbe2378db2

Observation 05ddfdf3-00fb-458a-aa00-d13e6dd877d0 · outbound

This paper cites LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:31:25.837388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:2837fd381a4e677e74da8d4176aa62b132aedd150550a19e0a851ce5ce2f4449

Observation 5adcbd4b-9be4-4c71-b461-ad3354c9de6d · outbound

This paper cites Palm-e: An embodied multimodal language model.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT Palm-e: An embodied multimodal language model

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T22:01:53.563110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:83d0d20e6960a69694aee43cb35c870a7bb3ef5d0f40822647262744d2e1a7b1

Observation 5be164d3-752e-442c-96f6-7562dc3ef2d8 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT RT-1: Robotics Transformer for Real-World Control at Scale

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:31:25.817150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:5c6d4819375ba40c82f57503882d6a83296f448611dabab7d49c00cb7d463177

Observation 03af8fc4-5078-4c72-9a37-4e7a4568dae6 · outbound

This paper cites Energy and policy con- siderations for deep learning in nlp.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT Energy and policy con- siderations for deep learning in nlp

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T22:01:53.555522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:29b960413d11a23785f1de7d1136ee95a14cd359c532ce002bf193927b1deaef

Observation a0717ceb-85e3-4d71-b6f8-a4c976ba5f8b · outbound

This paper cites Mofa: A model simplification roadmap for image restoration on mobile devices.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT Mofa: A model simplification roadmap for image restoration on mobile devices

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T22:01:53.599268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:4f0757abb8e502cc1cc62fc43bc0234ded9b9fe85803b248b50f27cd33e0e4a3

Observation b70e9e08-eb19-47bb-9a73-cce51385b84a · outbound

This paper cites Learning efficient vision transformers via fine-grained manifold distil- lation.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT Learning efficient vision transformers via fine-grained manifold distil- lation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T22:01:53.579260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:bec205c2485e89d05e97802ba2f4c08f4b2324761039aff4190e4d9711b3ae4b

Observation 867bcbf9-6916-4fd2-aa8f-67e6d9da22f3 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T22:01:53.569673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:35e688e4baf30a762c3846ec0dc62b62944a51ff74395b2030a045050fea2647

Observation 662e8702-89fd-43c8-b470-af7fa7a165fa · outbound

This paper cites Knowledge Distillation in Vision Transformers: A Critical Review.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT Knowledge Distillation in Vision Transformers: A Critical Review

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:31:25.793868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:49a7fad0772bd805661cfb61b50df374bcd572a96f40e421c63b45ec1117485d

Observation 0ea23fa1-5629-47a3-9ea8-95a88aaecce3 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT Distilling the Knowledge in a Neural Network

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:31:25.782086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:76583159f13d0f05d39834558bb0184d30bb731a6364a29c9d288a7182a6d9ea

Observation 925184fb-6836-4ae4-8527-70562bd4d48b · outbound

This paper cites Knowledge distillation: A survey.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT Knowledge distillation: A survey

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T22:01:53.585185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:f204983a1807016a71d1a27f1229135d3090d8799be03b367d26f82de807b4fb

Observation e25fc2f8-f1a3-4f17-9d91-5120e8c4284e · outbound

This paper cites Vggt: Visual geometry grounded transformer.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT Vggt: Visual geometry grounded transformer

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T22:01:53.559507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:fa80fe4396ebf23b84aa950a99f5f0c193c50679764c2a09fc57fad093c8d342

Observation 618d2b07-ca86-4709-ae9d-ece36449725b · outbound

This paper cites Pf3det: A prompted foundation feature assisted visual lidar 3d detector.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT Pf3det: A prompted foundation feature assisted visual lidar 3d detector

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T22:01:53.553031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:3d7f1422f16f9c8a9ed3aa198a7aa2303ee20e2635e6ca80f7ad6edcf2d50c2f

Observation 6a03fc50-7737-40db-89a7-ae434cdd3350 · outbound

This paper cites Visual instruction tuning.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT Visual instruction tuning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T22:01:53.591433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:53ff82bc1eaaed307c25a94280eb255bf0364ec1733a8fe997012fd18a0a10f7

Observation 20a36fce-7480-4358-a812-70cbe8a1c51c · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T22:01:53.576326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:bc2570fe5e9be74c6e341e4a988536ed43c09d5d6eda9d8c5a0b38d822aee306

Observation 91293638-1e97-4309-b372-4bcd8c441f33 · outbound

This paper cites Vqa: Visual question answering.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT Vqa: Visual question answering

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T22:01:53.572751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:9533fe7690b630c32e2870db26e8632669e2a66659abc55b4fadf2bacc3d2b29

Observation edae610a-5608-49f8-b3ab-43fe9fb6b450 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T22:01:53.608767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:4a461e9f0e735439e97c259a462eacf8aebaa04cb0c6a5f26e4c8260831ed179

Observation a3bdaee0-b542-432f-995c-81c4087057a2 · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT 3d-llm: Injecting the 3d world into large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T22:01:53.582327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:9cfd22a67030859944db86182df1f8aa05e8c2f595a3665135fc99cbe80acbda

Observation 5ddb7169-7dcb-4369-a3f1-e10f6a854a03 · outbound

This paper cites Distilling vision-language models on millions of videos.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT Distilling vision-language models on millions of videos

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T22:01:53.595574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:dd9ef561340b72d7a8aa67ad91e1ba7031e7d7c2ecc6c53ce492bfb1762823f9

Observation f48b52ab-4f4b-4fb0-a155-076208245d16 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT Chain-of-thought prompting elicits reasoning in large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T22:01:53.631866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:682e9195ad3cccae811b3944dcbd42b5eb0e1799530177d6fe52db321b647bd4

Observation 040dcedb-fba4-47a6-bdfd-92ea217dbded · outbound

This paper cites Large lan- guage models are zero-shot reasoners.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT Large lan- guage models are zero-shot reasoners

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T22:01:53.629063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:7419ad521005e2d314c67bca4b92a447c715ce995c785dcad95a72eee392fafb

Observation bd00bbb6-2003-42be-9397-05e27e3274be · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT Scanqa: 3d question answering for spatial scene understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T22:01:53.612338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:4d47a65acccc318621be2aa492ec341a03bef07ae1a6c20bcbe98106552f454e

Observation 7826bfbb-6c62-4c52-a68a-ac00e0869521 · outbound

This paper cites 3d-sps: Single-stage 3d visual grounding via referred point progressive selection.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT 3d-sps: Single-stage 3d visual grounding via referred point progressive selection

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T22:01:53.605490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:932faf2020c7c8ae585ca5931aa3fcb1cff38ede18d69fb621aae75d583581bc

Observation 48d7229e-af7a-4dea-adcc-ac08d9138960 · outbound

This paper cites Embodiedscan: A holistic multi-modal 3d perception suite towards embodied ai.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT Embodiedscan: A holistic multi-modal 3d perception suite towards embodied ai

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T22:01:53.618681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:fb70a353191fd36058295f3f818be59518044e15d1e7f56c91bfb8b6d8b61692

Observation 0978cd05-e33f-4d50-a9ac-d118983477c2 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT LLaMA: Open and Efficient Foundation Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:31:25.786761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:f06ce83a46acb2c0c49c0cb224ed2ce5c05fc9d84972286c056504e502d88221

Observation c31dc4fd-7d8e-495e-8093-c25fb4d29b3f · outbound

This paper cites Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:45:05.263509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:837b1a713d19a877d1b0cf8c14b244ef03d677090beb4fb697963b4c81a058b4

Observation 7e9b0384-ddb9-42dc-a6f0-6f7487a71550 · outbound

This paper cites Multi-task learning using uncer- tainty to weigh losses for scene geometry and semantics.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT Multi-task learning using uncer- tainty to weigh losses for scene geometry and semantics

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T22:01:53.615682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:bdbb3e74e3b20617c1159b90aed3582ecf7f4cdaffa30d230c1aff9790591b8d

Observation 305e958b-9484-47a3-9910-51df8bdcba37 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T22:01:53.566862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:43c20c0af8adfe8e0c9eff20cb710163dcb6bbdadb9b8cbaf9228483edcdd3d9

Observation ba99407b-ae00-4aae-b5f8-0f22e57927a7 · outbound

This paper cites Midi: Multi-instance diffusion for single image to 3d scene generation.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT Midi: Multi-instance diffusion for single image to 3d scene generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T22:01:53.622240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:9fb82f6efbc2c13298ea4be454262b0285d348289ad158b17598ee1abf550511

Observation 25791731-a687-4d60-9f62-bb59acbb3940 · outbound

This paper cites SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:31:25.822356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:53a200a3d76ff96efde74a403499f30c4cb6a124125192b510ed3b86114c0197

Observation 0bc4c46f-7df8-4b1d-a268-eed16a45a5f9 · outbound

This paper cites 3dsrbench: A comprehensive 3d spatial reasoning bench- mark.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT 3dsrbench: A comprehensive 3d spatial reasoning bench- mark

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T22:01:53.602712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:0af703b1c6203fe2a8853ef599ae29c5815aafbfc647781ef6d5c01a9b1b95d8

Observation 649bd923-e9ca-44a5-82c5-65922f10d106 · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.347525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:5571157db9aa201180adbc276fbab034daa01335f220ae7f8bd53bb46a6bfa18

Observation 8df35154-8f87-4997-9cb5-5fc8610a04a6 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT PaliGemma: A versatile 3B VLM for transfer

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:31:25.802386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:49aab8d88cf9ed0aef9675bd918d18f8f09c42b48ce6681e093fad93f3f1d852

Pith citing papers

No inbound Pith citation observations are available.