Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:25:48.557534Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 2 inbound Pith citation observations for arXiv:2506.05318.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:25:48.557534Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-12T18:56:43.374310Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T10:46:01.488781Z
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6bdcce7d-0b59-4581-bd7d-e0bd413fbd3e · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 100fc6bc-4584-4849-af90-bf20b0f5694f · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Scanqa: 3d question answering for spatial scene understanding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8f7d2a4b-5066-42dd-b54e-045aae48f131 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Scanrefer: 3d object localization in rgb-d scans using natural language
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bb3deee4-3811-4fa5-9fba-4664d28f59e2 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Language conditioned spatial relation reasoning for 3d object grounding.Advances in neural information processing systems, 35:20522–20535, 2022
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8deda4f2-318b-4b76-a7c3-b92d2c9a4db9 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs End-to-end 3d dense captioning with vote2cap-detr
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 041477f3-2ae1-4e60-a619-b58dd87b9c44 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5cb04d7e-a3d3-40a6-8bc7-69660d140e83 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs V ote2cap-detr++: Decoupling localization and describing for end-to-end 3d dense captioning.IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(11):7331–7347, 2024
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cf2dd138-cdb3-4e45-8ce3-5beb2f586286 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Grounded 3D-LLM with Referent Tokens
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30fdcd5a-1517-4913-a1f8-052590027ef2 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Scan2cap: Context-aware dense captioning in rgb-d scans
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b09f00b7-b29b-47c3-9230-584725cf07f2 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 82ef0af7-25ed-44b1-897b-505d4bbbee8a · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Scannet: Richly-annotated 3d reconstructions of indoor scenes
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f4da52c8-f29b-4cf7-bc32-9ff3de54291e · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dcc3e57-1339-43a5-89b4-aa16ea337daa · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bccfe3f7-61f1-463d-b7dc-a8cb42dad842 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs 3d-llm: Injecting the 3d world into large language models.Advances in Neural Information Processing Systems, 36:20482–20494, 2023
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 11268861-15b5-4cb8-8c3d-6479a7bbe137 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Chat-3d v2: Bridging 3d scene and large language models with object identifiers.CoRR, 2023
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1ffbeb72-5633-4309-a665-f5ac60618414 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs An Embodied Generalist Agent in 3D World
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 010ccd29-85dd-41c1-b224-e0d71ea926d9 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Clip2point: Transfer clip to point cloud classification with image-depth pre-training
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ce662654-1fb8-4b90-954d-4814eb2b5d09 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs EPCL: Frozen CLIP Transformer is An Efficient Point Cloud Encoder
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 47142fba-bd28-448d-9d75-888df2427826 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Pointgroup: Dual-set point grouping for 3d instance segmentation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation da47fbf2-09d1-41ea-959e-3625e160701c · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs UniGS: Unified Language-Image-3D Pretraining with Gaussian Splatting
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f4f92c8f-8fc5-44cf-9551-5a63b714e995 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 06a4d3a4-83f8-4397-b289-01023e0bfbc5 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs 3dmit: 3d multi-modal instruction tuning for scene understanding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 32568edd-6148-4a88-a6f9-17591953b7ef · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Rouge: A package for automatic evaluation of summaries
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 62e534d6-3bdc-43e4-8224-e3c23053b65f · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs DeepSeek-V3 Technical Report
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 399c8d97-cceb-4736-811f-4353f632abfe · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0afcfae9-9c24-4e28-88c5-92b07520c2a7 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Openshape: Scaling up 3d shape representation towards open-world understanding.Advances in neural information processing systems, 36:44860–44879, 2023
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b5709a7f-17aa-45f9-b753-dd1ce3348b3c · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Decoupled Weight Decay Regularization
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd4dec3b-1a3b-4b5f-b142-d4d15b470c63 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs SQA3D: Situated Question Answering in 3D Scenes
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8bb490b-426c-4d49-91f5-07d3e3abf549 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Image caption generation using vision transformer and gpt architecture
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 769730cb-d3d3-4aa4-a850-281de509c935 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs An end-to-end transformer model for 3d object detection
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2bfb3ae7-65b9-4250-8d73-eddabd07daad · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Bleu: a method for automatic evaluation of machine translation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b1e9ef53-b409-4ca9-8f1c-77dbf5b48b3c · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Pointnet++: Deep hierarchical feature learning on point sets in a metric space.Advances in neural information processing systems, 30, 2017
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fe6de52-0245-4260-8ada-a0899b626538 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Deep hough voting for 3d object detection in point clouds
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 23d4407d-e036-4616-8166-7cea9b9c1fd6 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Minigpt-3d: Efficiently aligning 3d point clouds with large language models using 2d priors
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 58c4ca19-8cce-43af-8c86-ce2b07162d9c · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Exploring the potential of encoder-free architectures in 3d lmms.arXiv preprint arXiv:2502.09620, 2025
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe3eadfa-8ce1-46d7-8376-4faa6ecd7b70 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs More text, less point: Towards 3d data-efficient point-language understanding
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 28f0aee3-b652-4e60-bcc8-976d1a5e2264 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Gemini: A Family of Highly Capable Multimodal Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8198f18e-47c9-4afa-9541-947ba2f1388d · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Consensus-based image description evaluation
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5dc01e8f-1205-4a86-b19b-341ee28b23ab · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Rio: 3d object instance re-localization in changing indoor environments
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1e8912ea-486e-4f5b-a969-fb59ab188393 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4104cf64-9a79-479a-8a95-aa966c494b5f · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c99e5d3e-9146-45d0-978a-a07c0c56dba5 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks.Advances in Neural Information Processing Systems, 37:69925–69975,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5102ccf4-b033-4b76-bb48-02f6ba6a1321 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Ulip: Learning a unified representation of language, images, and point clouds for 3d understanding
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3acf0840-5ce8-4a52-9c34-604e9fe48975 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Ulip-2: Towards scalable multimodal pre-training for 3d understanding
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0ebbd06f-6495-4e7b-88d9-a63067152ccd · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Qwen2 technical report, 2024
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b22b2eed-7e5a-4d9e-a226-523fdd129e06 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Point-bert: Pre-training 3d point cloud transformers with masked point modeling
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 958efc78-4c13-491e-8879-b37cd0ca0d39 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Clip2: Contrastive language-image-point pretraining from real-world point cloud data
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b5d08211-3d8a-4617-810f-467a6c6bbefc · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Pointclip: Point cloud understanding by clip
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bf31fa98-f0bf-47ed-ba37-04134f3fff41 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Learning 3d representations from 2d pre-trained models via image-to-point masked autoencoders
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 670d8369-f6cd-4e0a-8856-002805f573d3 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Multi3drefer: Grounding text description to multiple 3d objects
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 537570de-39ae-44c7-b1c1-77720a093e87 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb23a157-b6c3-4836-91c4-babda4b2c36f · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs Uni3D: Exploring Unified 3D Representation at Scale
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 264df855-651c-4abf-a942-66b432d95dc4 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56e5b207-47dd-4e9f-958f-b05bc7d1e707 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0715743-8459-4717-899e-aec8b4b8eaa5 · outbound
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs it is to the
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eefd24f7-3dc7-4112-82dc-548b3dead336 · inbound
Robotic Manipulation is Vision-to-Geometry Mapping ($f(v) \rightarrow G$): Vision-Geometry Backbones over Language and Video Models Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4d9868d3-4dde-4524-a927-dbd72da50379 · inbound
CAPruner: Conceptual-Adjacent Scene Graph Pruner for Enhancing 3D Spatial Reasoning of Large Language Models Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.