Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:22:11.053165Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 3 inbound Pith citation observations for arXiv:2505.04965.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:22:11.053165Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T13:00:00.925521Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-12T07:11:26.075563Z
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f288f369-cc55-45e4-acca-bb150dc631ab · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4251ca5d-7396-4451-8bff-d158bf5ed10f · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7bddf88b-0f75-43f9-a69c-0210b1b217ff · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Scanqa: 3d question answering for spatial scene understanding
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9713e6d0-eea1-4c96-9113-d949937aaef3 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a597cbf1-e571-4971-a419-c34b23e974d3 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding End-to-end object detection with transformers
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7d0bde6f-7e5c-45b3-80ce-d2faa052f357 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Matterport3D: Learning from RGB-D Data in Indoor Environments
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad7e99e0-51ee-43a4-a33f-89429b63c608 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Mikasa: Multi-key-anchor & scene-aware transformer for 3d visual grounding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cb6963df-46a9-47e6-8b20-c027afde004e · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Scanrefer: 3d object localization in rgb-d scans using natural language
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 42b599af-e5e0-4ae7-a52e-349d8306890a · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding End-to-end 3d dense captioning with vote2cap-detr
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3f2e2968-c6de-4608-8a1e-080400d10b10 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Real-Time Referring Expression Comprehension by Single-Stage Grounding Network
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71aec554-7ade-416a-9953-04dc80cf634d · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Scan2cap: Context-aware dense captioning in rgb-d scans
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3f1efb7a-1c9e-4555-863f-5c4873d38b7d · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Exploring contextual modeling with linear complexity for point cloud segmentation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9876815b-fa24-4792-a20c-3a4ff1e58826 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding 4d spatio-temporal convnets: Minkowski convolutional neural networks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35ec4a3d-68b3-4a82-bcc6-1a60072272be · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Scannet: Richly-annotated 3d reconstructions of indoor scenes
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation abde5baa-9a2e-4bd8-9858-6034120d1563 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Data Augmentation using Large Language Models: Data Perspectives, Learning Paradigms and Challenges
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc160f4a-20e8-406f-99c4-ad842558a5e0 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding GLM: General Language Model Pretraining with Autoregressive Blank Infilling
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fff69e8-67e9-41aa-aceb-5e8cd25740cc · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Diversify your vision datasets with automatic diffusion-based augmentation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a3a925fb-d63f-450b-a6bf-abae006dce3b · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Naturally supervised 3d visual grounding with language-regularized concept learners
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 165d4d81-6592-449c-af7b-62b78f0ee706 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Viewinfer3d: 3d visual grounding based on embodied viewpoint inference
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 25f9993a-03c8-4dc4-987a-e4cc055236c8 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Viewrefer: Grasp the multi-view knowledge for 3d visual grounding
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation db16b6e6-ad9b-41c2-90e9-ba8546c09857 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Deep residual learning for image recognition
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2cee5e1-95b3-46ec-92d9-a937226617e8 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding RefMask3D: Language-Guided Transformer for 3D Referring Segmentation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e050536f-246f-46c6-8aaf-da812f6dca6b · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Chat-scene: Bridging 3d scene and large language models with object identifiers
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bb067992-5177-40eb-a828-3ce133f95f14 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Segment3d: Learning fine-grained class-agnostic 3d segmentation without manual labels
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b35acb44-5fdf-4027-b599-4540ab3fa56c · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Training an open-vocabulary monocular 3d detection model without 3d data
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5005e507-dbae-48e3-a179-809fa931302f · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Multi-view transformer for 3d visual grounding
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1ebf1b95-300e-4cd2-97f4-bfb3d49495cf · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Bottom up top down detection transformers for language grounding in images and point clouds
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 12fc7542-c801-4fcf-92ac-9643ac6d2871 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Pointgroup: Dual-set point grouping for 3d instance segmentation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9a169e3d-fac9-4203-b343-2fbf2eb26ca1 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Q: How to specialize large vision-language models to data-scarce vqa tasks? a: Self-train on unlabeled images! In CVPR, 2023
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 493215af-6987-4018-acf5-84f9c21d2b92 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding OpenVLA: An Open-Source Vision-Language-Action Model
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 140639e6-0e5b-44ae-8596-55a56c2f7fb9 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding A real-time cross-modality correlation filtering method for referring expression comprehension
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 14d68064-2efb-44f4-94b7-05035f206b30 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Roberta: A robustly optimized bert pretraining approach, 2019
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79f19f0f-681b-4c9f-b66a-ed1ce2c6f948 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Bevfusion: Multi-task multi-sensor fusion with unified bird's-eye view representation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 519e48c6-5367-457b-9675-8cffdea805ff · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding 3d-sps: Single-stage 3d visual grounding via referred point progressive selection
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5fb46b7a-2e0f-40aa-9d53-f5f1619ac1b5 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Multi-modal understanding and generation for medical images and text via vision-language pre-training
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ea55a265-c636-423b-93df-ac8ae50a9f11 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Instruction Tuning with GPT-4
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b05a041-b4eb-4435-9b3a-a009a198aa98 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Zeetad: Adapting pretrained vision-language model for zero-shot end-to-end temporal action detection
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 359d9bb0-c333-48a3-98c2-71588051899e · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Learning transferable visual models from natural language supervision
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b80dccc8-8fe5-45a6-b1b3-5771bb4ed31c · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding LLaMA: Open and Efficient Foundation Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76b862f8-8088-4bd4-ac28-8796ff9376f1 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Rio: 3d object instance re-localization in changing indoor environments
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c881d1b7-44c6-4b1f-9560-5c9ac4a1b486 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Embodiedscan: A holistic multi-modal 3d perception suite towards embodied ai
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6cf46af8-9b8a-4847-b8a2-53df6057c9fa · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Data-Efficient 3D Visual Grounding via Order-Aware Referring
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b88d02ed-6f69-42c6-a945-98e4965c4d6d · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Point transformer v3: Simpler faster stronger
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8643aaf3-88f0-43fc-ad87-9995c58a503a · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Eda: Explicit text-decoupling and dense alignment for 3d visual grounding
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9370c879-6ad1-45cf-ba9a-ee061b1f89cf · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Exploiting contextual objects and relations for 3d visual grounding
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation eb052e49-edf9-436b-993a-4b753dfe60a7 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Dynamic graph attention for referring expression comprehension
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a7c0c1d3-1f92-4284-980a-d783f7d08ca1 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding 3dvg-transformer: Relation modeling for visual grounding on point clouds
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e7086f8d-467e-4228-b22c-bc9786ff61fd · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding 3D-VLA: A 3D Vision-Language-Action Generative World Model
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab5a6027-62bc-40b4-9d1f-d0c77e5880b4 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Denseg: Alleviating vision-language feature sparsity in multi-view 3d visual grounding
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 65ed72a9-fe48-4934-8431-2b39b3c849c1 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Class-balanced Grouping and Sampling for Point Cloud 3D Object Detection
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfb33ee9-56bd-4b8c-9476-271ce887608f · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Object2scene: Putting objects in context for open-vocabulary 3d detection, 2023 a
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ddb163f1-97f7-4aa6-9e2e-c6421e87e3c5 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding 3d-vista: Pre-trained transformer for 3d vision and text alignment
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3763223f-c46c-40d8-b731-f36d08a6ff4a · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding @esa (Ref
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3d174c6-d0d1-4e03-b211-419fb0e4df05 · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Unresolved cited work
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3da01b25-8898-41de-a7e0-c62c972cc1ce · outbound
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding e.g E.g i.e I.e cf Cf etc vs wrt d.o.f et al i.i.d * 90 [1] 0.5em #1
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 81f17d1f-e9b9-44ab-bf37-731031f23fee · inbound
OpenGround: Planning-based Online Perception for Open-World 3D Visual Grounding DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53605b2c-7b62-4b1e-ba4c-52323e074f0c · inbound
PriorVLA: Prior-Preserving Adaptation for Vision-Language-Action Models DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation de151fe1-60c0-4144-90b5-23aad1274deb · inbound
ThinkAfford: Affordance-Centric Reasoning for Fine-Grained 3D Grounding in Cluttered Scenes DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.