Pith. sign in

Paper Citation Record · LEDGER

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding

As of 22 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 3 inbound Pith citation observations for arXiv:2505.04965.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.04965 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:22:11.053165Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:00:00.925521Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T07:11:26.075563Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact2
  • verified fuzzy35
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f288f369-cc55-45e4-acca-bb150dc631ab · outbound

This paper cites GPT-4 Technical Report.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:10.829066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:22:10.829066Z digest=sha256:0abedf74ac241bcf262ecf05acb90612375d26c2dd172313e034d8c466371825

Observation 4251ca5d-7396-4451-8bff-d158bf5ed10f · outbound

This paper cites Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.852385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.833791Z digest=sha256:25cb981b8083a6430c5abf5b24ee31a3ba5be1f10962bd4879da6902e8273204

Observation 7bddf88b-0f75-43f9-a69c-0210b1b217ff · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Scanqa: 3d question answering for spatial scene understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:10.837723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:22:10.837723Z digest=sha256:9af038ae3e43e5632dbf20471d7cc21618d61e9e01d1563daede370c939f2e30

Observation 9713e6d0-eea1-4c96-9113-d949937aaef3 · outbound

This paper cites 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.826821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.841588Z digest=sha256:603461b7ccb703689301f6439f9158c0bcca383adc010e5f8debed38bbd859d5

Observation a597cbf1-e571-4971-a419-c34b23e974d3 · outbound

This paper cites End-to-end object detection with transformers.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding End-to-end object detection with transformers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.810566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.845459Z digest=sha256:a5675fe7cac6446440080fcafa59f8924db0f78e2fe8d729be000b8bf5848372

Observation 7d0bde6f-7e5c-45b3-80ce-d2faa052f357 · outbound

This paper cites Matterport3D: Learning from RGB-D Data in Indoor Environments.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Matterport3D: Learning from RGB-D Data in Indoor Environments

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:10.849534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:22:10.849534Z digest=sha256:e14a450126ff9eeb2b9a5885bb5726f772a43b7118591d860d129c1f26eb745c

Observation ad7e99e0-51ee-43a4-a33f-89429b63c608 · outbound

This paper cites Mikasa: Multi-key-anchor & scene-aware transformer for 3d visual grounding.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Mikasa: Multi-key-anchor & scene-aware transformer for 3d visual grounding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.795679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.855038Z digest=sha256:232df141b0b1054d64ac24a9a91b793232fd4da483dd4d8b0bbcc79859c557a7

Observation cb6963df-46a9-47e6-8b20-c027afde004e · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.780645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.859257Z digest=sha256:7762cb7422b34c47b329e5b8edddb9fe3670932e2e6312fb28b610dda569e0ba

Observation 42b599af-e5e0-4ae7-a52e-349d8306890a · outbound

This paper cites End-to-end 3d dense captioning with vote2cap-detr.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding End-to-end 3d dense captioning with vote2cap-detr

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.765373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.863434Z digest=sha256:19efdc215face26aba4c282a593143e72ab4ef64d98aab1da74b85e7e4b75b3e

Observation 3f2e2968-c6de-4608-8a1e-080400d10b10 · outbound

This paper cites Real-Time Referring Expression Comprehension by Single-Stage Grounding Network.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Real-Time Referring Expression Comprehension by Single-Stage Grounding Network

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:10.867686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:22:10.867686Z digest=sha256:f06c98bb8df578e846169eaea0a44c0801341b46ca4f30b90d4c74128e1b6261

Observation 71aec554-7ade-416a-9953-04dc80cf634d · outbound

This paper cites Scan2cap: Context-aware dense captioning in rgb-d scans.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Scan2cap: Context-aware dense captioning in rgb-d scans

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.749888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.872103Z digest=sha256:5c19e7f42a2f7fe9195d14582d709a7e3fb58e2d3587bc3c9e0415e49318b2a0

Observation 3f1efb7a-1c9e-4555-863f-5c4873d38b7d · outbound

This paper cites Exploring contextual modeling with linear complexity for point cloud segmentation.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Exploring contextual modeling with linear complexity for point cloud segmentation

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-15T23:22:11.221823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.876307Z digest=sha256:49ba852c5f2991c34abf5bdc19b06521ce212a8df9e566b1a7ce0c371b993a51

Observation 9876815b-fa24-4792-a20c-3a4ff1e58826 · outbound

This paper cites 4d spatio-temporal convnets: Minkowski convolutional neural networks.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding 4d spatio-temporal convnets: Minkowski convolutional neural networks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:10.880888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:22:10.880888Z digest=sha256:7d2a1cf7b6d86f615937442a03cc9bf95a554c57481a03861ac087516e202af2

Observation 35ec4a3d-68b3-4a82-bcc6-1a60072272be · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.723649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.885329Z digest=sha256:e6142d2dec1630e33d1db75eb207774d4a9963a5aeefbbbf8bc065799e07d181

Observation abde5baa-9a2e-4bd8-9858-6034120d1563 · outbound

This paper cites Data Augmentation using Large Language Models: Data Perspectives, Learning Paradigms and Challenges.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Data Augmentation using Large Language Models: Data Perspectives, Learning Paradigms and Challenges

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:10.889544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:22:10.889544Z digest=sha256:ed67ab31c66f7180a367c6e79541eae9959b6004c4c54e375987ac9dc5334a1c

Observation bc160f4a-20e8-406f-99c4-ad842558a5e0 · outbound

This paper cites GLM: General Language Model Pretraining with Autoregressive Blank Infilling.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding GLM: General Language Model Pretraining with Autoregressive Blank Infilling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:10.894178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:22:10.894178Z digest=sha256:0c10b8b79650a8c85b59546f6c83ec5beb11ce9935eeab99ef617c17d9d5683f

Observation 2fff69e8-67e9-41aa-aceb-5e8cd25740cc · outbound

This paper cites Diversify your vision datasets with automatic diffusion-based augmentation.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Diversify your vision datasets with automatic diffusion-based augmentation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.708547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.898993Z digest=sha256:5f209d97da502876dce2614c533f1860cceed9cd3ec4b32e7374bfb6878f2ba5

Observation a3a925fb-d63f-450b-a6bf-abae006dce3b · outbound

This paper cites Naturally supervised 3d visual grounding with language-regularized concept learners.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Naturally supervised 3d visual grounding with language-regularized concept learners

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.692368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.902932Z digest=sha256:bb523f6b02ee31585cbe95a44dd3b17de56f857115f65949aab6ed7d2316f797

Observation 165d4d81-6592-449c-af7b-62b78f0ee706 · outbound

This paper cites Viewinfer3d: 3d visual grounding based on embodied viewpoint inference.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Viewinfer3d: 3d visual grounding based on embodied viewpoint inference

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.676866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.906972Z digest=sha256:d1cec274547e5a78a8cd9f0c9491e93e976fba0c138f8ca8915ce3bd899d9210

Observation 25f9993a-03c8-4dc4-987a-e4cc055236c8 · outbound

This paper cites Viewrefer: Grasp the multi-view knowledge for 3d visual grounding.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Viewrefer: Grasp the multi-view knowledge for 3d visual grounding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.662346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.911013Z digest=sha256:6f547a66b7b9926ce55d14fb1430238100f8e94fc21ccbd43c03f78b507294c2

Observation db16b6e6-ad9b-41c2-90e9-ba8546c09857 · outbound

This paper cites Deep residual learning for image recognition.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Deep residual learning for image recognition

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:10.914790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:22:10.914790Z digest=sha256:f23cec469900beda45128cb69285dbcde0240273e6a4b51ba587dbad93204f5f

Observation e2cee5e1-95b3-46ec-92d9-a937226617e8 · outbound

This paper cites RefMask3D: Language-Guided Transformer for 3D Referring Segmentation.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding RefMask3D: Language-Guided Transformer for 3D Referring Segmentation

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-15T23:22:11.176082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.918546Z digest=sha256:0182332e8c4acc1338b4fd30e12813c9a6a22d71f999d707a88dc6c27cd2df90

Observation e050536f-246f-46c6-8aaf-da812f6dca6b · outbound

This paper cites Chat-scene: Bridging 3d scene and large language models with object identifiers.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Chat-scene: Bridging 3d scene and large language models with object identifiers

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.638244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.922686Z digest=sha256:b277c8ff4ef8363c88fa1a882250bb288dd1052dfd8916dd4c75f82d1c0878a4

Observation bb067992-5177-40eb-a828-3ce133f95f14 · outbound

This paper cites Segment3d: Learning fine-grained class-agnostic 3d segmentation without manual labels.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Segment3d: Learning fine-grained class-agnostic 3d segmentation without manual labels

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.623845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.926576Z digest=sha256:c86a521bf27b8572b31d8b0ee4274e05d5e2d2204dba391fccb57c0bf7ab2996

Observation b35acb44-5fdf-4027-b599-4540ab3fa56c · outbound

This paper cites Training an open-vocabulary monocular 3d detection model without 3d data.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Training an open-vocabulary monocular 3d detection model without 3d data

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.608915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.930381Z digest=sha256:8543baa7409afe7000c345246a15c8d4c2fccacc99fc2fdc36283d459f1be97c

Observation 5005e507-dbae-48e3-a179-809fa931302f · outbound

This paper cites Multi-view transformer for 3d visual grounding.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Multi-view transformer for 3d visual grounding

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.591559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.933896Z digest=sha256:f617c12cd274465b06f9a640c37d7828a51c940af98fee8a49684cccd9f83602

Observation 1ebf1b95-300e-4cd2-97f4-bfb3d49495cf · outbound

This paper cites Bottom up top down detection transformers for language grounding in images and point clouds.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Bottom up top down detection transformers for language grounding in images and point clouds

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.576073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.937361Z digest=sha256:24297337e7c57a02d5bb05419c6f769ae14162e1d0094f44de3b56f170052116

Observation 12fc7542-c801-4fcf-92ac-9643ac6d2871 · outbound

This paper cites Pointgroup: Dual-set point grouping for 3d instance segmentation.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Pointgroup: Dual-set point grouping for 3d instance segmentation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.560890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.940888Z digest=sha256:3169c772eb9dffe0af41136e538669612f56c77807612acf06e815f61e5582be

Observation 9a169e3d-fac9-4203-b343-2fbf2eb26ca1 · outbound

This paper cites Q: How to specialize large vision-language models to data-scarce vqa tasks? a: Self-train on unlabeled images! In CVPR, 2023.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Q: How to specialize large vision-language models to data-scarce vqa tasks? a: Self-train on unlabeled images! In CVPR, 2023

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.545880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.945195Z digest=sha256:eb8ef7d8e30d9e9df8817de6fa22baffeda55c869dee8f551adf9f1d1cd41590

Observation 493215af-6987-4018-acf5-84f9c21d2b92 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding OpenVLA: An Open-Source Vision-Language-Action Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:10.949601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:22:10.949601Z digest=sha256:da5e8b1b9d218a584aa996521f61d1597032276ca47dc983d4a03bf6f5146208

Observation 140639e6-0e5b-44ae-8596-55a56c2f7fb9 · outbound

This paper cites A real-time cross-modality correlation filtering method for referring expression comprehension.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding A real-time cross-modality correlation filtering method for referring expression comprehension

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.530724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.953540Z digest=sha256:0055e992cd75eb69edb60a59c1795b64cc6dcd7512dd26af90e5a0e006cc5db1

Observation 14d68064-2efb-44f4-94b7-05035f206b30 · outbound

This paper cites Roberta: A robustly optimized bert pretraining approach, 2019.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Roberta: A robustly optimized bert pretraining approach, 2019

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:10.957752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:22:10.957752Z digest=sha256:745a39d38f044c4b053a38c591c495c3084cf083d36509c8e8fe082fffcd2b12

Observation 79f19f0f-681b-4c9f-b66a-ed1ce2c6f948 · outbound

This paper cites Bevfusion: Multi-task multi-sensor fusion with unified bird's-eye view representation.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Bevfusion: Multi-task multi-sensor fusion with unified bird's-eye view representation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.506855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.962244Z digest=sha256:5c2f2b6d67b212f8ff84ad9cbca636133a6bf52378531947918603965ecf8816

Observation 519e48c6-5367-457b-9675-8cffdea805ff · outbound

This paper cites 3d-sps: Single-stage 3d visual grounding via referred point progressive selection.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding 3d-sps: Single-stage 3d visual grounding via referred point progressive selection

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.490036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.966349Z digest=sha256:77c740c715a6186fbfaf26dae38d3661f4e96cdf22f129eaa719a59cd6be8195

Observation 5fb46b7a-2e0f-40aa-9d53-f5f1619ac1b5 · outbound

This paper cites Multi-modal understanding and generation for medical images and text via vision-language pre-training.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Multi-modal understanding and generation for medical images and text via vision-language pre-training

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.475629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.970545Z digest=sha256:d78aad43a31bd28bcb17a8390b6e710817f6ecfa301d138c4f2747eb3fd0a126

Observation ea55a265-c636-423b-93df-ac8ae50a9f11 · outbound

This paper cites Instruction Tuning with GPT-4.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Instruction Tuning with GPT-4

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:10.975479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:22:10.975479Z digest=sha256:0c594a0b91f89c8a9945eb763614f9de26c9e3f98cb7f05d8520ce7854b81734

Observation 9b05a041-b4eb-4435-9b3a-a009a198aa98 · outbound

This paper cites Zeetad: Adapting pretrained vision-language model for zero-shot end-to-end temporal action detection.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Zeetad: Adapting pretrained vision-language model for zero-shot end-to-end temporal action detection

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.461351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.980052Z digest=sha256:916b0cfde0d16f1f0a455be545f819923760aac8037a9f109a5c86b93a1547d6

Observation 359d9bb0-c333-48a3-98c2-71588051899e · outbound

This paper cites Learning transferable visual models from natural language supervision.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Learning transferable visual models from natural language supervision

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:10.984132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:22:10.984132Z digest=sha256:8c4325f7ab5be8f614d6fc314b5bdf711003f8c19ba053e0e52a50adbcda1903

Observation b80dccc8-8fe5-45a6-b1b3-5771bb4ed31c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding LLaMA: Open and Efficient Foundation Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:10.988128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:22:10.988128Z digest=sha256:cf8356222ba417297ffac79cbb3f29b6d7ed84703a95f183e20b8f89150bba56

Observation 76b862f8-8088-4bd4-ac28-8796ff9376f1 · outbound

This paper cites Rio: 3d object instance re-localization in changing indoor environments.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Rio: 3d object instance re-localization in changing indoor environments

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.436912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.992150Z digest=sha256:116f75fd7701e7c3ddb9fe1bc92d481ce01fdb464e3e771088a071d83587fff2

Observation c881d1b7-44c6-4b1f-9560-5c9ac4a1b486 · outbound

This paper cites Embodiedscan: A holistic multi-modal 3d perception suite towards embodied ai.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Embodiedscan: A holistic multi-modal 3d perception suite towards embodied ai

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.422385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:10.996348Z digest=sha256:daf878deba4c25c8b88a0c7a0c3c95f1be01315f266642f9aedd2d66621b05f4

Observation 6cf46af8-9b8a-4847-b8a2-53df6057c9fa · outbound

This paper cites Data-Efficient 3D Visual Grounding via Order-Aware Referring.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Data-Efficient 3D Visual Grounding via Order-Aware Referring

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:11.000154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:22:11.000154Z digest=sha256:b656d466d815683d029665c9908103cbb68a5cb30e76d52f3ef0631b6f69fb3d

Observation b88d02ed-6f69-42c6-a945-98e4965c4d6d · outbound

This paper cites Point transformer v3: Simpler faster stronger.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Point transformer v3: Simpler faster stronger

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.407653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:11.004228Z digest=sha256:089a73143fc33ab460a4572f2cb68c5628fe80e64d68fac0112d1bfee8bd4bc2

Observation 8643aaf3-88f0-43fc-ad87-9995c58a503a · outbound

This paper cites Eda: Explicit text-decoupling and dense alignment for 3d visual grounding.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Eda: Explicit text-decoupling and dense alignment for 3d visual grounding

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.392445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:11.008363Z digest=sha256:a35120e6b5bdf08ceb3bc1bd06965363bf25e1d04f6fc3283dc9f8f47f447c40

Observation 9370c879-6ad1-45cf-ba9a-ee061b1f89cf · outbound

This paper cites Exploiting contextual objects and relations for 3d visual grounding.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Exploiting contextual objects and relations for 3d visual grounding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.378176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:11.012443Z digest=sha256:fa7db28806bd29dcf71929fbe59aeaa8047a04f896a442d6288582a118ec5cee

Observation eb052e49-edf9-436b-993a-4b753dfe60a7 · outbound

This paper cites Dynamic graph attention for referring expression comprehension.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Dynamic graph attention for referring expression comprehension

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.362859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:11.016043Z digest=sha256:ff80410bf73cf2189ba82c8b68a10601630b2880cb6ebe5b000ac0783414f6e7

Observation a7c0c1d3-1f92-4284-980a-d783f7d08ca1 · outbound

This paper cites 3dvg-transformer: Relation modeling for visual grounding on point clouds.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding 3dvg-transformer: Relation modeling for visual grounding on point clouds

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.348560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:11.019458Z digest=sha256:39769e392d7958a42688547539186d355071f48d9d4bd0fe9454df0a01507df9

Observation e7086f8d-467e-4228-b22c-bc9786ff61fd · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:11.023031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:22:11.023031Z digest=sha256:4553fa4471e6bd22d4f94e26dc58aacad2c14da8d65c41582e5c16424a33bbd3

Observation ab5a6027-62bc-40b4-9d1f-d0c77e5880b4 · outbound

This paper cites Denseg: Alleviating vision-language feature sparsity in multi-view 3d visual grounding.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Denseg: Alleviating vision-language feature sparsity in multi-view 3d visual grounding

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.334022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:11.027068Z digest=sha256:979b17341a88d94a515b233c39058961f8ed473abfc63804ad2cff448bf9a392

Observation 65ed72a9-fe48-4934-8431-2b39b3c849c1 · outbound

This paper cites Class-balanced Grouping and Sampling for Point Cloud 3D Object Detection.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Class-balanced Grouping and Sampling for Point Cloud 3D Object Detection

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:11.030711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:22:11.030711Z digest=sha256:fac0bda2e069da88d35c3eaf46d03f0425ca381195a3fac681ac735273c5a419

Observation bfb33ee9-56bd-4b8c-9476-271ce887608f · outbound

This paper cites Object2scene: Putting objects in context for open-vocabulary 3d detection, 2023 a.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Object2scene: Putting objects in context for open-vocabulary 3d detection, 2023 a

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.319863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:11.035146Z digest=sha256:f4fbd17e76ba6ca230c6cc88b9e2d89228a77700b8b1ef569fa3bc449c2a2e1e

Observation ddb163f1-97f7-4aa6-9e2e-c6421e87e3c5 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding 3d-vista: Pre-trained transformer for 3d vision and text alignment

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.305806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:11.039811Z digest=sha256:b5110fde4dafc9e91f3a51173fdc1939750c245e35f2137caff3daa8a4135275

Observation 3763223f-c46c-40d8-b731-f36d08a6ff4a · outbound

This paper cites @esa (Ref.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding @esa (Ref

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:11.044007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:22:11.044007Z digest=sha256:fa10f915a1bbde7b223aed5954d800cc71b839a38ec84f1c5c845744f6ae84bb

Observation d3d174c6-d0d1-4e03-b211-419fb0e4df05 · outbound

This paper cites an unresolved cited work.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:11.048868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:22:11.048868Z digest=sha256:117b10423552e6da879eb5a2fb5b9b45c8c3e8991f71d1c7c5be92f0720bc8fa

Observation 3da01b25-8898-41de-a7e0-c62c972cc1ce · outbound

This paper cites e.g E.g i.e I.e cf Cf etc vs wrt d.o.f et al i.i.d * 90 [1] 0.5em #1.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding e.g E.g i.e I.e cf Cf etc vs wrt d.o.f et al i.i.d * 90 [1] 0.5em #1

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:11.276105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:22:11.053165Z digest=sha256:b679cc07055c915d34656b924ef5cfe0ce4a718807235856140ee93fa0992ef5

Pith citing papers

Observation 81f17d1f-e9b9-44ab-bf37-731031f23fee · inbound

OpenGround: Planning-based Online Perception for Open-World 3D Visual Grounding cites this paper.

OpenGround: Planning-based Online Perception for Open-World 3D Visual Grounding DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-03T13:46:30.353680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:46:30.353680Z digest=sha256:ed815c7bcace9939e789d1606c1269b1943b3bed1532a4ed3699243e9d210a36

Observation 53605b2c-7b62-4b1e-ba4c-52323e074f0c · inbound

PriorVLA: Prior-Preserving Adaptation for Vision-Language-Action Models cites this paper.

PriorVLA: Prior-Preserving Adaptation for Vision-Language-Action Models DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:26.078635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T03:37:53.732205Z digest=sha256:14b08b80e808519673a3fbca74bc2d9c8c72dd65216863e954675003df055693

Observation de151fe1-60c0-4144-90b5-23aad1274deb · inbound

ThinkAfford: Affordance-Centric Reasoning for Fine-Grained 3D Grounding in Cluttered Scenes cites this paper.

ThinkAfford: Affordance-Centric Reasoning for Fine-Grained 3D Grounding in Cluttered Scenes DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T13:00:00.925521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:00:00.925521Z digest=sha256:f102d770c408edc8151ee1bd146978217183a91a13fab069ed2094e231edf761