Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:02:33.215747Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 2 inbound Pith citation observations for arXiv:2507.04047.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:02:33.215747Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T19:54:57.184897Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-13T19:03:08.599602Z
96 of 96 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e2a31111-e2e4-4890-a72a-2b5fe7745999 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scanents3d: Exploit- ing phrase-to-3d-object correspondences for improved visio- linguistic models in 3d scenes
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f2de7b2-1a14-431d-a5c8-30bbdd9a6843 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4f365fa-c59c-42b1-9e52-546e5375435c · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scanqa: 3d question answering for spatial scene understanding
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c49bc4df-6ccb-4a9e-9bbe-fd1b50ea2390 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Do as i can, not as i say: Grounding language in robotic affordances
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb7ccfdb-041b-411f-993c-23f9995f6668 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9dddfbd-8cbb-48ea-b292-57527f9aadad · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Emerg- ing properties in self-supervised vision transformers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation deda0b58-fc06-49a7-b970-3e239931f8ae · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Object goal naviga- tion using goal-oriented semantic exploration
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b823c8a9-85dc-4284-8432-c24334a7769b · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Object goal navi- gation using goal-oriented semantic exploration
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86802e41-5707-4c97-87d5-a88d5deb6b24 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scanrefer: 3d object localization in rgb-d scans using natu- ral language
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06333299-344d-4207-b617-3a79b9607d6c · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Language conditioned spatial relation reasoning for 3d object grounding
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 778803a5-506f-493f-99d9-949b414b8d32 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Ll3da: Visual interactive instruction tuning for omni-3d understand- ing reasoning and planning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4822169-e117-457f-8f5d-242e6d091ec3 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scan2cap: Context-aware dense captioning in rgb- d scans
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26c200ce-8196-49a5-a404-ede974461053 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Unit3d: A unified trans- former for 3d dense captioning and visual grounding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96a42d0e-704e-4d5e-b86c-f82703300d84 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Schwing, Alexan- der Kirillov, and Rohit Girdhar
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed1e8638-0b8a-4a26-9984-d39ecfa84001 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 4d spatio-temporal convnets: Minkowski convolutional neural networks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17ccbf1e-189e-4f93-b408-cafceb94bcaf · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scannet: Richly-annotated 3d reconstructions of indoor scenes
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23dd4bc8-824e-48b8-99bd-72233555fe15 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Embodied question answer- ing
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fec0ec8e-b139-451f-b6b8-91dcf8ba2993 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation A survey of embodied ai: From simulators to research tasks
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6db28803-3cf2-4251-ad77-654dcd228b67 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation The One RING: a Robotic Indoor Navigation Generalist
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a984db5-a1cb-4e85-842d-0a3a593d6b52 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Spoc: Imitating short- est paths in simulation enables effective navigation and ma- nipulation in the real world
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 713006fa-6b05-414d-a629-451cd2717320 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07a13e23-4b26-4210-91eb-1174c07e9df7 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Efficient graph-based image segmentation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1a3fd2f0-a809-4758-ad4d-a2be5e0be2cc · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Cows on pasture: Base- lines and benchmarks for language-driven zero-shot object navigation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8c68ce8f-f0a4-4e70-9234-da915752a457 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scaling open-vocabulary image segmentation with image- level labels
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 45f432e4-e104-4efe-9322-2a706db8a7fa · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Conceptgraphs: Open-vocabulary 3d scene graphs for per- ception and planning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation adf93ba6-943c-4359-bfef-d23cc8551d51 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Viewrefer: Grasp the multi-view knowledge for 3d visual grounding
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7d587f8b-2923-40f4-8826-daaabe53f5d4 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Transrefer3d: Entity-and- relation aware transformer for fine-grained 3d visual ground- ing
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0a0a1021-5c64-4d71-8425-ea1db88b305a · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Vln bert: A recurrent vision- and-language bert for navigation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d4c652b4-9733-408a-8d2e-c5eb0468f7ee · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 3d-llm: In- jecting the 3d world into large language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7f29a1b4-1c1e-4eb9-955e-9b9381b77dcd · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation A real-time occupancy map from multiple video streams
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9f43c2d6-7681-47f4-b4a6-54b61de930e7 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation An embodied generalist agent in 3d world
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4c5594b4-3ee3-422a-b71b-7dba27b78f5f · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Language models as zero-shot planners: Ex- tracting actionable knowledge for embodied agents
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 143deb32-0c82-488b-81f8-15730725a8d9 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation V oxposer: Composable 3d value maps for robotic manipulation with language models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 39bbe7cb-7eaf-46c9-a293-620806b5320b · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Bottom up top down detection transform- ers for language grounding in images and point clouds
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8f216072-19fc-4d8f-9511-ee0f109fc652 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Conceptfusion: Open-set multi- modal 3d mapping
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e52ef426-8946-4d81-856e-63b42e949343 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Sceneverse: Scaling 3d vision-language learning for grounded scene understanding
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b46b61ae-463b-42ff-a9ca-487841e009f3 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Goat-bench: A benchmark for multi-modal lifelong navigation
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9eeaf62f-5c89-4be2-be2a-02d3f7d04fb8 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Realfred: An em- bodied instruction following benchmark in photo-realistic environments
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 30f6df15-b03d-4f0c-b060-014a2e9d8e01 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Segment anything
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b1919572-cd83-449d-826b-e48b8505acfe · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation UniCLIP: Unified Framework for Contrastive Language-Image Pre-training
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8207f357-d3b8-407e-a437-96f5b9a07aa8 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Less is more: Clipbert for video-and-language learning via sparse sampling
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ad92fa97-2683-4468-9993-289fb62236ff · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d582015-4573-4d72-90f1-53a9abbd8ed0 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a1c475a-e27f-4406-9282-799d56d74131 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Panoptic segformer: Delving deeper into panoptic segmen- tation with transformers
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 53c90110-e900-4bf6-b726-6d8daf6d9d1e · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Code as policies: Language model programs for embodied con- trol
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b986deef-cc20-4681-a7cf-88509af4eabd · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d1902b8-80a2-4ddf-90eb-aa605a81c199 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c545085f-88f9-4e78-bb8c-50e4337486a6 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Sqa3d: Situated question answering in 3d scenes
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 41653663-ad6f-4d35-8548-2f546cc1f18e · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Zson: Zero-shot object-goal navigation using multimodal goal embeddings
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e621b6db-fe2c-4a9d-8df8-dbfe3a847b75 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Openeqa: Embodied question answering in the era of foun- dation models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9b824185-5025-4395-aca8-665e709e9b42 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Openeqa: Embodied question answering in the era of foun- dation models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bb3fbf5b-e554-44e3-9818-b6a3b50163c3 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Spatial memory
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6da5b11a-36a1-404a-b98c-a3b75be6c88e · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation DINOv2: Learning Robust Visual Features without Supervision
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 729ec64a-1e53-4c9e-af84-67004af85111 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Teach: Task-driven embodied agents that chat
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 91e13b98-5db9-4a06-b856-b57512660a26 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Openscene: 3d scene understanding with open vocabularies
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0bb53d3a-b4fb-4ce6-a6bd-76e809f22e22 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Learn- ing transferable visual models from natural language super- vision
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8337e663-f1b0-4a05-9243-ec3752d2a8ee · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Pirlnav: Pretraining with imitation and rl finetuning for objectnav
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 738a7c9a-cb3b-4a08-b6fb-3f9a641194f3 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Sayplan: Ground- ing large language models using 3d scene graphs for scalable task planning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 841e3bdb-4be8-4196-9e45-3ab0b52d98ca · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Explore until Confident: Efficient Exploration for Embodied Question Answering
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bd1b137-7252-436e-90e0-50c9dee394da · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Language- grounded indoor 3d semantic segmentation in the wild
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation acc76002-445d-4d6a-a6a7-d62271306133 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Habitat: A plat- form for embodied ai research
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7bc365d5-0260-4142-b6f4-cb24fdbc7afd · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Proximal Policy Optimization Algorithms
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46baae52-902b-4d3a-bb2d-173f6c84d29a · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Mask3d: Mask trans- former for 3d semantic instance segmentation
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0a686420-7872-4d1c-a62c-54e40f39f5ea · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Alfred: A benchmark for interpreting grounded instructions for everyday tasks
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 170add76-540e-4418-8d40-8792d0b32b8c · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Llm-planner: Few-shot grounded planning for embodied agents with large language models
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9bf74e8c-1c66-4fcf-9d3e-71c08d19150b · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Habitat 2.0: Training home assistants to rearrange their habitat
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9530d43-f3e2-4889-b87e-9045a8cd09d1 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Sumner, Marc Pollefeys, Federico Tombari, and Francis Engelmann
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c9361746-d768-4873-a170-e937b6194819 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Rio: 3d object instance re- localization in changing indoor environments
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c3958155-876a-4424-a8c8-c14914faa310 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Embodiedscan: A holistic multi- modal 3d perception suite towards embodied ai
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9840311b-3a6c-44c1-b628-fe68e2d051ac · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion Frames
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d855c8e7-24ce-4c92-bf00-1cde8f9bce5c · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Ver: Scaling on- policy rl leads to the emergence of navigation in embodied rearrangement
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c4dbf192-0e51-4d22-9090-66021d4bdc9b · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scenegraphfusion: Incre- mental 3d scene graph prediction from rgb-d sequences
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 73cebb80-6c31-4bd7-be82-84128aef8574 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation EmbodiedSAM: Online Segment Any 3D Thing in Real Time
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f204022-07ac-431f-b824-1607b07efda7 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Habitat-matterport 3d semantics dataset
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 66540fd2-5ebc-42da-8a41-49a9ceb41a84 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Frontier-based exploration using multiple robots
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1525a073-ef2b-44c6-bb44-cf0d9203795b · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b28ee52-6bb1-45e5-a211-0a0961a5e8fc · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 3d-mem: 3d scene memory for embodied exploration and reasoning
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f7902776-3426-4812-ba1d-835c3e02c367 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Vlfm: Vision-language frontier maps for zero-shot semantic navigation
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e846a639-0e12-4c7f-96d5-1bab80731938 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Hm3d-ovon: A dataset and bench- mark for open-vocabulary object goal navigation
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0e5784ad-8c37-4039-8d60-94a6b91f6d0e · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Frontier semantic exploration for visual target navigation
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aae9a2f4-215d-4950-8a5f-1f54cf5c022c · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Instancere- fer: Cooperative holistic understanding for visual ground- ing on point clouds through instance multi-level contextual referring
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2609d05e-91ce-4d80-b285-5d697180c3b5 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation PoliFormer: Scaling On-Policy RL with Transformers Results in Masterful Navigators
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 583e3afd-8e7f-47ef-9722-8da0e9f8a503 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6063f092-66ec-41ba-84c4-cb0f6f1045a0 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Vision-Language Pre-training with Object Contrastive Learning for 3D Scene Understanding
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5f9780c8-10d5-4986-8f9a-4d97c8147ca9 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Multi3drefer: Grounding text description to multiple 3d objects
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f7906a49-1d19-4666-ad00-b2842a1db9ca · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Microsoft kinect sensor and its effect
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f7fba5c5-b1d6-49a2-a672-f6d67140bc4e · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Task-oriented Sequential Grounding and Navigation in 3D Scenes
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0f133cb-4108-4947-b1fc-836c269d090e · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation To- wards explainable 3d grounded visual question answering: A new benchmark and strong baseline
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a10528fa-1096-4e34-adf2-8d1bfbaa1104 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Fast Segment Anything
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0dbc718-15b3-4fef-ba94-993a9eaf9604 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 3D-VLA: A 3D Vision-Language-Action Generative World Model
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83388a1b-427e-42dd-a5f5-0aa425a01d7c · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Dual memory units with uncertainty regulation for weakly supervised video anomaly detection
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 96ec4986-1805-49e8-ad9c-1def13a40b0b · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scanreason: Empowering 3d visual grounding with reasoning capabilities
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d4465aff-60cc-4572-b653-dfdb711f5ee1 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 3d-vista: Pre-trained transformer for 3d vision and text alignment
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 225f2ccb-d3ce-4433-9dbc-f8934a0e1169 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Unifying 3d vision-language understanding via prompt- able queries
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 17a85cb1-9423-4a72-994e-4fc5374ca733 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation TANGO: Training-free Embodied AI Agents for Open-world Tasks
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be143c56-3127-4de0-84c2-0274233b3c85 · outbound
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Generalized decoding for pixel, image, and language
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8b12b765-c17c-4532-8932-0c2f5ffe3a7d · inbound
SPG: Style-Prompting Guidance for Style-Specific Content Creation Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7363683b-1a7a-41e1-8f77-a918c1574094 · inbound
FSUNav: A Cerebrum-Cerebellum Architecture for Fast, Safe, and Universal Zero-Shot Goal-Oriented Navigation Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.