Pith. sign in

Paper Citation Record · LEDGER

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation

As of 7 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 2 inbound Pith citation observations for arXiv:2507.04047.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04047 v2

Coverage vector

measured 96 of 96 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:02:33.215747Z

measured 98 of 98 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T19:54:57.184897Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T19:03:08.599602Z

Reference resolution

96 of 96 outbound references displayed

  • verified exact2
  • verified fuzzy59
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e2a31111-e2e4-4890-a72a-2b5fe7745999 · outbound

This paper cites Scanents3d: Exploit- ing phrase-to-3d-object correspondences for improved visio- linguistic models in 3d scenes.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scanents3d: Exploit- ing phrase-to-3d-object correspondences for improved visio- linguistic models in 3d scenes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.156912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.156912Z digest=sha256:9c18185923ae7889440b81cc1505e8c37c7c6de581509942e8cc52f463cd1f5c

Observation 7f2de7b2-1a14-431d-a5c8-30bbdd9a6843 · outbound

This paper cites Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.251850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.251850Z digest=sha256:a3d08056755a1470b3071a8ac102cd9513c6a2eb95dacc07d3e4b181b113ee3d

Observation d4f365fa-c59c-42b1-9e52-546e5375435c · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scanqa: 3d question answering for spatial scene understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.317126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.317126Z digest=sha256:648d80628792e9c43b196f20d1f85e295e6faab00aaf129e0b010c4ac5d493f2

Observation c49bc4df-6ccb-4a9e-9bbe-fd1b50ea2390 · outbound

This paper cites Do as i can, not as i say: Grounding language in robotic affordances.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Do as i can, not as i say: Grounding language in robotic affordances

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.434833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.434833Z digest=sha256:b2f8077a5838c84aa7ded81523997b1b19d22c2f2203bd1f9a8131362dc8b286

Observation bb7ccfdb-041b-411f-993c-23f9995f6668 · outbound

This paper cites 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.497301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.497301Z digest=sha256:da014bfae0e5d5562e100f147600f5f367cfc5adbb06c0d8ace40e53a107e008

Observation f9dddfbd-8cbb-48ea-b292-57527f9aadad · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Emerg- ing properties in self-supervised vision transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.605674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.605674Z digest=sha256:b3691d940dc42a790df54d5c400ec7dc469954ee5e63e7ee7c66bcadfa9b9ded

Observation deda0b58-fc06-49a7-b970-3e239931f8ae · outbound

This paper cites Object goal naviga- tion using goal-oriented semantic exploration.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Object goal naviga- tion using goal-oriented semantic exploration

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.679473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.679473Z digest=sha256:06542bef39e8b68d4562acc094ddcf458a8aab49899f897fd39f3382c83c627f

Observation b823c8a9-85dc-4284-8432-c24334a7769b · outbound

This paper cites Object goal navi- gation using goal-oriented semantic exploration.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Object goal navi- gation using goal-oriented semantic exploration

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.763679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.763679Z digest=sha256:165848d98121afacf6ccaacaecc201a641d0d945b6d9afc989f88c35340f6764

Observation 86802e41-5707-4c97-87d5-a88d5deb6b24 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natu- ral language.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scanrefer: 3d object localization in rgb-d scans using natu- ral language

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.842833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.842833Z digest=sha256:a9179132600d885240ae4f44bffdeca9b519c1f4acb2dcbc3d8600c506284bd9

Observation 06333299-344d-4207-b617-3a79b9607d6c · outbound

This paper cites Language conditioned spatial relation reasoning for 3d object grounding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Language conditioned spatial relation reasoning for 3d object grounding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.924326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.924326Z digest=sha256:d4f104028ff0f6c66b5716e53dee36b7b2646ac9d4754c711769acc4c0963ac8

Observation 778803a5-506f-493f-99d9-949b414b8d32 · outbound

This paper cites Ll3da: Visual interactive instruction tuning for omni-3d understand- ing reasoning and planning.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Ll3da: Visual interactive instruction tuning for omni-3d understand- ing reasoning and planning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.014240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.014240Z digest=sha256:d234fec279def8838e2be266bc1b17da0ab05a28232ac1ef10118f008f6a45e5

Observation e4822169-e117-457f-8f5d-242e6d091ec3 · outbound

This paper cites Scan2cap: Context-aware dense captioning in rgb- d scans.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scan2cap: Context-aware dense captioning in rgb- d scans

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.057605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.057605Z digest=sha256:f7ed483922bed8474f81b4ac527cd3816dff63cbdfd36e6fd87a94c190d0da85

Observation 26c200ce-8196-49a5-a404-ede974461053 · outbound

This paper cites Unit3d: A unified trans- former for 3d dense captioning and visual grounding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Unit3d: A unified trans- former for 3d dense captioning and visual grounding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.194145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.194145Z digest=sha256:b56fd5d8cfa3bba17434ce39b5f521345fe674f73bca423d452083470068413e

Observation 96a42d0e-704e-4d5e-b86c-f82703300d84 · outbound

This paper cites Schwing, Alexan- der Kirillov, and Rohit Girdhar.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Schwing, Alexan- der Kirillov, and Rohit Girdhar

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.284849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.284849Z digest=sha256:5f913690298e290af7ad619dfd14c090232a1b29b46aec28e17753632000294c

Observation ed1e8638-0b8a-4a26-9984-d39ecfa84001 · outbound

This paper cites 4d spatio-temporal convnets: Minkowski convolutional neural networks.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 4d spatio-temporal convnets: Minkowski convolutional neural networks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.391830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.391830Z digest=sha256:9a3b1221a1c32425b7fefa1d936ff5c7bf14a8702f44932d96cdbf81353bc594

Observation 17ccbf1e-189e-4f93-b408-cafceb94bcaf · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.507880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.507880Z digest=sha256:c23da970a9ae0b9886aa5faf99538f7cf47d477fec00d4386f2f7e66c69bd2e8

Observation 23dd4bc8-824e-48b8-99bd-72233555fe15 · outbound

This paper cites Embodied question answer- ing.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Embodied question answer- ing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.662553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.605207Z digest=sha256:d65f76ed85b14236011da5c172b6983d6dd6538ec2c6c45d95d6454b88f38edf

Observation fec0ec8e-b139-451f-b6b8-91dcf8ba2993 · outbound

This paper cites A survey of embodied ai: From simulators to research tasks.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation A survey of embodied ai: From simulators to research tasks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.648665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.730308Z digest=sha256:548a7ce4f90d4e354b44fb252d85f73430332d66f846f79a4bf22b0b8d9f64df

Observation 6db28803-3cf2-4251-ad77-654dcd228b67 · outbound

This paper cites The One RING: a Robotic Indoor Navigation Generalist.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation The One RING: a Robotic Indoor Navigation Generalist

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.828420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.828420Z digest=sha256:206b4b0ad860616f9c10636685d2f010c562067c9e7605728a8395e02d36828a

Observation 4a984db5-a1cb-4e85-842d-0a3a593d6b52 · outbound

This paper cites Spoc: Imitating short- est paths in simulation enables effective navigation and ma- nipulation in the real world.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Spoc: Imitating short- est paths in simulation enables effective navigation and ma- nipulation in the real world

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.633572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.859253Z digest=sha256:488bfd7d378871d5a5eda623834601b03ec16dd033e13110913fda7ca74c73d5

Observation 713006fa-6b05-414d-a629-451cd2717320 · outbound

This paper cites Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.864163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.864163Z digest=sha256:bf02390e25caa6e7973438b99b934462240ba170de3914eeb3307f6826285556

Observation 07a13e23-4b26-4210-91eb-1174c07e9df7 · outbound

This paper cites Efficient graph-based image segmentation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Efficient graph-based image segmentation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.618261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.869619Z digest=sha256:d45afaaa90d753d6911d68b7ee67e59377b98db2e54d496827fd0f90aa8bd6ab

Observation 1a3fd2f0-a809-4758-ad4d-a2be5e0be2cc · outbound

This paper cites Cows on pasture: Base- lines and benchmarks for language-driven zero-shot object navigation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Cows on pasture: Base- lines and benchmarks for language-driven zero-shot object navigation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.604356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.873766Z digest=sha256:75539aeb9eb10fb696f2174347f175a7a9bd3870e7839fb5c700e043620b6ade

Observation 8c68ce8f-f0a4-4e70-9234-da915752a457 · outbound

This paper cites Scaling open-vocabulary image segmentation with image- level labels.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scaling open-vocabulary image segmentation with image- level labels

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.589762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.878891Z digest=sha256:9531dd0466229091ad46f0f657d4daa713474d1af9e3c0d96fcc4dca4c268f22

Observation 45f432e4-e104-4efe-9322-2a706db8a7fa · outbound

This paper cites Conceptgraphs: Open-vocabulary 3d scene graphs for per- ception and planning.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Conceptgraphs: Open-vocabulary 3d scene graphs for per- ception and planning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.575555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.883132Z digest=sha256:c0d6e251a2beb59d18dc861aaf39299690d0960821fbd424a63e921fc1bdb8e8

Observation adf93ba6-943c-4359-bfef-d23cc8551d51 · outbound

This paper cites Viewrefer: Grasp the multi-view knowledge for 3d visual grounding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Viewrefer: Grasp the multi-view knowledge for 3d visual grounding

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.561287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.887401Z digest=sha256:a5bbad8f7aa61d21d3b06730074d32756917555239a2662fe66169f93ee218f5

Observation 7d587f8b-2923-40f4-8826-daaabe53f5d4 · outbound

This paper cites Transrefer3d: Entity-and- relation aware transformer for fine-grained 3d visual ground- ing.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Transrefer3d: Entity-and- relation aware transformer for fine-grained 3d visual ground- ing

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.546556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.891844Z digest=sha256:af5552908cf3b42719fc39cdd35bf36a04fa92a36358a12b09f394891c0de580

Observation 0a0a1021-5c64-4d71-8425-ea1db88b305a · outbound

This paper cites Vln bert: A recurrent vision- and-language bert for navigation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Vln bert: A recurrent vision- and-language bert for navigation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.531842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.896182Z digest=sha256:c93c2bf8ff0fb619331358f78319b34bae50b7b7ac50e36b635782fa254655e2

Observation d4c652b4-9733-408a-8d2e-c5eb0468f7ee · outbound

This paper cites 3d-llm: In- jecting the 3d world into large language models.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 3d-llm: In- jecting the 3d world into large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.515695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.901425Z digest=sha256:b2f78fad1431bcd9e4142b5a673bdcde7c6c2fca3b3970032f5bcbf2b4a771be

Observation 7f29a1b4-1c1e-4eb9-955e-9b9381b77dcd · outbound

This paper cites A real-time occupancy map from multiple video streams.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation A real-time occupancy map from multiple video streams

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.498323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.906117Z digest=sha256:200bf8dbbdea15d7c53702347034f40d7664c4e617f5a1c2d32f4a4f4c572597

Observation 9f43c2d6-7681-47f4-b4a6-54b61de930e7 · outbound

This paper cites An embodied generalist agent in 3d world.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation An embodied generalist agent in 3d world

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.481002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.910518Z digest=sha256:1cbe9cdd6dcbfc96959c0e3e77e79d5ed7a0c6054653d69371617394ee04abe7

Observation 4c5594b4-3ee3-422a-b71b-7dba27b78f5f · outbound

This paper cites Language models as zero-shot planners: Ex- tracting actionable knowledge for embodied agents.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Language models as zero-shot planners: Ex- tracting actionable knowledge for embodied agents

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.463894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.914903Z digest=sha256:c9d0050f411556d90e4fcb6b5f3d785538d81256a5ef563c75c2a7b375ff511e

Observation 143deb32-0c82-488b-81f8-15730725a8d9 · outbound

This paper cites V oxposer: Composable 3d value maps for robotic manipulation with language models.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation V oxposer: Composable 3d value maps for robotic manipulation with language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.448269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.919522Z digest=sha256:1a599cfba127cc0d7a43c69f4fac1e5ad668f3490e8b8ed9d6ae0f4f44ad19cf

Observation 39bbe7cb-7eaf-46c9-a293-620806b5320b · outbound

This paper cites Bottom up top down detection transform- ers for language grounding in images and point clouds.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Bottom up top down detection transform- ers for language grounding in images and point clouds

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.433025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.924413Z digest=sha256:128528d23b9ba43bd4fadfa517a8c5d9d7452c3a92e206476cc21d62f835300c

Observation 8f216072-19fc-4d8f-9511-ee0f109fc652 · outbound

This paper cites Conceptfusion: Open-set multi- modal 3d mapping.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Conceptfusion: Open-set multi- modal 3d mapping

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.417777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.929001Z digest=sha256:e501600541d78a9f5ed74f7a8c9df3e49b39ea8a510977a0905e3897e9d296c9

Observation e52ef426-8946-4d81-856e-63b42e949343 · outbound

This paper cites Sceneverse: Scaling 3d vision-language learning for grounded scene understanding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Sceneverse: Scaling 3d vision-language learning for grounded scene understanding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.402368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.934381Z digest=sha256:6225f56eaf64680324cebac11528bd5946c3d5c403380ff55ec1d052fe49d83f

Observation b46b61ae-463b-42ff-a9ca-487841e009f3 · outbound

This paper cites Goat-bench: A benchmark for multi-modal lifelong navigation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Goat-bench: A benchmark for multi-modal lifelong navigation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.386201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.939241Z digest=sha256:93eb225398ac7ca0263d8c9710bdd88c3503633d4b26c322e4bb431a7f5879a5

Observation 9eeaf62f-5c89-4be2-be2a-02d3f7d04fb8 · outbound

This paper cites Realfred: An em- bodied instruction following benchmark in photo-realistic environments.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Realfred: An em- bodied instruction following benchmark in photo-realistic environments

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.371553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.943797Z digest=sha256:4f6436531ded0262274b4379da53a9741d94d562b8e813aa6d2d1a570165f502

Observation 30f6df15-b03d-4f0c-b060-014a2e9d8e01 · outbound

This paper cites Segment anything.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Segment anything

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.357025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.948191Z digest=sha256:cd9da1b0713c56062790ac92b7f73b232c15e680b2445659330dd02e741b82a6

Observation b1919572-cd83-449d-826b-e48b8505acfe · outbound

This paper cites UniCLIP: Unified Framework for Contrastive Language-Image Pre-training.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation UniCLIP: Unified Framework for Contrastive Language-Image Pre-training

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:02:33.577190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.952478Z digest=sha256:4ad4dcd4eca0db6a6a2dff4e0990c55d877a0d57170d747185669da977d7c92d

Observation 8207f357-d3b8-407e-a437-96f5b9a07aa8 · outbound

This paper cites Less is more: Clipbert for video-and-language learning via sparse sampling.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Less is more: Clipbert for video-and-language learning via sparse sampling

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.342337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.956776Z digest=sha256:3e0ab47a6806ad56b70d82ea016ac373fc846df5f775b4a242a2d6e1dd9cf735

Observation ad92fa97-2683-4468-9993-289fb62236ff · outbound

This paper cites BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.960840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.960840Z digest=sha256:9f307b0d1dd5dc483b1976a6858cf70a1caad21009f5bece27067cb76fa8a6ab

Observation 7d582015-4573-4d72-90f1-53a9abbd8ed0 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.966191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.966191Z digest=sha256:0376325358ae5ac95e1b152c4c318d47da089ba0c8e719b3e9183fd3aa36740f

Observation 2a1c475a-e27f-4406-9282-799d56d74131 · outbound

This paper cites Panoptic segformer: Delving deeper into panoptic segmen- tation with transformers.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Panoptic segformer: Delving deeper into panoptic segmen- tation with transformers

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.328116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.971408Z digest=sha256:dd34a72200895a78ef40e7d607ae49b8f244d3309775f91dec1f80b0b75e8e98

Observation 53c90110-e900-4bf6-b726-6d8daf6d9d1e · outbound

This paper cites Code as policies: Language model programs for embodied con- trol.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Code as policies: Language model programs for embodied con- trol

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.312598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.975605Z digest=sha256:8b6da54c8e2fe357ecb6238628b1361f8a5d3ddb7ee8c2f9e7da52a0750683a3

Observation b986deef-cc20-4681-a7cf-88509af4eabd · outbound

This paper cites NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.980063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.980063Z digest=sha256:dd2cbb513f50d411e1e510904e9926457df910e56594d85677876efcc5f2c1bf

Observation 2d1902b8-80a2-4ddf-90eb-aa605a81c199 · outbound

This paper cites Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.984728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.984728Z digest=sha256:323bb3928a6a0f30669484c3768e0be6407d2ad5cc68697f0c4b472ec5ab09ba

Observation c545085f-88f9-4e78-bb8c-50e4337486a6 · outbound

This paper cites Sqa3d: Situated question answering in 3d scenes.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Sqa3d: Situated question answering in 3d scenes

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.296444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.989899Z digest=sha256:a5c3849bfb7aed3ed33321de244aafc1382f0c6209964c58b12e1fd58271b0ac

Observation 41653663-ad6f-4d35-8548-2f546cc1f18e · outbound

This paper cites Zson: Zero-shot object-goal navigation using multimodal goal embeddings.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Zson: Zero-shot object-goal navigation using multimodal goal embeddings

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.280449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.994814Z digest=sha256:ed18b8dd1c35522e8e52151d5a1719b68871df42110dfe87425e83ab8067825b

Observation e621b6db-fe2c-4a9d-8df8-dbfe3a847b75 · outbound

This paper cites Openeqa: Embodied question answering in the era of foun- dation models.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Openeqa: Embodied question answering in the era of foun- dation models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.211271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:32.999861Z digest=sha256:63dcabb63ba978a08b7dcd3b6b2a1f1228aa274cd730753e5329bae5653954fb

Observation 9b824185-5025-4395-aca8-665e709e9b42 · outbound

This paper cites Openeqa: Embodied question answering in the era of foun- dation models.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Openeqa: Embodied question answering in the era of foun- dation models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.152079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.005435Z digest=sha256:da893551ca87cd77b53541f00d6200464bf84bd0bbf31c9a0ecb8cd41e59af84

Observation bb3fbf5b-e554-44e3-9818-b6a3b50163c3 · outbound

This paper cites Spatial memory.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Spatial memory

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.114012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.010136Z digest=sha256:adf928356d440a535c10ae4a6fcb124ab9957b789e5dac7871774849a171ce67

Observation 6da5b11a-36a1-404a-b98c-a3b75be6c88e · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation DINOv2: Learning Robust Visual Features without Supervision

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.014670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.014670Z digest=sha256:382ad3e8ddeb53853bc6d79721a537117f8f4ce5ad16b487f21a73d5a33147c7

Observation 729ec64a-1e53-4c9e-af84-67004af85111 · outbound

This paper cites Teach: Task-driven embodied agents that chat.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Teach: Task-driven embodied agents that chat

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.099440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.019013Z digest=sha256:859fbc4d109ed7aeaba44cb646acac575267e1198faa5526eb1f3be0b854f6a0

Observation 91e13b98-5db9-4a06-b856-b57512660a26 · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Openscene: 3d scene understanding with open vocabularies

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.082872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.023557Z digest=sha256:0438f6f8e4da50d4bdf528cbffd7eda75cade52fd1e834c0731fd20358e39e7a

Observation 0bb53d3a-b4fb-4ce6-a6bd-76e809f22e22 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Learn- ing transferable visual models from natural language super- vision

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.067457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.028017Z digest=sha256:9bfe02fa6b47d619438dd3ec87ebe807e167e52366a07171425ff2ec84eea67b

Observation 8337e663-f1b0-4a05-9243-ec3752d2a8ee · outbound

This paper cites Pirlnav: Pretraining with imitation and rl finetuning for objectnav.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Pirlnav: Pretraining with imitation and rl finetuning for objectnav

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.051867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.032520Z digest=sha256:cb8fabc87396ced414d6652bfc41a66864ee1cfd5f767d64aed611f603d1f335

Observation 738a7c9a-cb3b-4a08-b6fb-3f9a641194f3 · outbound

This paper cites Sayplan: Ground- ing large language models using 3d scene graphs for scalable task planning.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Sayplan: Ground- ing large language models using 3d scene graphs for scalable task planning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.036765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.037520Z digest=sha256:0b97a3abe69b1bbc93e717689152d0119ca82ee4165f1e4a1f119c2822caa744

Observation 841e3bdb-4be8-4196-9e45-3ab0b52d98ca · outbound

This paper cites Explore until Confident: Efficient Exploration for Embodied Question Answering.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Explore until Confident: Efficient Exploration for Embodied Question Answering

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.042822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.042822Z digest=sha256:96c88f90c04f54b6e5f66706e56ec58d0652733c4ad1b2d8fee76f3f4f560758

Observation 2bd1b137-7252-436e-90e0-50c9dee394da · outbound

This paper cites Language- grounded indoor 3d semantic segmentation in the wild.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Language- grounded indoor 3d semantic segmentation in the wild

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.019320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.047240Z digest=sha256:d387c3fe01acc1f8dded949a5c6a55366686e2dabb58813f1e51214c1bcc0ede

Observation acc76002-445d-4d6a-a6a7-d62271306133 · outbound

This paper cites Habitat: A plat- form for embodied ai research.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Habitat: A plat- form for embodied ai research

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.002845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.051322Z digest=sha256:0982230f22e43b3261dc712bb7abd27f65950fecd7e012687243a0d2ad36f565

Observation 7bc365d5-0260-4142-b6f4-cb24fdbc7afd · outbound

This paper cites Proximal Policy Optimization Algorithms.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Proximal Policy Optimization Algorithms

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.055692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.055692Z digest=sha256:093dc6f68f37b458051c4c71441912df64c38bb9b957ef49e57b1955bd393cc0

Observation 46baae52-902b-4d3a-bb2d-173f6c84d29a · outbound

This paper cites Mask3d: Mask trans- former for 3d semantic instance segmentation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Mask3d: Mask trans- former for 3d semantic instance segmentation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.987323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.059795Z digest=sha256:54f8a6c8c271c800a746c1ff0ceb65238bc9ee5ebed4fa27234dad9901cd3f2d

Observation 0a686420-7872-4d1c-a62c-54e40f39f5ea · outbound

This paper cites Alfred: A benchmark for interpreting grounded instructions for everyday tasks.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Alfred: A benchmark for interpreting grounded instructions for everyday tasks

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.971219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.064643Z digest=sha256:274db12b4fb45a95f117411739cfb572f2bd8a7ec321f8aa357a268e2a201a11

Observation 170add76-540e-4418-8d40-8792d0b32b8c · outbound

This paper cites Llm-planner: Few-shot grounded planning for embodied agents with large language models.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Llm-planner: Few-shot grounded planning for embodied agents with large language models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.954645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.069207Z digest=sha256:fe527a63b5b9708baaf227b10da16890d03b7557691fca33599a6dafc97242cd

Observation 9bf74e8c-1c66-4fcf-9d3e-71c08d19150b · outbound

This paper cites Habitat 2.0: Training home assistants to rearrange their habitat.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Habitat 2.0: Training home assistants to rearrange their habitat

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.073653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.073653Z digest=sha256:d20d240e9788ddcb24fb3dd45fb607092e9e7fa989241e15be126c400b821241

Observation c9530d43-f3e2-4889-b87e-9045a8cd09d1 · outbound

This paper cites Sumner, Marc Pollefeys, Federico Tombari, and Francis Engelmann.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Sumner, Marc Pollefeys, Federico Tombari, and Francis Engelmann

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.927417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.078435Z digest=sha256:5a07ec02dd7a40e35f06e71f678cbd1ae61bae3bc56a1ed125ebb2a5e2ecbd87

Observation c9361746-d768-4873-a170-e937b6194819 · outbound

This paper cites Rio: 3d object instance re- localization in changing indoor environments.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Rio: 3d object instance re- localization in changing indoor environments

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.911980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.083330Z digest=sha256:d8a852ddf7d11ad53f6ce9fee2713b0ce3f739228482f4f244b65d48067c8652

Observation c3958155-876a-4424-a8c8-c14914faa310 · outbound

This paper cites Embodiedscan: A holistic multi- modal 3d perception suite towards embodied ai.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Embodiedscan: A holistic multi- modal 3d perception suite towards embodied ai

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.895603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.088684Z digest=sha256:64e407488070ce33501a637f32e1cf29f6ed0a5c1b5807f35464dd8d40203d9a

Observation 9840311b-3a6c-44c1-b628-fe68e2d051ac · outbound

This paper cites DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion Frames.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion Frames

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.094358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.094358Z digest=sha256:bd336fc6854368b5f57b62ab9e0509892865624ec87cf076ec18d92e02c4b64e

Observation d855c8e7-24ce-4c92-bf00-1cde8f9bce5c · outbound

This paper cites Ver: Scaling on- policy rl leads to the emergence of navigation in embodied rearrangement.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Ver: Scaling on- policy rl leads to the emergence of navigation in embodied rearrangement

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.880023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.099421Z digest=sha256:b929b905717182481b39ecbb5718af3cb1e0930803b121ceebfdf6125aba28f0

Observation c4dbf192-0e51-4d22-9090-66021d4bdc9b · outbound

This paper cites Scenegraphfusion: Incre- mental 3d scene graph prediction from rgb-d sequences.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scenegraphfusion: Incre- mental 3d scene graph prediction from rgb-d sequences

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.865248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.103822Z digest=sha256:4b3af0abab59369c697b7b690a6a0ded9b95c4f0e7ce9be7b2326abf1238bc5f

Observation 73cebb80-6c31-4bd7-be82-84128aef8574 · outbound

This paper cites EmbodiedSAM: Online Segment Any 3D Thing in Real Time.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation EmbodiedSAM: Online Segment Any 3D Thing in Real Time

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.108579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.108579Z digest=sha256:dbff194bb0c29a9eef55ed235607960c68365595356ba971bb063658398c3332

Observation 8f204022-07ac-431f-b824-1607b07efda7 · outbound

This paper cites Habitat-matterport 3d semantics dataset.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Habitat-matterport 3d semantics dataset

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.850264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.113772Z digest=sha256:f60c97c0883518c6e0ef9f487a20b9967bb6ed022aed1056f06c7344ccb72dee

Observation 66540fd2-5ebc-42da-8a41-49a9ceb41a84 · outbound

This paper cites Frontier-based exploration using multiple robots.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Frontier-based exploration using multiple robots

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.835273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.118819Z digest=sha256:c81316ba5bd1bc527b3387b34b5e41c11671a197aa9b35f198df27aa3921b53e

Observation 1525a073-ef2b-44c6-bb44-cf0d9203795b · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.123605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.123605Z digest=sha256:bd69f93d1eaa425d2cd914059f97d788b2198be1bbad83a3986338fef2f87ed6

Observation 6b28ee52-6bb1-45e5-a211-0a0961a5e8fc · outbound

This paper cites 3d-mem: 3d scene memory for embodied exploration and reasoning.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 3d-mem: 3d scene memory for embodied exploration and reasoning

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.819523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.128531Z digest=sha256:c8037c04c838cd3dc04861d35d234386169f541ecdd0a9cc2fb4374a7cee6080

Observation f7902776-3426-4812-ba1d-835c3e02c367 · outbound

This paper cites Vlfm: Vision-language frontier maps for zero-shot semantic navigation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Vlfm: Vision-language frontier maps for zero-shot semantic navigation

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.804490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.133341Z digest=sha256:310ef89651e83df76773ff4bfb454818a564dd123955785e08d5ec1907464645

Observation e846a639-0e12-4c7f-96d5-1bab80731938 · outbound

This paper cites Hm3d-ovon: A dataset and bench- mark for open-vocabulary object goal navigation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Hm3d-ovon: A dataset and bench- mark for open-vocabulary object goal navigation

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.788045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.138214Z digest=sha256:2d29c56be198b9f3ab5a9d7a6f718867e261789efa5928628233fbaf6260f4c3

Observation 0e5784ad-8c37-4039-8d60-94a6b91f6d0e · outbound

This paper cites Frontier semantic exploration for visual target navigation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Frontier semantic exploration for visual target navigation

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.770186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.142761Z digest=sha256:647cce5bc8c3c3f29c7f0a4d2b14246b299089ded9e15c1f5abe753d3c123b65

Observation aae9a2f4-215d-4950-8a5f-1f54cf5c022c · outbound

This paper cites Instancere- fer: Cooperative holistic understanding for visual ground- ing on point clouds through instance multi-level contextual referring.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Instancere- fer: Cooperative holistic understanding for visual ground- ing on point clouds through instance multi-level contextual referring

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.754549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.147218Z digest=sha256:806edc3b0d1ec487e940b0ca9f61dd8cd40897c6f6dc0d4e8c6199bc229f5a70

Observation 2609d05e-91ce-4d80-b285-5d697180c3b5 · outbound

This paper cites PoliFormer: Scaling On-Policy RL with Transformers Results in Masterful Navigators.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation PoliFormer: Scaling On-Policy RL with Transformers Results in Masterful Navigators

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.151309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.151309Z digest=sha256:c8ac5db696e3514585aa2e4ca9ea4906c780e959489ab55242dbd7efd1927d86

Observation 583e3afd-8e7f-47ef-9722-8da0e9f8a503 · outbound

This paper cites Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.155998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.155998Z digest=sha256:e27546092f9b3024ae08c34fdb4f945d153cc9d59225a527dc8ea82ee6b61215

Observation 6063f092-66ec-41ba-84c4-cb0f6f1045a0 · outbound

This paper cites Vision-Language Pre-training with Object Contrastive Learning for 3D Scene Understanding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Vision-Language Pre-training with Object Contrastive Learning for 3D Scene Understanding

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:02:33.334560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.160598Z digest=sha256:6f030e72770e3d8f44983e0f64b17d6911a73d7f9626f53dfb6d8dbc973b46b0

Observation 5f9780c8-10d5-4986-8f9a-4d97c8147ca9 · outbound

This paper cites Multi3drefer: Grounding text description to multiple 3d objects.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Multi3drefer: Grounding text description to multiple 3d objects

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.740367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.165673Z digest=sha256:ae68f284c327d8191be57a1c8ff71b8894daf114949feaa229d501ea428c7710

Observation f7906a49-1d19-4666-ad00-b2842a1db9ca · outbound

This paper cites Microsoft kinect sensor and its effect.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Microsoft kinect sensor and its effect

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.723322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.170543Z digest=sha256:0605e761717eb9fcba4a7f7d6f370c9a33339289b080639d38378b6412242c27

Observation f7fba5c5-b1d6-49a2-a672-f6d67140bc4e · outbound

This paper cites Task-oriented Sequential Grounding and Navigation in 3D Scenes.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Task-oriented Sequential Grounding and Navigation in 3D Scenes

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.174687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.174687Z digest=sha256:a949e5fcfac252f7cf55b2fb5a8a5a1147cb4645e5111f591e80259bc0ccb8fb

Observation b0f133cb-4108-4947-b1fc-836c269d090e · outbound

This paper cites To- wards explainable 3d grounded visual question answering: A new benchmark and strong baseline.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation To- wards explainable 3d grounded visual question answering: A new benchmark and strong baseline

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.708124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.179080Z digest=sha256:898766c5e7d35c65a961be39889b5332b4a8d2f85907dd34d56ead266d7662ce

Observation a10528fa-1096-4e34-adf2-8d1bfbaa1104 · outbound

This paper cites Fast Segment Anything.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Fast Segment Anything

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.183289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.183289Z digest=sha256:625da61884f2c410210d970e8eae728c46bc897dd5f092b77edf7d994f42279b

Observation c0dbc718-15b3-4fef-ba94-993a9eaf9604 · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.187976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.187976Z digest=sha256:7d0f9a13e8d54d58b9cfd82a421784481ebecc0a8605c0704ff017199a01c131

Observation 83388a1b-427e-42dd-a5f5-0aa425a01d7c · outbound

This paper cites Dual memory units with uncertainty regulation for weakly supervised video anomaly detection.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Dual memory units with uncertainty regulation for weakly supervised video anomaly detection

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.690761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.193178Z digest=sha256:d5d664650ff49058eb1f22e78caa92f9df8182fd661704ab8e0ddba1fc0e6f53

Observation 96ec4986-1805-49e8-ad9c-1def13a40b0b · outbound

This paper cites Scanreason: Empowering 3d visual grounding with reasoning capabilities.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scanreason: Empowering 3d visual grounding with reasoning capabilities

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.676105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.197976Z digest=sha256:2628a216c64f5e3258f51bdc24a9df8a9adb34b93b4c315a9c40948cea8e04fc

Observation d4465aff-60cc-4572-b653-dfdb711f5ee1 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 3d-vista: Pre-trained transformer for 3d vision and text alignment

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.660133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.202562Z digest=sha256:490a722aeaf3956d7484e24c2a05f09d084f69904b035184fbeacd13c53c987e

Observation 225f2ccb-d3ce-4433-9dbc-f8934a0e1169 · outbound

This paper cites Unifying 3d vision-language understanding via prompt- able queries.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Unifying 3d vision-language understanding via prompt- able queries

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.644142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.206561Z digest=sha256:342f0701385b33be18bc7eaf05a38d292be70d1b1b828e1d4db5896fc221dd7a

Observation 17a85cb1-9423-4a72-994e-4fc5374ca733 · outbound

This paper cites TANGO: Training-free Embodied AI Agents for Open-world Tasks.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation TANGO: Training-free Embodied AI Agents for Open-world Tasks

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.210895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.210895Z digest=sha256:e5720a73d5934c7e665d12acb632175a7b25b9eaf53386cf15d68d269318feef

Observation be143c56-3127-4de0-84c2-0274233b3c85 · outbound

This paper cites Generalized decoding for pixel, image, and language.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Generalized decoding for pixel, image, and language

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.628398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:02:33.215747Z digest=sha256:3c60591b135c4f890d7ee0c08ac1c56cf72fc1171835315fcf4460fd591e048d

Pith citing papers

Observation 8b12b765-c17c-4532-8932-0c2f5ffe3a7d · inbound

SPG: Style-Prompting Guidance for Style-Specific Content Creation cites this paper.

SPG: Style-Prompting Guidance for Style-Specific Content Creation Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:57.184897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:57.184897Z digest=sha256:02fc2dec0298fc85e0b427e9380360649c144e9d0770238f807e0c028efde303

Observation 7363683b-1a7a-41e1-8f77-a918c1574094 · inbound

FSUNav: A Cerebrum-Cerebellum Architecture for Fast, Safe, and Universal Zero-Shot Goal-Oriented Navigation cites this paper.

FSUNav: A Cerebrum-Cerebellum Architecture for Fast, Safe, and Universal Zero-Shot Goal-Oriented Navigation Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:03:08.600996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T18:58:48.779518Z digest=sha256:b6ec0d8fae7790d687ce341ea6de58721c6a05adcd7827fab599d635a066ef09