Pith. sign in

Paper Citation Record · LEDGER

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation

As of 20 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 2 inbound Pith citation observations for arXiv:2507.04047.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04047 v2

Coverage vector

measured 96 of 96 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:02:33.215747Z

measured 98 of 98 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T19:54:57.184897Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T19:03:08.599602Z

Reference resolution

96 of 96 outbound references displayed

  • verified exact2
  • verified fuzzy59
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e2a31111-e2e4-4890-a72a-2b5fe7745999 · outbound

This paper cites Scanents3d: Exploit- ing phrase-to-3d-object correspondences for improved visio- linguistic models in 3d scenes.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scanents3d: Exploit- ing phrase-to-3d-object correspondences for improved visio- linguistic models in 3d scenes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.156912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.156912Z digest=sha256:53f74c7a87b323fd575b5cf979544e2ef4210c2c4ef2b34c8aae6973265569f1

Observation 7f2de7b2-1a14-431d-a5c8-30bbdd9a6843 · outbound

This paper cites Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.251850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.251850Z digest=sha256:715d7603b353fed0937905b455defd3af52d16c2b26732f89c157c98155ececf

Observation d4f365fa-c59c-42b1-9e52-546e5375435c · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scanqa: 3d question answering for spatial scene understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.317126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.317126Z digest=sha256:5244c471bdfb561339820b2e459713dbe0abc6e5834aa092ed0ee5ee79a0d2f2

Observation c49bc4df-6ccb-4a9e-9bbe-fd1b50ea2390 · outbound

This paper cites Do as i can, not as i say: Grounding language in robotic affordances.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Do as i can, not as i say: Grounding language in robotic affordances

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.434833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.434833Z digest=sha256:1fc06cbb12f0100954f9970cee2cfa07d593fd6f0cdff16265d012d27c7601e0

Observation bb7ccfdb-041b-411f-993c-23f9995f6668 · outbound

This paper cites 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.497301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.497301Z digest=sha256:c27b1b8b14e948689e2838995c35cc2dfc1f6e97ab4eb722d8bea2da5a5bcc40

Observation f9dddfbd-8cbb-48ea-b292-57527f9aadad · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Emerg- ing properties in self-supervised vision transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.605674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.605674Z digest=sha256:164bf90d3796f02520a7c747700c8d7b7067ce3e80f3a572825e1761741d3910

Observation deda0b58-fc06-49a7-b970-3e239931f8ae · outbound

This paper cites Object goal naviga- tion using goal-oriented semantic exploration.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Object goal naviga- tion using goal-oriented semantic exploration

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.679473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.679473Z digest=sha256:66fd42e11bbba9b526f3133dec9d088eea6985c1c9663b891e7239e2598b0ce8

Observation b823c8a9-85dc-4284-8432-c24334a7769b · outbound

This paper cites Object goal navi- gation using goal-oriented semantic exploration.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Object goal navi- gation using goal-oriented semantic exploration

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.763679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.763679Z digest=sha256:8456fe6fb042fe5c4d1f03928388590c1be0c6fc8376b7eba9da9f923029c017

Observation 86802e41-5707-4c97-87d5-a88d5deb6b24 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natu- ral language.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scanrefer: 3d object localization in rgb-d scans using natu- ral language

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.842833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.842833Z digest=sha256:e9b08b6b52bf894e2f593a65402cf44306f293be03d7cb0a7407ea68d4b19d1a

Observation 06333299-344d-4207-b617-3a79b9607d6c · outbound

This paper cites Language conditioned spatial relation reasoning for 3d object grounding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Language conditioned spatial relation reasoning for 3d object grounding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.924326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.924326Z digest=sha256:79b82707902326ca601587c774416d9fcc59775ce34b92c5ec18bec0ead004b2

Observation 778803a5-506f-493f-99d9-949b414b8d32 · outbound

This paper cites Ll3da: Visual interactive instruction tuning for omni-3d understand- ing reasoning and planning.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Ll3da: Visual interactive instruction tuning for omni-3d understand- ing reasoning and planning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.014240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.014240Z digest=sha256:56b61e2860177f3d9ddc1a1a1b2ec1e6ac3fa45f16783467a34b94c45f8a9dde

Observation e4822169-e117-457f-8f5d-242e6d091ec3 · outbound

This paper cites Scan2cap: Context-aware dense captioning in rgb- d scans.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scan2cap: Context-aware dense captioning in rgb- d scans

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.057605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.057605Z digest=sha256:aa342487ffbf1e0ef1c1e489eee46636a410e5fe54fc1a4de2794f865f910861

Observation 26c200ce-8196-49a5-a404-ede974461053 · outbound

This paper cites Unit3d: A unified trans- former for 3d dense captioning and visual grounding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Unit3d: A unified trans- former for 3d dense captioning and visual grounding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.194145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.194145Z digest=sha256:c649e335d41cc34b482c8c2ce63d16f3bcfc39fe93123bee24aa24a8ccad1326

Observation 96a42d0e-704e-4d5e-b86c-f82703300d84 · outbound

This paper cites Schwing, Alexan- der Kirillov, and Rohit Girdhar.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Schwing, Alexan- der Kirillov, and Rohit Girdhar

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.284849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.284849Z digest=sha256:11508da67ded574ffec66c948e8ca3da28b45bdffcf53b8f40e9f0a66aba963c

Observation ed1e8638-0b8a-4a26-9984-d39ecfa84001 · outbound

This paper cites 4d spatio-temporal convnets: Minkowski convolutional neural networks.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 4d spatio-temporal convnets: Minkowski convolutional neural networks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.391830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.391830Z digest=sha256:54e297d987be4bbd31bce76e0432e3018dcc6baa2bc69a33c8100410bd7a9ee1

Observation 17ccbf1e-189e-4f93-b408-cafceb94bcaf · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.507880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.507880Z digest=sha256:657cf93f586dde9b3035f7a97f35a70b711591c1f6e090716b12cd8cfc15e416

Observation 23dd4bc8-824e-48b8-99bd-72233555fe15 · outbound

This paper cites Embodied question answer- ing.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Embodied question answer- ing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.662553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.605207Z digest=sha256:961876f3b0d5cd47cb72bd6bdd41606119db35d7e523c5c1404982951fe224a8

Observation fec0ec8e-b139-451f-b6b8-91dcf8ba2993 · outbound

This paper cites A survey of embodied ai: From simulators to research tasks.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation A survey of embodied ai: From simulators to research tasks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.648665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.730308Z digest=sha256:0128d528fe930ecbd9d7dff2030774c28432987df59dc041bf9856838f9997a0

Observation 6db28803-3cf2-4251-ad77-654dcd228b67 · outbound

This paper cites The One RING: a Robotic Indoor Navigation Generalist.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation The One RING: a Robotic Indoor Navigation Generalist

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.828420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.828420Z digest=sha256:8266986b28e42fbcb1d35b95daf312ea59e63ac7b79ef33e7e669b29e365c38b

Observation 4a984db5-a1cb-4e85-842d-0a3a593d6b52 · outbound

This paper cites Spoc: Imitating short- est paths in simulation enables effective navigation and ma- nipulation in the real world.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Spoc: Imitating short- est paths in simulation enables effective navigation and ma- nipulation in the real world

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.633572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.859253Z digest=sha256:14f7bf3dddbeae32e9679dce1022dd83d7ed1bd53a096de3b29cf5a8b063efce

Observation 713006fa-6b05-414d-a629-451cd2717320 · outbound

This paper cites Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.864163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.864163Z digest=sha256:e91ef7ab246fdf8afc1fc0e46dd3e4e79382c9b0c1e9169c09ccf34aee6cd709

Observation 07a13e23-4b26-4210-91eb-1174c07e9df7 · outbound

This paper cites Efficient graph-based image segmentation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Efficient graph-based image segmentation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.618261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.869619Z digest=sha256:6a11cf4663c118c783adec111c71310b01562f001cb95d4e7466e1b895e847e9

Observation 1a3fd2f0-a809-4758-ad4d-a2be5e0be2cc · outbound

This paper cites Cows on pasture: Base- lines and benchmarks for language-driven zero-shot object navigation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Cows on pasture: Base- lines and benchmarks for language-driven zero-shot object navigation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.604356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.873766Z digest=sha256:f4d519cc252f5b719e638e852cb94e75e2dbadfa3b5cb9cb18ba74a48eef7f8b

Observation 8c68ce8f-f0a4-4e70-9234-da915752a457 · outbound

This paper cites Scaling open-vocabulary image segmentation with image- level labels.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scaling open-vocabulary image segmentation with image- level labels

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.589762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.878891Z digest=sha256:ffcbfaee8a98a0384b3428571956d09ebd4631102b5794ffd3d9f39ecb1e005d

Observation 45f432e4-e104-4efe-9322-2a706db8a7fa · outbound

This paper cites Conceptgraphs: Open-vocabulary 3d scene graphs for per- ception and planning.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Conceptgraphs: Open-vocabulary 3d scene graphs for per- ception and planning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.575555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.883132Z digest=sha256:dd8cd7ef7f33994eb3f40a8a38b64f8327a900c491bd32327154c36bb44e4541

Observation adf93ba6-943c-4359-bfef-d23cc8551d51 · outbound

This paper cites Viewrefer: Grasp the multi-view knowledge for 3d visual grounding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Viewrefer: Grasp the multi-view knowledge for 3d visual grounding

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.561287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.887401Z digest=sha256:3b21292e34883656cc4f1e6977a40539ae2f767596c389dd2c31be8db46440ca

Observation 7d587f8b-2923-40f4-8826-daaabe53f5d4 · outbound

This paper cites Transrefer3d: Entity-and- relation aware transformer for fine-grained 3d visual ground- ing.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Transrefer3d: Entity-and- relation aware transformer for fine-grained 3d visual ground- ing

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.546556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.891844Z digest=sha256:b7d0f561ad30c9eb2135bce500f619e1963e1bbea70e0a8c192af88b7ccd31bd

Observation 0a0a1021-5c64-4d71-8425-ea1db88b305a · outbound

This paper cites Vln bert: A recurrent vision- and-language bert for navigation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Vln bert: A recurrent vision- and-language bert for navigation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.531842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.896182Z digest=sha256:b297823a55234d31e4388eef97418c804661ee63742912f26174594e7ce23c64

Observation d4c652b4-9733-408a-8d2e-c5eb0468f7ee · outbound

This paper cites 3d-llm: In- jecting the 3d world into large language models.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 3d-llm: In- jecting the 3d world into large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.515695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.901425Z digest=sha256:b9fc330cc2b2d60120e7dd0db2eff12a73aa1c80899d631d201831589b045e65

Observation 7f29a1b4-1c1e-4eb9-955e-9b9381b77dcd · outbound

This paper cites A real-time occupancy map from multiple video streams.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation A real-time occupancy map from multiple video streams

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.498323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.906117Z digest=sha256:385857ed29000e1360f17936a18b5a0c09280023548bc9e4a4488457919a5027

Observation 9f43c2d6-7681-47f4-b4a6-54b61de930e7 · outbound

This paper cites An embodied generalist agent in 3d world.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation An embodied generalist agent in 3d world

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.481002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.910518Z digest=sha256:4ff7cc06e59380051187159e08756ea59c4a35b130167c7b4913d3bf704ff9f9

Observation 4c5594b4-3ee3-422a-b71b-7dba27b78f5f · outbound

This paper cites Language models as zero-shot planners: Ex- tracting actionable knowledge for embodied agents.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Language models as zero-shot planners: Ex- tracting actionable knowledge for embodied agents

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.463894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.914903Z digest=sha256:c337060541ed5560f80ef8fed519aabbf67498a340605875ac47238f498fbd59

Observation 143deb32-0c82-488b-81f8-15730725a8d9 · outbound

This paper cites V oxposer: Composable 3d value maps for robotic manipulation with language models.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation V oxposer: Composable 3d value maps for robotic manipulation with language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.448269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.919522Z digest=sha256:9489d0944493718ec03edf3f3585f007c243efeff47956b06744c55fd9f3d292

Observation 39bbe7cb-7eaf-46c9-a293-620806b5320b · outbound

This paper cites Bottom up top down detection transform- ers for language grounding in images and point clouds.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Bottom up top down detection transform- ers for language grounding in images and point clouds

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.433025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.924413Z digest=sha256:347dc7e5f016597e266abc28ecad704d1f7ce5560d7f5a7c4d6aa5822a514977

Observation 8f216072-19fc-4d8f-9511-ee0f109fc652 · outbound

This paper cites Conceptfusion: Open-set multi- modal 3d mapping.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Conceptfusion: Open-set multi- modal 3d mapping

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.417777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.929001Z digest=sha256:d214be647aae618b4a672d4b5fb7a485792aba5eb6454d4a0311844b93aefabc

Observation e52ef426-8946-4d81-856e-63b42e949343 · outbound

This paper cites Sceneverse: Scaling 3d vision-language learning for grounded scene understanding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Sceneverse: Scaling 3d vision-language learning for grounded scene understanding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.402368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.934381Z digest=sha256:8a4c156e115ce32dc904ecfdb8630ce1982c99159bab59fa16d6402c165fe5fb

Observation b46b61ae-463b-42ff-a9ca-487841e009f3 · outbound

This paper cites Goat-bench: A benchmark for multi-modal lifelong navigation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Goat-bench: A benchmark for multi-modal lifelong navigation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.386201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.939241Z digest=sha256:764f01db04baaa6d2dc27ba5be8092ce718e926c0ebd1c025e85d651fa86ce76

Observation 9eeaf62f-5c89-4be2-be2a-02d3f7d04fb8 · outbound

This paper cites Realfred: An em- bodied instruction following benchmark in photo-realistic environments.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Realfred: An em- bodied instruction following benchmark in photo-realistic environments

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.371553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.943797Z digest=sha256:ee7fccd403ff540342190dd911844ecf69838d1b056acce0bd5521340e21a883

Observation 30f6df15-b03d-4f0c-b060-014a2e9d8e01 · outbound

This paper cites Segment anything.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Segment anything

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.357025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.948191Z digest=sha256:f5fa224b964fee92dc8bf35d1c4ef2a3c91f3a075850b12bb9bf307921378a61

Observation b1919572-cd83-449d-826b-e48b8505acfe · outbound

This paper cites UniCLIP: Unified Framework for Contrastive Language-Image Pre-training.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation UniCLIP: Unified Framework for Contrastive Language-Image Pre-training

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:02:33.577190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.952478Z digest=sha256:4201e55952016bccfada9deb64277d5b018928bcc143abeb10fcb7f20389590b

Observation 8207f357-d3b8-407e-a437-96f5b9a07aa8 · outbound

This paper cites Less is more: Clipbert for video-and-language learning via sparse sampling.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Less is more: Clipbert for video-and-language learning via sparse sampling

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.342337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.956776Z digest=sha256:1bcb628478bd8ca908cfae81263879c692a51bd8bb99b1668c64ed56031175db

Observation ad92fa97-2683-4468-9993-289fb62236ff · outbound

This paper cites BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.960840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.960840Z digest=sha256:6e13f225c4235ec922071897cd2ae7f459e245b8466b75b1e7958e75df4696b8

Observation 7d582015-4573-4d72-90f1-53a9abbd8ed0 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.966191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.966191Z digest=sha256:0e2ed4b1ac8f28c56d47c412a28cb21151e4afb04b8f7da5e77fc958df405cd4

Observation 2a1c475a-e27f-4406-9282-799d56d74131 · outbound

This paper cites Panoptic segformer: Delving deeper into panoptic segmen- tation with transformers.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Panoptic segformer: Delving deeper into panoptic segmen- tation with transformers

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.328116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.971408Z digest=sha256:3186d55ea987646e1ece5d7ca0aca362434446faec2996deec6ba6a1cc0c323d

Observation 53c90110-e900-4bf6-b726-6d8daf6d9d1e · outbound

This paper cites Code as policies: Language model programs for embodied con- trol.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Code as policies: Language model programs for embodied con- trol

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.312598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.975605Z digest=sha256:cb1c92cb6ea4f8f899856e3bf2a29a1250b959dd869588b450366469ba0e4855

Observation b986deef-cc20-4681-a7cf-88509af4eabd · outbound

This paper cites NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.980063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.980063Z digest=sha256:979fc0bceacb9c527360b216e3cd00b54e0b2fef84a7fd02795bf7e73990acd7

Observation 2d1902b8-80a2-4ddf-90eb-aa605a81c199 · outbound

This paper cites Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.984728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.984728Z digest=sha256:c893c25aa77a278664d078ceb4d6f1fbae46b5cc9ecdbed146f47743be687512

Observation c545085f-88f9-4e78-bb8c-50e4337486a6 · outbound

This paper cites Sqa3d: Situated question answering in 3d scenes.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Sqa3d: Situated question answering in 3d scenes

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.296444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.989899Z digest=sha256:641c43eeea3f674a5594b543027294154ca58a8fadf559e6ef0ad6e4bb6dfe47

Observation 41653663-ad6f-4d35-8548-2f546cc1f18e · outbound

This paper cites Zson: Zero-shot object-goal navigation using multimodal goal embeddings.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Zson: Zero-shot object-goal navigation using multimodal goal embeddings

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.280449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.994814Z digest=sha256:bb2c06cca7385f97a873d29f111bb12b7de30684a9d30d03a1a3d57036e71495

Observation e621b6db-fe2c-4a9d-8df8-dbfe3a847b75 · outbound

This paper cites Openeqa: Embodied question answering in the era of foun- dation models.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Openeqa: Embodied question answering in the era of foun- dation models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.211271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:32.999861Z digest=sha256:66ceb2c54cedc453f4b86aa6df448acea5bc336a27d0eb05468b89a7d5004e31

Observation 9b824185-5025-4395-aca8-665e709e9b42 · outbound

This paper cites Openeqa: Embodied question answering in the era of foun- dation models.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Openeqa: Embodied question answering in the era of foun- dation models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.152079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.005435Z digest=sha256:e60b020969719796bd0fda3900da05fbb8f4af4e54565833e7b91a86869b353c

Observation bb3fbf5b-e554-44e3-9818-b6a3b50163c3 · outbound

This paper cites Spatial memory.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Spatial memory

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.114012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.010136Z digest=sha256:83a024376f64ce2acea67fdecf9063242e55d0a1807ae03d4cee4cfdb2300c97

Observation 6da5b11a-36a1-404a-b98c-a3b75be6c88e · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation DINOv2: Learning Robust Visual Features without Supervision

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.014670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.014670Z digest=sha256:8cb8e36859f260fa7c04a2eb482c64c59169f43a6ff33cd1eea6e83d2f238839

Observation 729ec64a-1e53-4c9e-af84-67004af85111 · outbound

This paper cites Teach: Task-driven embodied agents that chat.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Teach: Task-driven embodied agents that chat

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.099440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.019013Z digest=sha256:cb90864295325ed29ca915d2d5016cc3601a610ecbe14ebb87cace18d567b147

Observation 91e13b98-5db9-4a06-b856-b57512660a26 · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Openscene: 3d scene understanding with open vocabularies

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.082872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.023557Z digest=sha256:9b872a4b6bee4604fc409539b1f76fec19f95075a2a03349b74c45159e149b4c

Observation 0bb53d3a-b4fb-4ce6-a6bd-76e809f22e22 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Learn- ing transferable visual models from natural language super- vision

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.067457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.028017Z digest=sha256:76b888f735ddb49ee8048ab4bb8441fb209b9f56ae544a8f3aef5243d3345681

Observation 8337e663-f1b0-4a05-9243-ec3752d2a8ee · outbound

This paper cites Pirlnav: Pretraining with imitation and rl finetuning for objectnav.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Pirlnav: Pretraining with imitation and rl finetuning for objectnav

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.051867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.032520Z digest=sha256:c93df675980947b20aeaccadbc76925decfb0d95b218afecfcf62759a24128d8

Observation 738a7c9a-cb3b-4a08-b6fb-3f9a641194f3 · outbound

This paper cites Sayplan: Ground- ing large language models using 3d scene graphs for scalable task planning.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Sayplan: Ground- ing large language models using 3d scene graphs for scalable task planning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.036765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.037520Z digest=sha256:aa872e4119a63cdffd7c1b0fa50b08461a70f4a276986513fcfcf1713004b85f

Observation 841e3bdb-4be8-4196-9e45-3ab0b52d98ca · outbound

This paper cites Explore until Confident: Efficient Exploration for Embodied Question Answering.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Explore until Confident: Efficient Exploration for Embodied Question Answering

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.042822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.042822Z digest=sha256:31ede6fbfbe86e16cb52eadbf14843271ec2d5e7957e66b25d79ec0c8421d5fc

Observation 2bd1b137-7252-436e-90e0-50c9dee394da · outbound

This paper cites Language- grounded indoor 3d semantic segmentation in the wild.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Language- grounded indoor 3d semantic segmentation in the wild

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.019320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.047240Z digest=sha256:9aa72046c915eff38623122cf636562a57b70fda3e72de9b66a5a23f635cb055

Observation acc76002-445d-4d6a-a6a7-d62271306133 · outbound

This paper cites Habitat: A plat- form for embodied ai research.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Habitat: A plat- form for embodied ai research

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.002845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.051322Z digest=sha256:fc217ae9295c9c256fa69bb91f221ae2587e49a888b378a67d3767dd322c354e

Observation 7bc365d5-0260-4142-b6f4-cb24fdbc7afd · outbound

This paper cites Proximal Policy Optimization Algorithms.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Proximal Policy Optimization Algorithms

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.055692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.055692Z digest=sha256:a222d624974f416b5480fcb2508f2ba58176f2244a3b1cb2a336a5676bff5123

Observation 46baae52-902b-4d3a-bb2d-173f6c84d29a · outbound

This paper cites Mask3d: Mask trans- former for 3d semantic instance segmentation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Mask3d: Mask trans- former for 3d semantic instance segmentation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.987323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.059795Z digest=sha256:d88cae9abbada3e68c2d969c65f8f0172c8e303e9ecd8fd939f51997cd8ac461

Observation 0a686420-7872-4d1c-a62c-54e40f39f5ea · outbound

This paper cites Alfred: A benchmark for interpreting grounded instructions for everyday tasks.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Alfred: A benchmark for interpreting grounded instructions for everyday tasks

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.971219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.064643Z digest=sha256:1d89369e20a1d5e6ab4164af0dc042566e357ba414f5256325e9e35e34ccd61d

Observation 170add76-540e-4418-8d40-8792d0b32b8c · outbound

This paper cites Llm-planner: Few-shot grounded planning for embodied agents with large language models.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Llm-planner: Few-shot grounded planning for embodied agents with large language models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.954645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.069207Z digest=sha256:bbbd56c2844d61eb6245e7d7500ec141ade8c8e9079b13a77f75f808b89873ea

Observation 9bf74e8c-1c66-4fcf-9d3e-71c08d19150b · outbound

This paper cites Habitat 2.0: Training home assistants to rearrange their habitat.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Habitat 2.0: Training home assistants to rearrange their habitat

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.073653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.073653Z digest=sha256:8abc929e80c4d36a65d6c74d1c9666a41c78877b69b77cb7479bf79c22caf8b8

Observation c9530d43-f3e2-4889-b87e-9045a8cd09d1 · outbound

This paper cites Sumner, Marc Pollefeys, Federico Tombari, and Francis Engelmann.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Sumner, Marc Pollefeys, Federico Tombari, and Francis Engelmann

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.927417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.078435Z digest=sha256:1dd08b5e0b4b3b4e2c2ed6d7d4475a9ff7652e9d58b7844247991c6a00a91067

Observation c9361746-d768-4873-a170-e937b6194819 · outbound

This paper cites Rio: 3d object instance re- localization in changing indoor environments.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Rio: 3d object instance re- localization in changing indoor environments

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.911980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.083330Z digest=sha256:982f486e781360ed81832aa8d9a92b75db12f0769f46b20eca3ce2f2248e1423

Observation c3958155-876a-4424-a8c8-c14914faa310 · outbound

This paper cites Embodiedscan: A holistic multi- modal 3d perception suite towards embodied ai.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Embodiedscan: A holistic multi- modal 3d perception suite towards embodied ai

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.895603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.088684Z digest=sha256:71153b51c590b2ac8fc8a9cec56d5936180c9b90afcbcdf4bc7005d4ba037aac

Observation 9840311b-3a6c-44c1-b628-fe68e2d051ac · outbound

This paper cites DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion Frames.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion Frames

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.094358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.094358Z digest=sha256:6e19313941c80e7ba429750c6c55498dffb2c10cf341f0ac9eaf12c6f2eae403

Observation d855c8e7-24ce-4c92-bf00-1cde8f9bce5c · outbound

This paper cites Ver: Scaling on- policy rl leads to the emergence of navigation in embodied rearrangement.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Ver: Scaling on- policy rl leads to the emergence of navigation in embodied rearrangement

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.880023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.099421Z digest=sha256:0da24de27ce1f482b1850bfd6ecdabbc496fbd7a0e6d6aefef652931c44be298

Observation c4dbf192-0e51-4d22-9090-66021d4bdc9b · outbound

This paper cites Scenegraphfusion: Incre- mental 3d scene graph prediction from rgb-d sequences.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scenegraphfusion: Incre- mental 3d scene graph prediction from rgb-d sequences

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.865248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.103822Z digest=sha256:e8252eafbc6aad8a200af0df3df28975768fff0d2313bf426908289704abbbcd

Observation 73cebb80-6c31-4bd7-be82-84128aef8574 · outbound

This paper cites EmbodiedSAM: Online Segment Any 3D Thing in Real Time.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation EmbodiedSAM: Online Segment Any 3D Thing in Real Time

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.108579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.108579Z digest=sha256:099a27fb9495f44b221c2091606d6d2d12c6079edb00ec3e6d7ea506e804eb25

Observation 8f204022-07ac-431f-b824-1607b07efda7 · outbound

This paper cites Habitat-matterport 3d semantics dataset.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Habitat-matterport 3d semantics dataset

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.850264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.113772Z digest=sha256:76686f10fdac383ae550794c9330244172ce914e676c008c39efe2b206b2cfee

Observation 66540fd2-5ebc-42da-8a41-49a9ceb41a84 · outbound

This paper cites Frontier-based exploration using multiple robots.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Frontier-based exploration using multiple robots

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.835273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.118819Z digest=sha256:6ae1ad69fa4c889f1c50e91f3c01f18d1178957ef3254fc693922fbe264cbc8e

Observation 1525a073-ef2b-44c6-bb44-cf0d9203795b · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.123605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.123605Z digest=sha256:ef25c1ed6ebcaf75613e940e4c57e1b8b209a01e8e430b9e8e6ace68c78ff351

Observation 6b28ee52-6bb1-45e5-a211-0a0961a5e8fc · outbound

This paper cites 3d-mem: 3d scene memory for embodied exploration and reasoning.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 3d-mem: 3d scene memory for embodied exploration and reasoning

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.819523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.128531Z digest=sha256:67e9601144bc9b421a634b9ef6c66b06abeda50c1dd183793cb61dba1fa91bdd

Observation f7902776-3426-4812-ba1d-835c3e02c367 · outbound

This paper cites Vlfm: Vision-language frontier maps for zero-shot semantic navigation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Vlfm: Vision-language frontier maps for zero-shot semantic navigation

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.804490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.133341Z digest=sha256:6f67cd2b34ad2e9ec317ddfb2f9cdbecae8e2a30562c2c1eb236a9c8d3e0c800

Observation e846a639-0e12-4c7f-96d5-1bab80731938 · outbound

This paper cites Hm3d-ovon: A dataset and bench- mark for open-vocabulary object goal navigation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Hm3d-ovon: A dataset and bench- mark for open-vocabulary object goal navigation

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.788045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.138214Z digest=sha256:55a5348555e31ae862d93fd3b864a04d690adbbf31a6066a11e84bc79de6959c

Observation 0e5784ad-8c37-4039-8d60-94a6b91f6d0e · outbound

This paper cites Frontier semantic exploration for visual target navigation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Frontier semantic exploration for visual target navigation

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.770186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.142761Z digest=sha256:c0fb4f7c758b55baafca37c6c0e32ee6de4bc952e63089607f0834e161a491f3

Observation aae9a2f4-215d-4950-8a5f-1f54cf5c022c · outbound

This paper cites Instancere- fer: Cooperative holistic understanding for visual ground- ing on point clouds through instance multi-level contextual referring.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Instancere- fer: Cooperative holistic understanding for visual ground- ing on point clouds through instance multi-level contextual referring

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.754549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.147218Z digest=sha256:6c03449debc4794594483f2fd11c764096892ea523410667f737cf61915484ec

Observation 2609d05e-91ce-4d80-b285-5d697180c3b5 · outbound

This paper cites PoliFormer: Scaling On-Policy RL with Transformers Results in Masterful Navigators.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation PoliFormer: Scaling On-Policy RL with Transformers Results in Masterful Navigators

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.151309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.151309Z digest=sha256:b8c19a4d26fd97d7ecfbf2ad48c0519ddc420ac0a240facfdb4c6e284f1d87e3

Observation 583e3afd-8e7f-47ef-9722-8da0e9f8a503 · outbound

This paper cites Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.155998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.155998Z digest=sha256:301c5a8b0b85c3ef4c7b607405df079abe5e2de598370c51b20533797e6dc2f5

Observation 6063f092-66ec-41ba-84c4-cb0f6f1045a0 · outbound

This paper cites Vision-Language Pre-training with Object Contrastive Learning for 3D Scene Understanding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Vision-Language Pre-training with Object Contrastive Learning for 3D Scene Understanding

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:02:33.334560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.160598Z digest=sha256:a3f15279c2f5002b44f7edb099ee82ade05cf43655c041d06687b7d2cb0e53bc

Observation 5f9780c8-10d5-4986-8f9a-4d97c8147ca9 · outbound

This paper cites Multi3drefer: Grounding text description to multiple 3d objects.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Multi3drefer: Grounding text description to multiple 3d objects

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.740367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.165673Z digest=sha256:4d26287d36e9cce367043b19bd0bc89c415bc9b3e7724945cc7adbbb88d9680e

Observation f7906a49-1d19-4666-ad00-b2842a1db9ca · outbound

This paper cites Microsoft kinect sensor and its effect.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Microsoft kinect sensor and its effect

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.723322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.170543Z digest=sha256:79d0c495fc0ea5682ade85b917f347f98d50d6fe52a5b229115a8c32d8e0c53e

Observation f7fba5c5-b1d6-49a2-a672-f6d67140bc4e · outbound

This paper cites Task-oriented Sequential Grounding and Navigation in 3D Scenes.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Task-oriented Sequential Grounding and Navigation in 3D Scenes

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.174687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.174687Z digest=sha256:ea3db42f38eb55d5750d4ab66a7e662b7c0b7dc446cc80d5c15917e34d37f3af

Observation b0f133cb-4108-4947-b1fc-836c269d090e · outbound

This paper cites To- wards explainable 3d grounded visual question answering: A new benchmark and strong baseline.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation To- wards explainable 3d grounded visual question answering: A new benchmark and strong baseline

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.708124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.179080Z digest=sha256:7ea20e1b2c5860638dc3a8fa9c3f8d087e4e518a0aef2e77bc9f7278d091fa8d

Observation a10528fa-1096-4e34-adf2-8d1bfbaa1104 · outbound

This paper cites Fast Segment Anything.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Fast Segment Anything

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.183289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.183289Z digest=sha256:a41da08f42fc9692c82854eb3e4b4ea6f8391b5bb18a3e199cbe9f9e3e6e7852

Observation c0dbc718-15b3-4fef-ba94-993a9eaf9604 · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.187976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.187976Z digest=sha256:8484c4947f7b224e0ff174a16bcfeeb9d24198356b7b7ffedf2e0ff95238f7ba

Observation 83388a1b-427e-42dd-a5f5-0aa425a01d7c · outbound

This paper cites Dual memory units with uncertainty regulation for weakly supervised video anomaly detection.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Dual memory units with uncertainty regulation for weakly supervised video anomaly detection

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.690761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.193178Z digest=sha256:3a1e7c06e770f8e8f2d657a470531d3b152ff5235252e6cdd7877327e788bad3

Observation 96ec4986-1805-49e8-ad9c-1def13a40b0b · outbound

This paper cites Scanreason: Empowering 3d visual grounding with reasoning capabilities.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scanreason: Empowering 3d visual grounding with reasoning capabilities

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.676105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.197976Z digest=sha256:4ee9dfcaf3f472a5e609a8d0364357a16e95b9898a2fb9d60fd606579217ce08

Observation d4465aff-60cc-4572-b653-dfdb711f5ee1 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 3d-vista: Pre-trained transformer for 3d vision and text alignment

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.660133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.202562Z digest=sha256:0c4f93c15eb4be5b7ed26c6017c3da44995bcc47f8145d45ea412cdd2dff131d

Observation 225f2ccb-d3ce-4433-9dbc-f8934a0e1169 · outbound

This paper cites Unifying 3d vision-language understanding via prompt- able queries.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Unifying 3d vision-language understanding via prompt- able queries

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.644142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.206561Z digest=sha256:1852a3d22e96f9964a3fcbcdaa14bd6fdbde8108b627988557979fa9fdf326b2

Observation 17a85cb1-9423-4a72-994e-4fc5374ca733 · outbound

This paper cites TANGO: Training-free Embodied AI Agents for Open-world Tasks.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation TANGO: Training-free Embodied AI Agents for Open-world Tasks

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.210895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.210895Z digest=sha256:31498ec124b6f7ef5774f432450076c48ca779ce9954c66f444a3e778bb8f7ad

Observation be143c56-3127-4de0-84c2-0274233b3c85 · outbound

This paper cites Generalized decoding for pixel, image, and language.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Generalized decoding for pixel, image, and language

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.628398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:02:33.215747Z digest=sha256:252a5369d8acd71ed1178aa7134093e39b57f1bb28140bd40265273b3698db3a

Pith citing papers

Observation 8b12b765-c17c-4532-8932-0c2f5ffe3a7d · inbound

SPG: Style-Prompting Guidance for Style-Specific Content Creation cites this paper.

SPG: Style-Prompting Guidance for Style-Specific Content Creation Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:57.184897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:57.184897Z digest=sha256:6fcfd388202a0cb0717715cb38c2caf2a698f217a344b3b358c3f6e33d48eb89

Observation 7363683b-1a7a-41e1-8f77-a918c1574094 · inbound

FSUNav: A Cerebrum-Cerebellum Architecture for Fast, Safe, and Universal Zero-Shot Goal-Oriented Navigation cites this paper.

FSUNav: A Cerebrum-Cerebellum Architecture for Fast, Safe, and Universal Zero-Shot Goal-Oriented Navigation Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:03:08.600996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T18:58:48.779518Z digest=sha256:7869eb1ee82df095c9e2fdcb0066a4786d3227be31c77b82e211d14a490e8bfe