Pith. sign in

Paper Citation Record · LEDGER

NavBench: Probing Multimodal Large Language Models for Embodied Navigation

As of 23 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 9 inbound Pith citation observations for arXiv:2506.01031.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01031 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:56:34.381892Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:34:35.166153Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:16:27.244688Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact1
  • verified fuzzy35
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 21215b46-a989-48d6-a31a-5c189b1c4039 · outbound

This paper cites Visual instruction tuning.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Visual instruction tuning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:30.989880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:30.989880Z digest=sha256:8587bfb0046007052cc06c157cd26b61b9816946c1fb9813e2bc6228a53f4463

Observation d6dcc617-bf87-405b-8a38-f2ac8f5e3a98 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:31.036599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:31.036599Z digest=sha256:692f4dd62496dc0959805bf4a56978f965d4fe3ba7eba836af29f249a01de214

Observation 7bf7b956-907d-44c0-83bb-03a43bee0878 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:31.156886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:31.156886Z digest=sha256:9df42ffe3e51cbefce3f32759fd11faaad71e7bba82cb725933c64f7df4b52be

Observation c51a4e45-7c00-4f36-9154-1f413d56c8c6 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:31.269579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:31.269579Z digest=sha256:4a9ef5458365239e8dca593ffbfbcdfc99634ca5430f884f58f2d5dbafec84b7

Observation c4c0c367-e3e5-464f-a8e7-002e1df9fd07 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:31.360091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:31.360091Z digest=sha256:f11f3896a43a7a8a69c0932a85676858f5bf9975d20d0a2f29f35e46ff690036

Observation ada2f096-3cec-4aec-97fd-2809ac0113fc · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:31.435423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:31.435423Z digest=sha256:55c6990ec8d8c1f5d6a47987c0d1ccb2e911681436e75257bf906f0a2be451ea

Observation 7b4c1524-e616-4184-8a8f-eaa4f92420c5 · outbound

This paper cites Spatialbot: Precise spatial understanding with vision language models.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Spatialbot: Precise spatial understanding with vision language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.839389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:31.538043Z digest=sha256:e77c3d9cf7fd4b36b7ce6b62504553c35fce5e745b0065908fa77a55113d8e01

Observation a0924245-280f-45f5-899b-f0f7bd9a23f3 · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember and Recall Spaces.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Thinking in Space: How Multimodal Large Language Models See, Remember and Recall Spaces

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.833229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:31.656410Z digest=sha256:ddb55d075ccd7a6fe974bb293855002c45c9e4ed8c30be93e95a12ab355d7ce3

Observation 030c5f51-3ec7-4655-9446-10e0009e8a9b · outbound

This paper cites Reid, Stephen Gould, and Anton van den Hengel.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Reid, Stephen Gould, and Anton van den Hengel

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.826606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:31.740055Z digest=sha256:07556e5c705454d9753d1273110f75bc5e984832175193ef50a3554ebb7bb4e3

Observation 9bd105d3-efdd-45e2-b6d8-8310cd3906f0 · outbound

This paper cites Object goal navigation using goal-oriented semantic exploration.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Object goal navigation using goal-oriented semantic exploration

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:31.853061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:31.853061Z digest=sha256:09116d39abf15677aaeb4d5531c77d33d5c00a639e65295ec37dba2623baf478

Observation 163a60bc-f90e-4919-971f-bc8a1bf2d7fa · outbound

This paper cites The spatial semantic hierarchy.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation The spatial semantic hierarchy

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.815799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:31.993816Z digest=sha256:a1a2a047f9598ccb0ca708519d8badd6f2c0b045ad86fb4e6cc074dc776097a1

Observation 956b1190-8d93-4d6e-b833-e2abbe250619 · outbound

This paper cites Intern vl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Intern vl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.809216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:32.094607Z digest=sha256:1760f9c506f9f8a78ec52133a94b3518290f04ebb4e3d216495f0ca8b29b389c

Observation 43a30cec-1dad-4cb1-9574-8989f357a51c · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Flamingo: a Visual Language Model for Few-Shot Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:32.233009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:32.233009Z digest=sha256:1e20b880d3fa4f4d47945efe4a79ec6e64e18280c7bf12c3e72bb604a6602d21

Observation 1ae2745b-9f2c-4dde-a16d-bd98599fb154 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:32.381639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:32.381639Z digest=sha256:ba2201561a08936b487bc8167c46e62cd3c21283cfdb991ef6d666d85540ad39

Observation 50155778-7ad6-4327-93f5-5d026fd1ce09 · outbound

This paper cites Improved baselines with visual instruction tuning.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Improved baselines with visual instruction tuning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.802388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:32.448570Z digest=sha256:d080fe4a1c145379ef21db8934b4e8971f3845abc5c24437ab0c9de9b656ad04

Observation 4c307aed-a278-4a95-be65-6c1593812cc2 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:32.535038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:32.535038Z digest=sha256:e5fb48cd40ca4905ea8ed981994176ca1394b4ae150ccff22308bc588906ce55

Observation 666f5b79-3766-45a8-a37d-1c3bb08eef14 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:32.613705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:32.613705Z digest=sha256:1b2cfb6188272637427d14467edb16928a74c2fdf99b1e77fa0ae7771d9e3b56

Observation 0e8687be-74f3-4bb6-970e-1e0b67838695 · outbound

This paper cites Towards vqa models that can read.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Towards vqa models that can read

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:32.700436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:32.700436Z digest=sha256:1da7590ec05ba820768a899e6c183c1996c44a17e563a1fd4613c38684dff104

Observation 7a65ad5c-e12b-4aa9-9f5e-735fd19daff0 · outbound

This paper cites A Survey on Multimodal Large Language Models.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation A Survey on Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:32.778551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:32.778551Z digest=sha256:710ac460186d0f1cd5bde6d34404c027400759751258536d38cffaa065dcf3af

Observation 50ae15e1-5e0c-4f5b-9528-576352a5248b · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:32.876652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:32.876652Z digest=sha256:d94c682dc035bf415f276dc2075de652039eb09b57f47239b94cb8925cc3d634

Observation e4bf5719-35c4-4031-9ea6-a8bb600ab2c0 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:32.992111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:32.992111Z digest=sha256:ea55476b91ed784f54fccd127e85bc3a73c7deb34aba6f5a4c0c269ff534a576

Observation a39b7188-b251-4424-80d0-02ee0fbfd917 · outbound

This paper cites Scanreason: Empowering 3d visual grounding with reasoning capabilities.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Scanreason: Empowering 3d visual grounding with reasoning capabilities

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.779899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:33.138520Z digest=sha256:ba1bef3970f46fcf7bc81b8a451fa15728e567c0e968d797bb683cbc4f5ac518

Observation 0ee7bb89-df2f-4513-9e83-3155b4816b88 · outbound

This paper cites REVERIE: remote embodied visual referring expression in real indoor environments.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation REVERIE: remote embodied visual referring expression in real indoor environments

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.710186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:33.218505Z digest=sha256:40ce122ac35f63d472af8ef2408cd21126d75a89b840d5138f4d21870025472e

Observation 9ab1db93-f232-4bd4-aa0f-7d98aaeb8838 · outbound

This paper cites Navigating the Nuances: A Fine-grained Evaluation of Vision-Language Navigation.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Navigating the Nuances: A Fine-grained Evaluation of Vision-Language Navigation

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:56:34.461252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:33.316054Z digest=sha256:70337042d240101886d615c91527b2b58cc1b251902bce6f9ddb0057e44aa628

Observation b7fa6440-7214-4ee8-a7fc-9be78f94c935 · outbound

This paper cites Target-driven visual navigation in indoor scenes using deep reinforcement learning.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Target-driven visual navigation in indoor scenes using deep reinforcement learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.704008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:33.442990Z digest=sha256:d0c736b732657c0643ee7c2d785696cebbba7c06fcbdd347fdd4ff74bc011563

Observation 0a07806e-9fd6-482a-a798-a6ea9258cc60 · outbound

This paper cites Instance-Specific Image Goal Navigation: Training Embodied Agents to Find Object Instances.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Instance-Specific Image Goal Navigation: Training Embodied Agents to Find Object Instances

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:33.633441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:33.633441Z digest=sha256:63495409a87bd1933490e25986cacdc72551580674a2635e834224d08b5eba67

Observation d0a96ec0-53bb-487f-8be0-5d67cb8cc933 · outbound

This paper cites Habitat: A platform for embodied ai research.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Habitat: A platform for embodied ai research

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.697575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:33.786563Z digest=sha256:8530de54f40a1d2002924374c6b49f1cd2e25813cddfc478f780ba5fe9798210

Observation d28e1034-482c-4f62-ac0d-51f97579a780 · outbound

This paper cites History aware multimodal transformer for vision-and-language navigation.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation History aware multimodal transformer for vision-and-language navigation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.691481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:33.926155Z digest=sha256:77e82579517bd3bdccd2c07c9971f651a4b667f14f6b545963abe19f89a628e0

Observation ffb1cc7d-e6cb-40f0-9866-b62c3758316a · outbound

This paper cites Vision-and-language navigation today and tomorrow: A survey in the era of foundation models.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Vision-and-language navigation today and tomorrow: A survey in the era of foundation models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.685086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:33.931292Z digest=sha256:3c7e0e6e53a74852ba72f4421cce03a5e5fc01cbc106a025abd8ab678b211449

Observation ff5f2acb-1254-47a2-8950-80c4aa765228 · outbound

This paper cites Room-across-room: Mul- tilingual vision-and-language navigation with dense spatiotemporal grounding.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Room-across-room: Mul- tilingual vision-and-language navigation with dense spatiotemporal grounding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.678492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:34.016628Z digest=sha256:8e29b8ab4065d6996befbdeed328c6adacdfee73b42c2eb172c0398a7c3d2a74

Observation da9e8b29-1c88-4862-a82f-6446963e2b00 · outbound

This paper cites Vision-and-dialog navigation.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Vision-and-dialog navigation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.671826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:34.129147Z digest=sha256:8f7cf60ed1073d2f3356c98a977b7f0eef8043517a9aa1be639c1c0946fe88e5

Observation 3b45aee4-bb4f-48f3-b4ff-f721f41f20bb · outbound

This paper cites Find what you want: Learning demand-conditioned object attribute space for demand-driven navigation.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Find what you want: Learning demand-conditioned object attribute space for demand-driven navigation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.665891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:34.239516Z digest=sha256:26d8c85c765ac7af4b0ee74de9e6e6aeba785f5da2075bbcd9e98bbb2ddd0fb7

Observation dbd5fa5b-0761-4392-9c2a-2798e5b23685 · outbound

This paper cites Object-and-action aware model for visual language navigation.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Object-and-action aware model for visual language navigation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.659387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:34.302998Z digest=sha256:3c89fb3720fd449624477e6d57dc6851600f2e55e64724a6a12fbac12008e618

Observation 6fdfc8e7-7c36-4f64-b176-5eac9219a851 · outbound

This paper cites Language and visual entity relationship graph for agent navigation.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Language and visual entity relationship graph for agent navigation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.652728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:34.305366Z digest=sha256:3ea47f0d4c016080e2c43756b59db397177edca5605c9cfe8dd1aa5c8cfedb1c

Observation ffe8f20b-1dee-48f8-8931-bdb2673dad8d · outbound

This paper cites Neighbor-view enhanced model for vision and language navigation.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Neighbor-view enhanced model for vision and language navigation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.645702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:34.308363Z digest=sha256:b9b87f9973ccdd13d9621c7990ee1de1465491fc15c712d2cceb15f4cda0c0b5

Observation d6d414e9-6abf-4031-be40-b86b044e4d54 · outbound

This paper cites Towards learning a generic agent for vision-and-language navigation via pre-training.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Towards learning a generic agent for vision-and-language navigation via pre-training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.311272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.311272Z digest=sha256:df316ae1d6f0ca44a9a3da52cc2de3d7b69b7b69d99f7e8610c49a3a6107c06b

Observation eb74ac20-a0b7-459f-90f3-1de2bbe0d4b7 · outbound

This paper cites Airbert: In- domain pretraining for vision-and-language navigation.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Airbert: In- domain pretraining for vision-and-language navigation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.635480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:34.313903Z digest=sha256:38f60154c16a7a4317891fff94ce6f4e7808cafcd77c0d5c7791cb40699291eb

Observation 797d208f-725c-4563-9f38-ad6d6a2eb10a · outbound

This paper cites Improving vision-and-language navigation with image-text pairs from the web.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Improving vision-and-language navigation with image-text pairs from the web

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.628863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:34.316051Z digest=sha256:761169b662cee2906e0b3b99ea30930746b946ac749c3b50f5ad6620c063f3bd

Observation f4007912-7675-4df8-ade3-55b576e94169 · outbound

This paper cites History aware multimodal transformer for vision-and-language navigation.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation History aware multimodal transformer for vision-and-language navigation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.318070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.318070Z digest=sha256:c9e8df4975553e20bd4fdb8f2033366c4fe3476ff415f9c45a41dd5bc648292e

Observation e69bee68-35b1-427e-b8ea-e294a8136d1c · outbound

This paper cites Hop: History-and-order aware pre-training for vision-and-language navigation.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Hop: History-and-order aware pre-training for vision-and-language navigation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.619052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:34.320729Z digest=sha256:cf54cb438f8c8cfbed4194db6e4e9337fefa1230fb8feefad1521908d7490d4f

Observation 0cc9a409-1da3-46a0-b1c3-e5cb4312d527 · outbound

This paper cites Hop+: History-enhanced and order-aware pre-training for vision-and-language navigation.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Hop+: History-enhanced and order-aware pre-training for vision-and-language navigation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.612135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:34.322616Z digest=sha256:6251532911c595e20136a3802431824b71d1bdfdfed77566531a5b4157b34d19

Observation cab423a7-5761-4eff-9ef3-0c661c24d1d2 · outbound

This paper cites Bevbert: Mul- timodal map pre-training for language-guided navigation.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Bevbert: Mul- timodal map pre-training for language-guided navigation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.605762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:34.324822Z digest=sha256:9e791fa846c5b3895cf543a2cf8c84015f2ac0f629eb836cc7f127c9ecbe6e25

Observation ab3890ed-ae3b-4910-b9a6-8a119a827766 · outbound

This paper cites Bird’s-eye-view scene graph for vision-language navigation.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Bird’s-eye-view scene graph for vision-language navigation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.598706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:34.327508Z digest=sha256:2a53cd8f749623f50acbf89af3e63338392fff9ea408210b7ebe526d256ace7a

Observation 869c5aa6-ead2-4a95-b715-14fba6052152 · outbound

This paper cites Gridmm: Grid memory map for vision-and-language navigation.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Gridmm: Grid memory map for vision-and-language navigation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.329612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.329612Z digest=sha256:fb87ecb100ceb15b3e366ff804f5e62b76810e94051de6b6683f3e2d59f74435

Observation b4148ee3-1211-4057-a773-47d20b81f475 · outbound

This paper cites Scaling data generation in vision-and-language navigation.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Scaling data generation in vision-and-language navigation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.588372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:34.332220Z digest=sha256:f36060fa47d6a30d8928d3522263733f0e0334df9247311e390ebfce8a93b593

Observation fbd31e1a-43f8-4c87-8d75-681bf43297a7 · outbound

This paper cites A new path: Scaling vision-and-language navigation with synthetic instructions and imitation learning.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation A new path: Scaling vision-and-language navigation with synthetic instructions and imitation learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.581922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:34.334452Z digest=sha256:cc5fda9b77409959980c80a6a083a620bb58b115fdac0a4509c3a4019df9616c

Observation fd583f3a-7b3f-48da-b60c-cef11520fcf8 · outbound

This paper cites NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.336469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.336469Z digest=sha256:9543de37bd009d038b37a09acef730f1c4187d64fa2aaf5fa27caa9ccc53fc12

Observation 4c5f0d30-1b6d-4dce-952a-b5db2d52d47d · outbound

This paper cites Esc: Exploration with soft commonsense constraints for zero-shot object navigation.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Esc: Exploration with soft commonsense constraints for zero-shot object navigation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.338831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.338831Z digest=sha256:1425ba53ca7342e6990be6c2fbad3a164ff3652dab6014738949c746b4cbaa60

Observation 4c23622b-1486-448a-8ddf-6ff00fd9303f · outbound

This paper cites Cows on pasture: Baselines and benchmarks for language-driven zero-shot object navigation.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Cows on pasture: Baselines and benchmarks for language-driven zero-shot object navigation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.570863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:34.341023Z digest=sha256:780efa25701fa6660f231e80d9f48e3ba75f2aa6692392a3415342fd38dc2f33

Observation 13a87f84-42db-4e4b-924c-f64727b4c4f3 · outbound

This paper cites Vlfm: Vision- language frontier maps for zero-shot semantic navigation.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Vlfm: Vision- language frontier maps for zero-shot semantic navigation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.564674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:34.343415Z digest=sha256:1087dc9380a1d076af9b1134dc82728e598b6afc1a0edef727d9cd8de0b0b366

Observation f24c66b4-811a-44a7-b80f-3dca4d9ed00a · outbound

This paper cites InstructNav: Zero-shot System for Generic Instruction Navigation in Unexplored Environment.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation InstructNav: Zero-shot System for Generic Instruction Navigation in Unexplored Environment

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.345807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.345807Z digest=sha256:c073606bb88fc799783aff8dc997be78d27f602b44e5dce71d98b58d59003295

Observation e0db8a61-4362-467e-a5e9-b777b260f7d7 · outbound

This paper cites Navgpt: Explicit reasoning in vision-and-language navigation with large language models.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Navgpt: Explicit reasoning in vision-and-language navigation with large language models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.558219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:34.348198Z digest=sha256:bb32243b5d11908d84de0ab7dfb1557a740e334af771043b986817f706e4eac1

Observation 73da22ea-1bd2-4d23-a7d2-299ecd103800 · outbound

This paper cites Open-Nav: Exploring Zero-Shot Vision-and-Language Navigation in Continuous Environment with Open-Source LLMs.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Open-Nav: Exploring Zero-Shot Vision-and-Language Navigation in Continuous Environment with Open-Source LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.350222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.350222Z digest=sha256:6575dcbde751c9de5145710d01c917a226f322164a6924568987988539f7bdbc

Observation ad13d19e-961a-4947-9563-355054a51117 · outbound

This paper cites Mapgpt: Map- guided prompting with adaptive path planning for vision-and-language navigation.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Mapgpt: Map- guided prompting with adaptive path planning for vision-and-language navigation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.551288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:34.352853Z digest=sha256:b1152482cd982564f6a1634223e65ac35cabcbcd98fcc65d96aed9d4c8d6b2f1

Observation 6edd141a-afbb-4880-acea-c47a3dc02043 · outbound

This paper cites Chang, Angela Dai, Thomas A.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Chang, Angela Dai, Thomas A

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.544972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:34.355084Z digest=sha256:fbc9f40606b46fcff3cf73575fc8486dd1b257b3b695266698474cb0734a6094

Observation e159f15b-d07c-49cd-8886-5a0df644850e · outbound

This paper cites Grounded entity-landmark adaptive pre-training for vision-and-language navigation.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Grounded entity-landmark adaptive pre-training for vision-and-language navigation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.538432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:34.357892Z digest=sha256:af02bf3a8f628de0ba8263258543d8648074930eef85e429b6069450426e5507

Observation 8849c3ca-7649-4007-87cb-a264cde8c80f · outbound

This paper cites Sub-instruction aware vision-and-language navigation.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Sub-instruction aware vision-and-language navigation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.531915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:34.360525Z digest=sha256:eb9619d0ef1f3596cb4911310a37806e3fc23c1c4ac8e95be6a14e966aa10519

Observation 1a66783c-cad6-41fa-a3f5-806567c558ea · outbound

This paper cites Are large vision language models good game players? In Proceedings of the International Conference on Learning Representations, 2025.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Are large vision language models good game players? In Proceedings of the International Conference on Learning Representations, 2025

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.525823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:34.362980Z digest=sha256:fa5c7da40117e1fecff96ded3d891ac2085efdb4a41330e02fbe1b1fc6821662

Observation 06ac8df7-1398-4aa9-98f8-72ced6995343 · outbound

This paper cites Benchmarking Complex Instruction-Following with Multiple Constraints Composition.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Benchmarking Complex Instruction-Following with Multiple Constraints Composition

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.365773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.365773Z digest=sha256:377ced9858ce04acc3236ed76012e96346524b141ac9a69bb3021b1ab270fd8b

Observation c1723077-1fe6-4428-b60f-2d9391758425 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.368716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.368716Z digest=sha256:1044cb2c755eadfc6430c4c4a754fd405443cc63a91af713d9495eb4478c1529

Observation c024a099-385d-4309-b3a3-2bd1e3852bad · outbound

This paper cites Qwen2.5-VL Technical Report.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Qwen2.5-VL Technical Report

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.371350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.371350Z digest=sha256:73712168f5e2d8f8c0422d60a7e8260fed88c077bd4d4ce36eadd2df94027cb5

Observation 88b45deb-1d22-4ed4-9504-36f2cf5eab46 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation LLaVA-OneVision: Easy Visual Task Transfer

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.374202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.374202Z digest=sha256:8dc018da6f91985a0af4f20d8db1ebcbd957060ff25ef463ef626ca9bdc9ba54

Observation 01e28082-e430-4a01-8720-f98fd4e5c1a1 · outbound

This paper cites The Llama 3 Herd of Models.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation The Llama 3 Herd of Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.377086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.377086Z digest=sha256:9f611e6c0740c6360f8c988ad3b197d1a742e7b7178bb0d3755b363e78fabefe

Observation e7a884cf-2ce2-4223-8b8d-c19b043549b3 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Gonzalez, Hao Zhang, and Ion Stoica

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.379660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.379660Z digest=sha256:883c7253015b484dbf90667062cbcbcdd1b3a853791f46cc3e7de51964d1341e

Observation 71935fba-e03a-4d0a-b192-16795e431d8c · outbound

This paper cites Lmdeploy: A toolkit for compressing, deploying, and serving llm.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Lmdeploy: A toolkit for compressing, deploying, and serving llm

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.513685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:56:34.381892Z digest=sha256:e1dc5ae655a3eadf66c64326fd7eef0af8027837bdf1e362504ab9e9682a694a

Pith citing papers

Observation 5a62e989-4c3d-4ef9-95d9-ea7cb53c7295 · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs NavBench: Probing Multimodal Large Language Models for Embodied Navigation

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:16:34.582778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T20:12:46.385646Z digest=sha256:5604b769076a9d8041daf62231ba4cd2fb39e02cc0ba6779c5f2fd152dde9730

Observation 8f94560f-6d13-458d-9dc2-364341ad59d5 · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs NavBench: Probing Multimodal Large Language Models for Embodied Navigation

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:31:25.399311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T10:30:06.829915Z digest=sha256:4b38cc468f5d2320a1c56bd016398619df9c3888d3736f10c5e5780bf786a036

Observation 9e96874f-a3bc-4662-ac17-b6d6368fe404 · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs NavBench: Probing Multimodal Large Language Models for Embodied Navigation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T22:00:08.208686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:00:08.208686Z digest=sha256:872d8537d17385e292a310673ac93909d332eb031abc528f258c9bc58226565d

Observation 63e27b80-1468-4662-91e6-5979965a309b · inbound

MirrorBench: Evaluating Self-centric Intelligence in MLLMs by Introducing a Mirror cites this paper.

MirrorBench: Evaluating Self-centric Intelligence in MLLMs by Introducing a Mirror NavBench: Probing Multimodal Large Language Models for Embodied Navigation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:55:03.971070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T10:53:48.374637Z digest=sha256:f5eb3c127ede6946921fb48a5d569ab74abcf33aed808a364a9010d4e1e0e506

Observation 309f744d-0ad3-465b-94b4-4d66cfe3cf21 · inbound

Plan in Sandbox, Navigate in Open Worlds: Learning Physics-Grounded Abstracted Experience for Embodied Navigation cites this paper.

Plan in Sandbox, Navigate in Open Worlds: Learning Physics-Grounded Abstracted Experience for Embodied Navigation NavBench: Probing Multimodal Large Language Models for Embodied Navigation

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:11:27.618343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-12T03:36:24.941205Z digest=sha256:449728af2c1c800700b0ed488f2d4da426f99f4eb376070c2615e18014d806cf

Observation 1bcaed83-10ac-41c5-986b-486efaa7ddc1 · inbound

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding cites this paper.

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding NavBench: Probing Multimodal Large Language Models for Embodied Navigation

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:43:38.820539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T18:39:11.904941Z digest=sha256:c7d6292c3fdc4a78cdf9f0159a7ec5bbbca37565861f39f4f255a5ddd8436597

Observation 94f19bab-96e5-4f0a-9326-b856187046dd · inbound

Ask When It Pays: Cost-Aware Open-Ended Interaction for Instance Goal Navigation cites this paper.

Ask When It Pays: Cost-Aware Open-Ended Interaction for Instance Goal Navigation NavBench: Probing Multimodal Large Language Models for Embodied Navigation

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:27.247643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T11:01:08.668612Z digest=sha256:958e31b344e78e0ecf28273830af6c96c5df439b0c6a7c6947a62b1849100c3f

Observation 046acf49-1f28-4c56-b115-bcc16b03c8cd · inbound

NavVerse: Benchmarking Indoor-to-Outdoor Embodied Navigation in Continuous Robot Simulation cites this paper.

NavVerse: Benchmarking Indoor-to-Outdoor Embodied Navigation in Continuous Robot Simulation NavBench: Probing Multimodal Large Language Models for Embodied Navigation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T12:04:42.722556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:04:42.722556Z digest=sha256:a540e1cdc322c2a640418b62f23c62c1be1eebf618ed471b4b293f17931d7ff9

Observation 81cd374f-3024-449e-aadb-4a19f7f1b7ae · inbound

Goal-oriented Navigation Instruction Generation with Tour Video Priors cites this paper.

Goal-oriented Navigation Instruction Generation with Tour Video Priors NavBench: Probing Multimodal Large Language Models for Embodied Navigation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:35.166153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:34:35.166153Z digest=sha256:a96ab719e0084905fcd5fe788107e18c8b2d433d0ad50be77940ab9f98620793