Pith. sign in

Paper Citation Record · LEDGER

ABot-N1: Toward a General Visual Language Navigation Foundation Model

As of 4 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 2 inbound Pith citation observations for arXiv:2607.10383.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.10383 v3

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T07:19:46.694802Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T03:24:25.215191Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

80 of 80 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved80
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation adb8d8a6-5568-4bf9-b57d-d7983547f11e · outbound

This paper cites 1st Place Solutions for RxR-Habitat Vision-and-Language Navigation Competition (CVPR 2022).

ABot-N1: Toward a General Visual Language Navigation Foundation Model 1st Place Solutions for RxR-Habitat Vision-and-Language Navigation Competition (CVPR 2022)

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:36.808445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:36.808445Z digest=sha256:7f9f0955cb0722fd774de6ad926c95e82c797c01ac33a6a4cb2b3f9fbfeacab3

Observation b6304653-9a51-4724-b28a-20db22fe21ce · outbound

This paper cites Etpnav: Evolving topological planning for vision-language navigation in continuous environments.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Etpnav: Evolving topological planning for vision-language navigation in continuous environments.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:36.880741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:36.880741Z digest=sha256:216b6a96874c64679cb3c9aa4b35da5557d7502e5ee7f95d0aed55276548f4ab

Observation ec2a35d4-3b96-4de9-9bfe-a61f76517456 · outbound

This paper cites Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:36.938706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:36.938706Z digest=sha256:c9e40f8a375d79a54af42e2fd2856a336c18890e689122449108f61f0e0a2813

Observation 53affc0d-cfba-4b9c-9336-f00b74a31862 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:36.990464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:36.990464Z digest=sha256:fccf065667e2a19587d7d5efeb212fc03c0dd2542940ed2b4f2100773da8ea17

Observation 0f34bf93-085e-4bd1-863e-09d364244fbe · outbound

This paper cites ObjectNav Revisited: On Evaluation of Embodied Agents Navigating to Objects.

ABot-N1: Toward a General Visual Language Navigation Foundation Model ObjectNav Revisited: On Evaluation of Embodied Agents Navigating to Objects

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:37.070208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:37.070208Z digest=sha256:f96b8373e9cfb82b34ecbfe4758d690e79874f5306e48a14db91cd89edc73d18

Observation 826ef829-ab27-4aef-b0e6-544bf1774399 · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:37.117575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:37.117575Z digest=sha256:634a80e038560b71dc8c6d286fa6455012b5133b662b2e1b4317527d6dc531a8

Observation 122a4698-8e1a-4fce-917b-11a3a3fad552 · outbound

This paper cites Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:37.181342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:37.181342Z digest=sha256:e98c01bf3cfb7a9ab418023f8ec2af221c48f215bbcd29b408a570743f2554f1

Observation 86df4f93-2068-4319-ac8f-a1a774594377 · outbound

This paper cites Affordances-oriented planning using foundation models for continuous vision-language navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Affordances-oriented planning using foundation models for continuous vision-language navigation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:37.263414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:37.263414Z digest=sha256:6abc926328d55b3400b59a10deb4ff990350eb96a3caa51eb152686c63bc87fd

Observation 5125dade-d346-4bec-bbc3-7e6a973086a1 · outbound

This paper cites AstraNav-World: World Model for Foresight Control and Consistency.

ABot-N1: Toward a General Visual Language Navigation Foundation Model AstraNav-World: World Model for Foresight Control and Consistency

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:37.335878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:37.335878Z digest=sha256:7ab83a4be2b5846956aa4aa035a0499f8752876d5e8e60576d5651fb81c17fc3

Observation c1e8090a-5e9b-4c25-bf3f-76865e3d3d62 · outbound

This paper cites Topological planning with transformers for vision-and-language navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Topological planning with transformers for vision-and-language navigation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:37.400923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:37.400923Z digest=sha256:d6a7da174dc54a7d1c6c6e84b77f73920d92035b605676670cc23be2cbba3b76

Observation ebb3fbdb-5ac6-4751-a227-0b18a806e801 · outbound

This paper cites Weakly- supervised multi-granularity map learning for vision-and-language navigation.Advances in Neural Information Processing Systems, 35:38149–38161, 2022.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Weakly- supervised multi-granularity map learning for vision-and-language navigation.Advances in Neural Information Processing Systems, 35:38149–38161, 2022

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:37.487932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:37.487932Z digest=sha256:8215730e2766313896f60e57436b81d3a438b23dc699671e0e26bb2679d53d41

Observation 9525adab-4ff2-4b9a-9be7-4ae88c9da44b · outbound

This paper cites Explore like humans: Autonomous exploration with online sg-memo construction for embodied agents,.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Explore like humans: Autonomous exploration with online sg-memo construction for embodied agents,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:37.589776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:37.589776Z digest=sha256:6d2dfd3d856424825173b240136371a5178c28c453d032fdf50a5cfd13e006d2

Observation ad62a71f-2a0d-447f-b742-d70398530c9a · outbound

This paper cites Socialnav: Training human-inspired foundation model for socially-aware embodied navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Socialnav: Training human-inspired foundation model for socially-aware embodied navigation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:37.754962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:37.754962Z digest=sha256:3484510ae54952bafe8461bf3dd1b2b77c18f11c4b768267523968baea9f2f48

Observation d4cd4a0c-1303-4220-bb7c-e8301ad32e1f · outbound

This paper cites Socialnav: Training human-inspired foundation model for socially-aware embodied navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Socialnav: Training human-inspired foundation model for socially-aware embodied navigation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:37.827587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:37.827587Z digest=sha256:bd01320891587da1f56d62c521ab760306bf1cfe9c33beb43d025368d1adaac8

Observation 111dc39e-bfdb-495e-839a-0f743d99968b · outbound

This paper cites Spatialrgpt: Grounded spatial reasoning in vision-language models.Advances in Neural Information Processing Systems, 37:135062–135093, 2024.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Spatialrgpt: Grounded spatial reasoning in vision-language models.Advances in Neural Information Processing Systems, 37:135062–135093, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:37.918214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:37.918214Z digest=sha256:d5fcc677cdbed3b902ae3ce88652e694e7335ca7a68c216cba81fe1681b33b4e

Observation 2816f1cb-6124-4612-887f-33fc2ba2146c · outbound

This paper cites Navila: Legged robot vision-language-action model for navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Navila: Legged robot vision-language-action model for navigation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:37.960793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:37.960793Z digest=sha256:45000cfa8d6ce67a07f5943b6e65c7e0c89a17426879315ab5f7d1fe0248b40d

Observation b23e3ce2-1839-463e-808a-8e25af104d0d · outbound

This paper cites Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:38.021123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:38.021123Z digest=sha256:e8ce3bd015e4fd15f275a7cf778ee759ff844463e10fac077ff7fbcea9bd12dd

Observation 4601b2dd-ff75-40ca-bb6f-ae604f0de3c0 · outbound

This paper cites Abot-n0: Technical report on the vla foundation model for versatile embodied navigation, 2026.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Abot-n0: Technical report on the vla foundation model for versatile embodied navigation, 2026

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:38.081301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:38.081301Z digest=sha256:98cf6b4878ce1da3c6c8b391db3d1356e11a3c68ff930ed7760d3bf587922f8c

Observation b3dd50d3-e7b1-43d0-88b9-d4473b4f9f1d · outbound

This paper cites Vpn: Visual prompt navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Vpn: Visual prompt navigation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:38.136227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:38.136227Z digest=sha256:e2a88ce8b996ee32784564c7b0ec897e4a6cb2b1d765a07f7b80a82f661ff28e

Observation 891b3320-bfeb-4dff-ad9b-346ee9c97ece · outbound

This paper cites Helix: A vision-language-action model for generalist humanoid control.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Helix: A vision-language-action model for generalist humanoid control

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:38.216170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:38.216170Z digest=sha256:1ace6fc24df26a68e09c10bfb9ffa41570678d527b382b5e05ae391a525f7a48

Observation 3bbe455c-b2ac-4aee-81e2-8f0e24068f23 · outbound

This paper cites POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:38.321058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:38.321058Z digest=sha256:3792c233435ab330a286047175351ed334d0598cba5ff00e42372491f1b22184

Observation 84dc7f60-b8c4-40fa-9c95-7575481f4859 · outbound

This paper cites Vision-and-language navigation: A survey of tasks, methods, and future directions.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Vision-and-language navigation: A survey of tasks, methods, and future directions

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:38.370727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:38.370727Z digest=sha256:e08e63dad9e0eae42914d449f6276b1568ed40c2a0b259b397c14520ea60a8fe

Observation 83394366-1244-42e5-bc9a-e593ebae063d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

ABot-N1: Toward a General Visual Language Navigation Foundation Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:38.433524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:38.433524Z digest=sha256:ae3d35c8daf95e6fa42aa47f982289979c14fa5336206ef5dc1108d155b9a8dd

Observation faec5ddf-a971-4758-bd72-5d3f25e535ab · outbound

This paper cites A novel vision-based tracking algorithm for a human-following mobile robot.IEEE Transactions on Systems, Man, and Cybernetics: Systems, 47(7):1415–1427, 2016.

ABot-N1: Toward a General Visual Language Navigation Foundation Model A novel vision-based tracking algorithm for a human-following mobile robot.IEEE Transactions on Systems, Man, and Cybernetics: Systems, 47(7):1415–1427, 2016

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:38.441120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:38.441120Z digest=sha256:2acee5ff94afdbd62b3e9cfaeaf3a084cee9b6a004a218284c0a751da39deb41

Observation 1bfa2779-a80d-4126-a669-60a79fa00688 · outbound

This paper cites Bridging the gap between learning in discrete and continuous environments for vision-and-language navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Bridging the gap between learning in discrete and continuous environments for vision-and-language navigation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:38.449475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:38.449475Z digest=sha256:439f84e56a32ded6dd81f3b3dc36288c6b25d20aa7a8f8cc881c0e7dd78c2aab

Observation 74343176-14e1-443d-bf2a-69a4e3d1a491 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

ABot-N1: Toward a General Visual Language Navigation Foundation Model $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:38.498013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:38.498013Z digest=sha256:e9db5c814134d81de79728a22605d721074ab32aad1a7fa490e63a0db2ef2c45

Observation eb18c568-df6a-4dcd-b467-c79204643797 · outbound

This paper cites A comprehensive review of recent advancements in vision-and-language navigation.Discover Computing, 29(1):167, 2026.

ABot-N1: Toward a General Visual Language Navigation Foundation Model A comprehensive review of recent advancements in vision-and-language navigation.Discover Computing, 29(1):167, 2026

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:38.592299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:38.592299Z digest=sha256:e2d74d61799460bae960265f460ecad85d1b3a4d97f67fdfb0cf396ac5b026b4

Observation a10459ff-3da6-4574-b3a3-2d57746b6d7e · outbound

This paper cites Sim-2-sim transfer for vision-and-language navigation in continuous environments.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Sim-2-sim transfer for vision-and-language navigation in continuous environments

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:38.749263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:38.749263Z digest=sha256:b1b617ffce1384f3c9601122a01bde1d130311adbebf9bb725e656badc0a0e54

Observation 92565456-8fbb-4f8b-ba79-4abbee239075 · outbound

This paper cites Beyond the nav-graph: Vision- and-language navigation in continuous environments.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Beyond the nav-graph: Vision- and-language navigation in continuous environments

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:38.893334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:38.893334Z digest=sha256:7b677e1df0e49f0992b31dff03a517955119af3a0ddf580b25087f745525a154

Observation 2e19769c-7526-40be-986e-9b7f4eaa3641 · outbound

This paper cites Waypoint models for instruction-guided navigation in continuous environments.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Waypoint models for instruction-guided navigation in continuous environments

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:39.022007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:39.022007Z digest=sha256:18878a9fb5220ba2664dff53b21207cb49a1fb9771dc8f2bbd91cacf26226f3d

Observation 302cfabe-c7cb-4ee1-a8da-240a88e2ad90 · outbound

This paper cites Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:39.190648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:39.190648Z digest=sha256:d9ee30bbad13a9505f41da3a0891748e087898a590df07a1d56e8c5cb079d32a

Observation 394ab79b-b505-4b3b-8678-4aa890450fdc · outbound

This paper cites Large-scale model-enhanced vision-language navigation: Recent advances, practical applications, and future challenges.Sensors, 26(7):2022, 2026.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Large-scale model-enhanced vision-language navigation: Recent advances, practical applications, and future challenges.Sensors, 26(7):2022, 2026

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:39.270648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:39.270648Z digest=sha256:a5053eeaff842cd2d2c28fa79908ee72d5b67d6c5137a57d513df9dad92ba7b3

Observation ad0bed1f-a59d-45a7-ae5d-5c9f82e49d7e · outbound

This paper cites Navcot: Boosting llm-based vision-and-language navigation via learning disentangled reasoning.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Navcot: Boosting llm-based vision-and-language navigation via learning disentangled reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:39.467980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:39.467980Z digest=sha256:e756062f7bfe96e8486fed05f88b617a51acb3b2a7c28c2b29b03147af95f31f

Observation 156c63b7-23f7-4127-b6c8-34bb149e5a5d · outbound

This paper cites Conflict-averse gradient descent for multi-task learning.Advances in neural information processing systems, 34:18878–18890, 2021.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Conflict-averse gradient descent for multi-task learning.Advances in neural information processing systems, 34:18878–18890, 2021

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:39.574382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:39.574382Z digest=sha256:9e19f89c38c234315385239f788169343ced0cf38e533ca7ce0d6ee0f506ade8

Observation 50477096-bb03-49b9-8cb8-4eaab3decd3a · outbound

This paper cites Navforesee: A unified vision-language world model for hierarchical planning and dual-horizon navigation prediction.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Navforesee: A unified vision-language world model for hierarchical planning and dual-horizon navigation prediction

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:39.691873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:39.691873Z digest=sha256:eb10faeafd2795bcd9299169f3211059375f5975edf04a9b9b6de204f98d4c22

Observation 90a00645-7917-432c-8dd8-0bb8d514a38d · outbound

This paper cites Navforesee: A unified vision-language world model for hierarchical planning and dual-horizon navigation prediction.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Navforesee: A unified vision-language world model for hierarchical planning and dual-horizon navigation prediction

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:39.799646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:39.799646Z digest=sha256:b16758994d2f3222685245800c12bee1b3879861dc1da7d3b7b264ab0ca36733

Observation ad2bda31-0b80-4a07-b4ce-0494b10127f7 · outbound

This paper cites Trackvla++: Unleashing reasoning and memory capabilities in vla models for embodied visual tracking.arXiv preprint arXiv:2510.07134, 2025.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Trackvla++: Unleashing reasoning and memory capabilities in vla models for embodied visual tracking.arXiv preprint arXiv:2510.07134, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:39.883480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:39.883480Z digest=sha256:6785ebef06c75300933572c2651c6b71a8677929e5d60cf430d7ae9f60071418

Observation 2bd60e77-0852-4b8a-8f26-986ca9ea610c · outbound

This paper cites Citywalker: Learning embodied urban navigation from web-scale videos.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Citywalker: Learning embodied urban navigation from web-scale videos

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:39.946903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:39.946903Z digest=sha256:e2fa943539c4c5cb0e5dc409c5a625c1afec6fb19d95b93a9570dc5ca3401914

Observation 5b425577-fa8e-4f88-904d-a268d96ded91 · outbound

This paper cites InstructNav: Zero-shot System for Generic Instruction Navigation in Unexplored Environment.

ABot-N1: Toward a General Visual Language Navigation Foundation Model InstructNav: Zero-shot System for Generic Instruction Navigation in Unexplored Environment

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:40.032137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:40.032137Z digest=sha256:fcf9ed66c9b2ef62792e37dffc93778bb0f1af9a157efae127748b0b3238927c

Observation f1c82efe-76e2-4865-85bc-b0e94dae1bf4 · outbound

This paper cites Pivot: Iterative visual prompting elicits actionable knowledge for vlms.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Pivot: Iterative visual prompting elicits actionable knowledge for vlms

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:40.091709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:40.091709Z digest=sha256:e5c6deb44fa6adc2a5e81aba5cc8d98296a3caff05a95389d1db9d1fdba27e3c

Observation 187b3f55-7677-4926-9bf6-ff4120c9f179 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

ABot-N1: Toward a General Visual Language Navigation Foundation Model GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:40.319034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:40.319034Z digest=sha256:b640aa42fb426b14a1da16673a3d435b21fa16064350b8225697cf857e8ce1fe

Observation 1e0dc0c5-6860-45b1-8596-bf19425790df · outbound

This paper cites OpenFrontier: General Navigation with Visual-Language Grounded Frontiers.

ABot-N1: Toward a General Visual Language Navigation Foundation Model OpenFrontier: General Navigation with Visual-Language Grounded Frontiers

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:40.471855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:40.471855Z digest=sha256:5611be07966e7105d061f942ec64183c81b1f4038fd6fc49b24159d603b5f5c3

Observation d087e6c2-8e6a-45eb-8ee1-879b33f5ea76 · outbound

This paper cites Habitat: A platform for embodied ai research.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Habitat: A platform for embodied ai research

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:40.601537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:40.601537Z digest=sha256:665d55f5b8173e5e21d147bc994d29a39c5db86bac7620e97fb707efb1d5a547

Observation 59ddb180-bbda-45af-b020-75ea7f9aff43 · outbound

This paper cites Gnm: A general navigation model to drive any robot.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Gnm: A general navigation model to drive any robot

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:40.794162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:40.794162Z digest=sha256:6a3c7e3c1a901202bad44741adeca9de9dbf40f23e8837aab0eb319fb0fee38d

Observation 597f69b9-bd76-43ca-b68d-50b2c5ad3eac · outbound

This paper cites Vint: A foundation model for visual navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Vint: A foundation model for visual navigation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:41.001578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:41.001578Z digest=sha256:1d3708738928e45380c498be8b35f90ef60a50e9329a22066b9330eb6b9cf31c

Observation 7cadf22a-03b2-4143-8bf3-e1a1e5b5c8d7 · outbound

This paper cites Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:41.156501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:41.156501Z digest=sha256:747fdb2759141ec92cca0c6592f78175c4d690761ffa8d00394be1427011be53

Observation 5efd68fd-27c8-462e-ab83-5827f5d3a342 · outbound

This paper cites Hume: Introducing System-2 Thinking in Visual-Language-Action Model.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:41.336447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:41.336447Z digest=sha256:e50ac424bebb121128cef3eba43ce812a9870ccd58f472e23e391cc46b873b4e

Observation 91b3af87-24f6-42c2-9a9f-42ec0f4198b5 · outbound

This paper cites Nomad: Goal masked diffusion policies for navigation and exploration.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Nomad: Goal masked diffusion policies for navigation and exploration

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:41.436082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:41.436082Z digest=sha256:6621e3df01af24c4b133c1e0208cc3f776398b0bdc8d4d705be06a417ee4b211

Observation 8a3f2dc4-a7f6-4331-a51e-a42b46d0a661 · outbound

This paper cites Emma-x: An embodied multimodal action model with grounded chain of thought and look-ahead spatial reasoning.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Emma-x: An embodied multimodal action model with grounded chain of thought and look-ahead spatial reasoning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:41.559440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:41.559440Z digest=sha256:2bbbfd985926349db7f83d429c14cb190096946e5f45f65b511821be2b5faa6a

Observation 0ddff957-a98c-47dc-a4ff-74feac8f942d · outbound

This paper cites Robobrain 2.0 technical report.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Robobrain 2.0 technical report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:41.747678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:41.747678Z digest=sha256:68b9d5877034d2da28cd2dd9344da55d3129ce5b47ce0a9a566093d6f8fc2a5d

Observation 146d8fdc-2eca-462f-98c9-4d0503f45d03 · outbound

This paper cites Robobrain 2.5: Depth in sight, time in mind.arXiv preprint arXiv:2601.14352, 2026.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Robobrain 2.5: Depth in sight, time in mind.arXiv preprint arXiv:2601.14352, 2026

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:41.949051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:41.949051Z digest=sha256:fb678eb516d2bac5e9e56367f2db507da69aa56577f78d98f98122c93a21c8c5

Observation c17b63ee-6b9e-4f45-9987-cd14d64924f9 · outbound

This paper cites Drivevlm: The convergence of autonomous driving and large vision-language models.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Drivevlm: The convergence of autonomous driving and large vision-language models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:42.085991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:42.085991Z digest=sha256:e20edeadfb23323ffe95f00fb1436a809da7e0e5ba537e120603bc4827c2c2e4

Observation 66d353e3-7077-4572-b75d-540e94739956 · outbound

This paper cites Dreamwalker: Mental planning for continuous vision-language navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Dreamwalker: Mental planning for continuous vision-language navigation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:42.224826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:42.224826Z digest=sha256:e2606146dbd7e28f0710e7efa0d49c271fcc483d80aa12eb9557ee1d3c46d9d8

Observation 441e92bc-1557-4a7d-a989-e32f160bb4ef · outbound

This paper cites Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:42.400030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:42.400030Z digest=sha256:578c22e52abe7704721ff19122c4aebad2d50487e754830d8c13bd30cffcf239

Observation 7066ddf5-e0c3-4f18-b96d-5b2649a436d7 · outbound

This paper cites Moge: Unlocking accurate monocular geometry estimation for open-domain images with optimal training supervision.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Moge: Unlocking accurate monocular geometry estimation for open-domain images with optimal training supervision

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:42.537690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:42.537690Z digest=sha256:56edd09edcc64cb3f8571d3ca81c6b5d80b70086ad4136a41b7655d07479ccec

Observation d2a5e682-9c0f-4b86-b398-958e6ccf2a58 · outbound

This paper cites TrackVLA: Embodied Visual Tracking in the Wild.

ABot-N1: Toward a General Visual Language Navigation Foundation Model TrackVLA: Embodied Visual Tracking in the Wild

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:42.675991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:42.675991Z digest=sha256:87afa8454c25d8622616d70d012162495d3997d2af69c1bfec07a4c5c33d4662

Observation 1d498e91-58a5-4e3f-abee-5514a3f38ab6 · outbound

This paper cites Gridmm: Grid memory map for vision-and-language navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Gridmm: Grid memory map for vision-and-language navigation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:42.887647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:42.887647Z digest=sha256:fb60894059aacded1de6dd6e9a3d652efc300fdf2fc003293419143a4a06c390

Observation f5bd57f7-9021-4d13-9c35-84914e1681f7 · outbound

This paper cites Lookahead exploration with neural radiance representation for continuous vision-language navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Lookahead exploration with neural radiance representation for continuous vision-language navigation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:43.042916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:43.042916Z digest=sha256:7a188d588a5b7a39a08df159f8eea2765e576f2604fc11e0142fe31cead7f89a

Observation 343b73d3-f947-49f5-af83-c92c35158c7b · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:43.223398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:43.223398Z digest=sha256:7f5ced8b46a6b15621909c5580bf87beab8a32dab22e7ac84f7984f190450b9a

Observation 13e2aa2c-7a73-4bdb-8fe8-13be17d0adbe · outbound

This paper cites Ground slow, move fast: A dual-system foundation model for generalizable vision-and-language navigation.arXiv preprint arXiv:2512.08186, 2025.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Ground slow, move fast: A dual-system foundation model for generalizable vision-and-language navigation.arXiv preprint arXiv:2512.08186, 2025

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:43.356041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:43.356041Z digest=sha256:ca247677b1e853e136fca7b1db889c4d7c215b6501fda04f92bbe57a8a9b793b

Observation 35ace54b-82f0-409c-bcfd-8e99ef9ab06b · outbound

This paper cites StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling.

ABot-N1: Toward a General Visual Language Navigation Foundation Model StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:43.503753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:43.503753Z digest=sha256:bf993c28f5eb164715b7a9ffaea9344fd6bcde812ebf17d06d70960af808b9c7

Observation 4bc63c77-f5ad-48e3-b10e-f20deda7c58c · outbound

This paper cites Nav-r2 dual-relation reasoning for generalizable open-vocabulary object-goal navigation, 2025.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Nav-r2 dual-relation reasoning for generalizable open-vocabulary object-goal navigation, 2025

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:43.638832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:43.638832Z digest=sha256:6f1df8b770edf6d157a842cb6b78dce605667a94e4c6e325a521370ec5b0f5ef

Observation f240ba48-7982-4785-a6eb-3933a3dfd00e · outbound

This paper cites Omninav: A unified framework for prospective exploration and visual-language navigation,.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Omninav: A unified framework for prospective exploration and visual-language navigation,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:43.887148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:43.887148Z digest=sha256:f7cad122202586ff2ac8fd45cfcca9b55d7af651039b00727dafcd6ecf198ded

Observation b3c11b26-4533-4378-870a-ed559ff17874 · outbound

This paper cites an unresolved cited work.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:44.163112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:44.163112Z digest=sha256:142c3780d5e58198ce156a6f643390e15b807d908282d6d7ec3eaa873852acad

Observation f87437b6-e352-4f3e-88cb-186f17837e35 · outbound

This paper cites AsyncShield: A Plug-and-Play Edge Adapter for Asynchronous Cloud-based VLA Navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model AsyncShield: A Plug-and-Play Edge Adapter for Asynchronous Cloud-based VLA Navigation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:44.848057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:44.848057Z digest=sha256:a9914edcd64df146a165750e797e738595a0ecc82778fb37feb3df771bedd9a7

Observation fb7e05a8-f520-459b-ae16-e2e899d97805 · outbound

This paper cites Ce-nav: Flow-guided reinforcement refinement for cross-embodiment local navigation.arXiv preprint arXiv:2509.23203, 2025.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Ce-nav: Flow-guided reinforcement refinement for cross-embodiment local navigation.arXiv preprint arXiv:2509.23203, 2025

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:44.595755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:44.595755Z digest=sha256:0d39dbc2018e43b8ed5dc7ef5e3dabe9ad5049add52084984edd5a462d6ef0ac

Observation c5740e3e-0e1e-4ceb-84d7-b7bf2ee97dee · outbound

This paper cites Gradient surgery for multi-task learning.Advances in neural information processing systems, 33:5824–5836, 2020.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Gradient surgery for multi-task learning.Advances in neural information processing systems, 33:5824–5836, 2020

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:45.185975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:45.185975Z digest=sha256:62f76d0d78884050426a5cfdb67743a9782626fd60a2d994720b8e40047322b2

Observation dd6cb1f0-d359-4913-aeb3-9fb9a95af901 · outbound

This paper cites Hm3d-ovon: A dataset and benchmark for open-vocabulary object goal navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Hm3d-ovon: A dataset and benchmark for open-vocabulary object goal navigation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:45.033252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:45.033252Z digest=sha256:07506c4c5c3c936b300caf6ea0b5fcb8e4613f2159ea2557cfaa7c522eab1f35

Observation 0032cf60-d307-4307-a2a7-0583942824a2 · outbound

This paper cites Robotic control via embodied chain-of-thought reasoning.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Robotic control via embodied chain-of-thought reasoning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:45.353049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:45.353049Z digest=sha256:33c7100c03188302f654a4a3584a3c8930bcff40764327c9f1a29a89ba7e5396

Observation c1ffa229-320e-430d-820c-a04ef96401db · outbound

This paper cites Correctnav: Self-correction flywheel empowers vision-language-action navigation model.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Correctnav: Self-correction flywheel empowers vision-language-action navigation model

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:45.272916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:45.272916Z digest=sha256:b2758066bf48fae344501012e50dabb0872692e3ae48c4d09da912bbd3638f9d

Observation bfdbe320-9b50-442c-89dc-f11389031e3d · outbound

This paper cites Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:45.554365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:45.554365Z digest=sha256:ad3ac31276620d3c58860b12ee58f0116ca53cb115a5526daf1915b6e38069b3

Observation f45fc376-bbdd-4e1b-bf69-de16055f0d5f · outbound

This paper cites PoliFormer: Scaling On-Policy RL with Transformers Results in Masterful Navigators.

ABot-N1: Toward a General Visual Language Navigation Foundation Model PoliFormer: Scaling On-Policy RL with Transformers Results in Masterful Navigators

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:45.435794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:45.435794Z digest=sha256:b950d469afdbbb151d2dbfad88a652a41e2566e2890835239a75aca256bd0c08

Observation 510c865f-b4da-407a-b17f-b9ddbceb467d · outbound

This paper cites Embodied navigation foundation model, 2025.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Embodied navigation foundation model, 2025

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:45.893363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:45.893363Z digest=sha256:32409f6871ba05f78e4b13b7cb240e8567eded2d62dab52650100c13713ec53c

Observation 84b45b46-bdc3-4a0e-bfa2-c00f9278b93b · outbound

This paper cites NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:45.694354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:45.694354Z digest=sha256:17203de6eb9c8109e827e5e69b22c009fd90d9357f6b370c7d7b53bca1310f32

Observation c4a90c1c-ddee-426a-bd76-fac3191b2434 · outbound

This paper cites Cot-vla: Visual chain-of-thought reasoning for vision-language-action models.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Cot-vla: Visual chain-of-thought reasoning for vision-language-action models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:46.111431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:46.111431Z digest=sha256:99051c201565ad103083c40922919b27997bf8cb1db427ecf23e6624b508098e

Observation 59865f02-98db-450f-8038-67d1c7f55cd0 · outbound

This paper cites Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation System.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation System

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:45.977650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:45.977650Z digest=sha256:69af8683fcc4bf27062a6848b384962f16d8b85b07cfd6752d35445dcb68af90

Observation 1b2f08d2-583e-427c-bc98-06e325f44c00 · outbound

This paper cites Navgpt-2: Unleashing navigational reasoning capability for large vision-language models.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Navgpt-2: Unleashing navigational reasoning capability for large vision-language models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:46.494248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:46.494248Z digest=sha256:9e67536ed72b377d86994ad53e9e7718cdc69b157f6984f0f455b1c994c7b705

Observation 3c065a18-e89e-4221-b477-1683adbf7e21 · outbound

This paper cites Empowering embodied visual tracking with visual foundation models and offline rl.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Empowering embodied visual tracking with visual foundation models and offline rl

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:46.276765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:46.276765Z digest=sha256:0fa669483541252b3792915b8a6b05acdc2078e0363f764b264a1c0c407d0013

Observation e57645ce-5f1f-40f3-8174-76c56f8fb83f · outbound

This paper cites Fantasyvln: Unified multimodal chain-of-thought reasoning for vision-and-language navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Fantasyvln: Unified multimodal chain-of-thought reasoning for vision-and-language navigation

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:46.694802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:46.694802Z digest=sha256:88e7142a27556e1b0dc8aa74c6ee0d173dd20ca9bd17ce825f7e431611390270

Observation d5937e12-c0c7-4f15-b192-086bebb1c23e · outbound

This paper cites Explore Like Humans: Autonomous Exploration with Online SG-Memo Construction for Embodied Agents.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Explore Like Humans: Autonomous Exploration with Online SG-Memo Construction for Embodied Agents

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:37.659711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:37.659711Z digest=sha256:de0817b0b7ce52c28d449910368c35392a4ba6c7063792a2d611203c80ba55d7

Pith citing papers

Observation 25261757-7e38-47d7-8624-7e27c07703b7 · inbound

Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation cites this paper.

Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation ABot-N1: Toward a General Visual Language Navigation Foundation Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T00:43:27.294478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:43:27.294478Z digest=sha256:0284d04f6d13345d7455bd255151d21299d78afd553fd4e37dbfa956146a0ee4

Observation 3a8e0699-478c-4ebc-ae1a-20becba05d77 · inbound

Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation cites this paper.

Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation ABot-N1: Toward a General Visual Language Navigation Foundation Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T03:24:25.215191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:24:25.215191Z digest=sha256:28700bb2e49b69d4282117c36c0da234065f41a56c157f02002013c89a0929e2