Pith. sign in

Paper Citation Record · LEDGER

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies

As of 21 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2607.05122.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.05122 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T08:29:01.477958Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 619c6f72-e8cc-4123-a220-525fa1e712da · outbound

This paper cites Contrastive Variational Autoencoder Enhances Salient Features.

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies Contrastive Variational Autoencoder Enhances Salient Features

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-11T08:29:01.477958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:29:01.477958Z digest=sha256:f71de58ab5c20236546ada4324ad37c680f97919415d3baaa57d9822071a0462

Observation 95981706-a1f6-4313-9a39-272a257540ec · outbound

This paper cites GPT-4 Technical Report.

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T08:29:01.477958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:29:01.477958Z digest=sha256:5f6f79a3d2965982395f854287ccfcc57fc10a5c71f4eb8c52175d935fd1180a

Observation b6e50527-e702-4417-aa23-4552559563ad · outbound

This paper cites NaVILA: Legged Robot Vision-Language-Action Model for Navigation.

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-11T08:29:01.477958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:29:01.477958Z digest=sha256:d873853c293d8709dc7cc19391fcf237ed7d99767758a1c83a530b6dd9fb1b17

Observation e010aa7e-79b6-4b1e-9885-a936b56077a3 · outbound

This paper cites Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs.

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T08:29:01.477958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:29:01.477958Z digest=sha256:7f3a9fb3d4d51424aaa79b0c00a969ea54f26fe0d1d2a8dbcb7b7cfb83bd4e0b

Observation 1730f73e-b5d0-4755-aa67-b7726b034c66 · outbound

This paper cites Gpt-3: Its nature, scope, limits, and consequences.Minds and machines, 30(4):681–694, 2020.

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies Gpt-3: Its nature, scope, limits, and consequences.Minds and machines, 30(4):681–694, 2020

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T08:29:01.477958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:29:01.477958Z digest=sha256:e4f33e837df63bcf6d54fd102552d0f6a4c4cd8785dabfddb1a90b48ac7a3c80

Observation a3b959a2-c5a7-4942-b205-63fae6310e2e · outbound

This paper cites GrandTour: A Legged Robotics Dataset in the Wild for Multi-Modal Perception and State Estimation.

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies GrandTour: A Legged Robotics Dataset in the Wild for Multi-Modal Perception and State Estimation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T08:29:01.477958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:29:01.477958Z digest=sha256:f5fd8d4ffb9789e23e1ec49c09acb66b961803afc050bd6977ec5f670860d348

Observation b6ca0b24-073a-4443-bb2f-b6814daef2e6 · outbound

This paper cites CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models.

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T08:29:01.477958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:29:01.477958Z digest=sha256:72a1fd46982963aabd67d5b20c1c9930b53d3f6aa4d63cd022c5c9c03621689b

Observation 8eacd32a-40a9-4412-af01-7699a5566722 · outbound

This paper cites Run-time observation interventions make vision- language-action models more visually robust.

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies Run-time observation interventions make vision- language-action models more visually robust

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T08:29:01.477958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:29:01.477958Z digest=sha256:5fae1fc4db28c3bfed5a12f832660425c6a1c6bb1c0410424267e80b0981e43d

Observation fd7bb761-36fe-486e-9797-a71e6e109af2 · outbound

This paper cites Omnivla: An omni-modal vision- language-action model for robot navigation.arXiv preprint arXiv:2509.19480, 2025.

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies Omnivla: An omni-modal vision- language-action model for robot navigation.arXiv preprint arXiv:2509.19480, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T08:29:01.477958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:29:01.477958Z digest=sha256:f906fcb4939387b865376f0b7ce432ad5575d988b92d9806ff821e4e96f5c1b9

Observation 07f8d2b6-620e-4201-bec6-d11316cae1e4 · outbound

This paper cites Visual language maps for robot navigation.

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies Visual language maps for robot navigation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T08:29:01.477958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:29:01.477958Z digest=sha256:831baa78a633af2f3ea02f599f9b2fe8f9fe14dc87800bb92959cf4d9c90522d

Observation cf0d5938-d2f0-448b-a87f-98b9b5aaa8c3 · outbound

This paper cites Language is not all you need: Aligning perception with language models.Advances in Neural Information Processing Systems, 36:72096–72109, 2023.

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies Language is not all you need: Aligning perception with language models.Advances in Neural Information Processing Systems, 36:72096–72109, 2023

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T08:29:01.477958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:29:01.477958Z digest=sha256:afd7d6b080f6bdd75cfc34c309d8f9c3b97f1044fd57af034cf69e03ca3a02b3

Observation 92b39c66-40fe-4edc-b114-a2a062a06390 · outbound

This paper cites Segment anything.

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies Segment anything

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T08:29:01.477958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:29:01.477958Z digest=sha256:dd2b4c59a3e0b6f20ad6442feaf5161bb233402715d2e75afe59733223d9589c

Observation cd4561a9-4371-447c-8fda-5928d69a8e9c · outbound

This paper cites PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs.

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T08:29:01.477958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:29:01.477958Z digest=sha256:8ca6f32315bfbbff99b98aba4f158ae1656ce545c9766210720aa221a4a34826

Observation 4ccb6bed-6e12-4596-9726-cdc3e94a3ff1 · outbound

This paper cites Learning transferable visual models from natural lan- guage supervision.

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies Learning transferable visual models from natural lan- guage supervision

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-11T08:29:01.477958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:29:01.477958Z digest=sha256:489d09d2d3ed42c5d190ea86ab503d2585f560bbafca3c91a267d8648cd25460

Observation 30b4e02d-41ea-49de-9572-de954ec5420a · outbound

This paper cites Lm- nav: Robotic navigation with large pre-trained models of language, vision, and action.

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies Lm- nav: Robotic navigation with large pre-trained models of language, vision, and action

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T08:29:01.477958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:29:01.477958Z digest=sha256:97be42b428a1ac31d57dcc4e8d070243084c63bfd03d52f3c4b94cad46567332

Observation 0ae1ffc2-1c5a-4431-a8ef-68508a0b0abd · outbound

This paper cites Vl-tgs: Trajectory generation and selection using vision language models in mapless outdoor envi- ronments.IEEE Robotics and Automation Letters, 2025.

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies Vl-tgs: Trajectory generation and selection using vision language models in mapless outdoor envi- ronments.IEEE Robotics and Automation Letters, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-11T08:29:01.477958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:29:01.477958Z digest=sha256:e02a03a3de702b97717b942909848bd0f4e1761b2dab56ce163907283170b6ed

Observation 11562855-92c8-455f-ba19-b9ce3b119783 · outbound

This paper cites A comprehensive review of yolo architectures in computer vision: From yolov1 to yolov8 and yolo-nas.Machine learning and knowledge extraction, 5(4):1680–1716, 2023.

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies A comprehensive review of yolo architectures in computer vision: From yolov1 to yolov8 and yolo-nas.Machine learning and knowledge extraction, 5(4):1680–1716, 2023

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T08:29:01.477958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:29:01.477958Z digest=sha256:ea05fe0f88d742c899149cbb7d5167022acdbee97c286e2b787834defc13e2c3

Observation bec28576-1e4e-49ab-9be9-8a08b6153636 · outbound

This paper cites Segformer: Sim- ple and efficient design for semantic segmentation with transformers.Advances in neural information processing systems, 34:12077–12090, 2021.

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies Segformer: Sim- ple and efficient design for semantic segmentation with transformers.Advances in neural information processing systems, 34:12077–12090, 2021

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-11T08:29:01.477958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:29:01.477958Z digest=sha256:7db167413e09b7e64ec5654cd45d844f0c1773c12bbdd3e66901598e6197cfca

Observation 348aaec1-2dd5-47bd-b44c-df6a9852b279 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T08:29:01.477958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:29:01.477958Z digest=sha256:4fcd01ef1733943fc2fefbfebce56a114e3dfd104cbfb2d20b685982c849bc6d

Observation 247dbdb6-fb27-471a-8678-ad10435ccfcb · outbound

This paper cites Segment everything everywhere all at once.Advances in neural information processing sys- tems, 36:19769–19782, 2023.

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies Segment everything everywhere all at once.Advances in neural information processing sys- tems, 36:19769–19782, 2023

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T08:29:01.477958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:29:01.477958Z digest=sha256:cd6060d1c235dd21d885d0299597d1106688694e16b2fc81622678b4e0630a3f

Pith citing papers

No inbound Pith citation observations are available.