Pith. sign in

Paper Citation Record · LEDGER

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision

As of 22 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 9 inbound Pith citation observations for arXiv:2509.05578.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.05578 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:26:15.704069Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T05:06:33.852576Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:48:39.596474Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a4b40f5c-cdb8-4a1f-95d9-83410d760f0d · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision PaliGemma: A versatile 3B VLM for transfer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:15.585479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:26:15.585479Z digest=sha256:b736efeb5e731af2a9c10f6fc412db8901c4d4dd60a6a6f1737c145f796c81f4

Observation 819333d0-c886-42c1-9d99-6b3e344ae725 · outbound

This paper cites VAD: Vectorized Scene Representation for Efficient Autonomous Driving.

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision VAD: Vectorized Scene Representation for Efficient Autonomous Driving

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:15.602694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:26:15.602694Z digest=sha256:a496619aae247a0a27ecd5e3aa56229f75feef2a56bf96a8d8aadec316206ca9

Observation 73d694dd-ebb0-4090-a1af-96ff63c461aa · outbound

This paper cites OccTransformer: Improving BEVFormer for 3D camera-only occupancy prediction.

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision OccTransformer: Improving BEVFormer for 3D camera-only occupancy prediction

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:15.614268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:26:15.614268Z digest=sha256:a66678cf134cf1d8631cacafdfbbfd4faa5ec87bfaa75b0e32bb606aedd77486

Observation 9d3f5af6-f056-44de-949c-e41152da9c8f · outbound

This paper cites Decoupled Weight Decay Regularization.

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision Decoupled Weight Decay Regularization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:15.619993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:26:15.619993Z digest=sha256:00a247460b0fc1de013756b93560976bd5e011490db340687793f544c26b7e1a

Observation fa17e969-aa3e-46e7-b6e7-506c6ea1da85 · outbound

This paper cites URLhttps: //aclanthology.org/2023.emnlp-demo.13.

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision URLhttps: //aclanthology.org/2023.emnlp-demo.13

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:26:16.210402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:26:15.634695Z digest=sha256:8b18351d6da050cf1b874bcc16d0a96a9b3892b34bd27388ff48eee0b10e9eaf

Observation 25d3b5ef-29ff-4d55-80cb-aa3b36177bc3 · outbound

This paper cites PaliGemma 2: A Family of Versatile VLMs for Transfer.

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision PaliGemma 2: A Family of Versatile VLMs for Transfer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:15.649617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:26:15.649617Z digest=sha256:f56569b104ecd299de142a2eb2436fa171de5dcd43f47b7e80f5665a3a542fc9

Observation 082a111f-0e5b-48a0-b8b0-563b5fea803a · outbound

This paper cites DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models.

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:15.654582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:26:15.654582Z digest=sha256:0fb415b82fbaf9c288a20b239213890977bdbacfb3c15fff6d78df8aae7305f4

Observation 6e7633ad-96e6-4c6f-afa0-43398143f7d1 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision LLaMA: Open and Efficient Foundation Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:15.659943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:26:15.659943Z digest=sha256:65834d475a34166da31fb1d1ffb5aeb38ae57b23c90a7e1fc0958b410c4fde7d

Observation 9529e215-2148-429d-b64d-a27a92768539 · outbound

This paper cites 12 Preprint.

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision 12 Preprint

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:15.664533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:26:15.664533Z digest=sha256:c123a4b25800adf0f941694b959553b351b68bef8d2a42c2870b3cece2256e4f

Observation 9c5496ad-3f5b-4efd-bd2b-2e6cf53373d4 · outbound

This paper cites OccLLaMA: An Occupancy-Language-Action Generative World Model for Autonomous Driving.

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision OccLLaMA: An Occupancy-Language-Action Generative World Model for Autonomous Driving

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:15.669024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:26:15.669024Z digest=sha256:66bb85da0019e82cb8138c396b1835b90e522abbf61045a79719cc9f3db7cc5b

Observation a4577109-337f-4ca3-a1f6-62ee9d179133 · outbound

This paper cites Neural Map Prior for Autonomous Driving.

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision Neural Map Prior for Autonomous Driving

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T16:26:15.903844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:26:15.673775Z digest=sha256:ce2c9dd39bb29ce8f65060c2d569525513650adc91abf3f16d060e17683e6b94

Observation 88242e4b-b939-4d61-8a3c-37e140382b10 · outbound

This paper cites DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model.

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:15.678713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:26:15.678713Z digest=sha256:a59614df02cd55fc64109387cc1e120b7b259581e0b4629e4df2f33698beafcb

Observation 60ce2586-7d5f-403a-87c1-2c073a800c28 · outbound

This paper cites LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding.

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:15.684056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:26:15.684056Z digest=sha256:149bd0260bf8d98d96e85533f6d23a71356320d1bbae09b0f2126a87ba9c0189

Observation 6b063f9d-131a-4842-b71c-41bda42fa56b · outbound

This paper cites Rethinking the Open-Loop Evaluation of End-to-End Autonomous Driving in nuScenes.

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision Rethinking the Open-Loop Evaluation of End-to-End Autonomous Driving in nuScenes

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:15.689086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:26:15.689086Z digest=sha256:0576cff87272f6553b93511430d7c7686a26b1c65b5ea5e908e6aedea2e4326f

Observation 9a38f73e-77f9-459d-b1a7-197c79fd9cbf · outbound

This paper cites OccFormer: Dual-path Transformer for Vision-based 3D Semantic Occupancy Prediction.

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision OccFormer: Dual-path Transformer for Vision-based 3D Semantic Occupancy Prediction

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:15.693715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:26:15.693715Z digest=sha256:d7811717f6284a8839525fdb5498514ce1f54d3e73d136f415b9beb87eb4c222

Observation eb094116-c308-4135-9b69-edf76732ea58 · outbound

This paper cites DriveAgent-R1: Advancing VLM-based Autonomous Driving with Active Perception and Hybrid Thinking.

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision DriveAgent-R1: Advancing VLM-based Autonomous Driving with Active Perception and Hybrid Thinking

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:15.699566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:26:15.699566Z digest=sha256:33afb6b2b0dd09d2f973ed93c8876c6322f8f3fae5f74c763c5ec0cd10d13a8f

Observation f5f2124b-ce8d-45d3-99ee-fc272b619a11 · outbound

This paper cites an unresolved cited work.

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:15.704069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:26:15.704069Z digest=sha256:69cfb904f1fbe5fe03d8ab021266c70d64427d2aefef40bc2e159cdb61172e5a

Observation 18b6d784-41e2-4682-ad8a-83096becfd5e · outbound

This paper cites NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario.

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:15.639601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:26:15.639601Z digest=sha256:f77f3a7ada0703a65ad8b5575b9d7e9d53425191cf62f7c83dda538077a238ca

Observation e9d7a048-9be0-4e12-a0d1-5dcbd3522690 · outbound

This paper cites Bridging Stereo Geometry and BEV Representation with Reliable Mutual Interaction for Semantic Scene Completion.

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision Bridging Stereo Geometry and BEV Representation with Reliable Mutual Interaction for Semantic Scene Completion

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:15.607853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:26:15.607853Z digest=sha256:d9468ad37f76191685f2f4f35184e81b55de751ba2d20abebf3568767fcb4841

Observation 805e8997-b38a-4cd5-b1b7-7b2e99b2ea31 · outbound

This paper cites Adapters: A unified li- brary for parameter-efficient and modular transfer learning.

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision Adapters: A unified li- brary for parameter-efficient and modular transfer learning

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:26:16.226552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:26:15.629681Z digest=sha256:7d31907923d7224734e5023b6b654eda64a530cbd744ab7dbabe5336cc08ded9

Observation 089e9a7c-596e-4ced-a9e9-c6c33fdb4d75 · outbound

This paper cites ST-P3: End-to-end Vision-based Autonomous Driving via Spatial-Temporal Feature Learning.

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision ST-P3: End-to-end Vision-based Autonomous Driving via Spatial-Temporal Feature Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:15.591617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:26:15.591617Z digest=sha256:90ee0c163b3923afff0572aad610cf9ab13916249bee7743455d1918a3ff4d0c

Observation eea1e51a-4e1e-435e-b1f1-86b1b4a37f72 · outbound

This paper cites Tri-Perspective View for Vision-Based 3D Semantic Occupancy Prediction.

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision Tri-Perspective View for Vision-Based 3D Semantic Occupancy Prediction

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:15.597320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:26:15.597320Z digest=sha256:6c1bfba6470ade0e78de37c0e9e0009da411832bf8622402c6800a7b8ba44eee

Observation a37b57bd-8291-4f65-96e3-5f1846e09a0d · outbound

This paper cites GPT-4 Technical Report.

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision GPT-4 Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:15.625080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:26:15.625080Z digest=sha256:a24cf7a7eb0723ffbe11170500d0bc6bbde62348163a54c409b1bd8d452168c1

Observation fc05f3a1-bb1c-445f-8eda-3a47119e9884 · outbound

This paper cites LMFusion: Adapting Pretrained Language Models for Multimodal Generation.

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:15.644541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:26:15.644541Z digest=sha256:93f7f1ddbd450d5f7fc3532f06d3f734f558d093709bad91c5216b6bd7196c7b

Pith citing papers

Observation 5f1683f6-57be-4be2-9d80-8e61963892de · inbound

ShelfGaussian: Shelf-Supervised Open-Vocabulary Gaussian-based 3D Scene Understanding cites this paper.

ShelfGaussian: Shelf-Supervised Open-Vocabulary Gaussian-based 3D Scene Understanding OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:28:58.098776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T03:24:32.577420Z digest=sha256:aba79b1268d9f49ffad5f02f19f483a26546aad7e00bca25d417f198ea9f3cf8

Observation 9f063d58-4fb3-4dd6-8835-ef025ac71ba9 · inbound

Monocular Open Vocabulary Occupancy Prediction for Indoor Scenes cites this paper.

Monocular Open Vocabulary Occupancy Prediction for Indoor Scenes OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:54:09.204266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T11:50:37.423645Z digest=sha256:1dd6375a5c6168f09e05c481d5df10fcbb24ff3cd29351327083a4073c086348

Observation 5ffcdbe2-573d-4277-9544-3875d8f2d94a · inbound

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook cites this paper.

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision

Reference 124

Resolution
unresolved
no resolver link, observed 2026-07-13T14:03:01.974171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:03:01.974171Z digest=sha256:b0deca941b9d0ea9c4549b82a3be0dc7af779cdc1ff8be524fb82c4e051c6f41

Observation 129cd057-4381-4156-be54-a34a21c1153c · inbound

Height-Guided Projection Reparameterization for Camera-LiDAR Occupancy cites this paper.

Height-Guided Projection Reparameterization for Camera-LiDAR Occupancy OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:36:06.388360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T17:24:36.149795Z digest=sha256:4041e44406191e364b079be3fc27c86b73d9cf118c1164f3365d43dc0bc0b75a

Observation 698acbad-fd1c-4fd1-8a42-201e9edaf122 · inbound

Height-Guided Projection Reparameterization for Camera-LiDAR Occupancy cites this paper.

Height-Guided Projection Reparameterization for Camera-LiDAR Occupancy OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:41:43.960049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T04:02:19.480937Z digest=sha256:f9ffdffc3bb7b7c9ef2c9efb2f0bb1f40fbadee9b6dfd8fb5f6c9167cbccdec8

Observation 207aa84a-c5bf-4bc7-96ca-eaf65bf772c0 · inbound

TPS-Drive: Task-Guided Representation Purification for VLM-based Autonomous Driving cites this paper.

TPS-Drive: Task-Guided Representation Purification for VLM-based Autonomous Driving OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:43:45.704784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:42:49.902997Z digest=sha256:e190bb8fd7475d095fc0bc6ee99e70210a3e7c9f8b9c703d5f55cd4c8fd041d2

Observation 673c8019-968e-4da9-89d5-713c18b52b3e · inbound

M$^3$Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks cites this paper.

M$^3$Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:16:47.758810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T06:16:07.090870Z digest=sha256:8a6f1ea2f5b1562a1dbe2f1ff95b1d0126e73196aadfd927e35b580a2ac956ae

Observation 1a5a9899-08ed-4210-ba36-436a01495968 · inbound

Teaching Vision-Language-Action Models What to See and Where to Look cites this paper.

Teaching Vision-Language-Action Models What to See and Where to Look OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:48:39.597762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-03T16:42:13.520913Z digest=sha256:2f8692ef769e7d0f221315151268090146d4d2d48098b3a1b039711ac340f6a7

Observation 8d12a04d-ee7a-4419-9c93-78de1d3c9166 · inbound

GPOcc++: Unified Sparse Gaussian Occupancy Prediction with Visual Geometry Priors cites this paper.

GPOcc++: Unified Sparse Gaussian Occupancy Prediction with Visual Geometry Priors OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T05:06:33.852576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:06:33.852576Z digest=sha256:5852aed031305d75889846fcc844770a8b4a8c09d1a2b71dbc022ab00b3d73f2