Pith. sign in

Paper Citation Record · LEDGER

EdgeVLA: Efficient Vision-Language-Action Models

As of 16 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 10 inbound Pith citation observations for arXiv:2507.14049.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.14049 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:14:51.731026Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:26:09.241377Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:09:43.333762Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fcc61c4b-b9e6-4380-a4cb-ea1bec758bb6 · outbound

This paper cites an unresolved cited work.

EdgeVLA: Efficient Vision-Language-Action Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:14:52.114122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T16:14:51.643918Z digest=sha256:037c96ce9654fa0e339bbd1f4b5fcce83ba0411c0576a18ae2394912f7a08b29

Observation 7cabeaff-8ea3-401f-bf07-ce88bd2fdb57 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control, 2023.

EdgeVLA: Efficient Vision-Language-Action Models Rt-2: Vision-language-action models transfer web knowledge to robotic control, 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:14:52.093783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T16:14:51.649176Z digest=sha256:4b360415cfebab6041527528b5b288580f08862c526d02f282c825b595f986ad

Observation 4b55040a-e5e3-4fa9-bc2f-ea89ea386e74 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion, 2024.

EdgeVLA: Efficient Vision-Language-Action Models Diffusion policy: Visuomotor policy learning via action diffusion, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:51.653392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:14:51.653392Z digest=sha256:5cd6d879cfef8b19df3056e0651d4e5c7468644775c4cfa6ca89b3c9c0208cd4

Observation b4339069-68af-4cb4-90d1-244af27c6055 · outbound

This paper cites an unresolved cited work.

EdgeVLA: Efficient Vision-Language-Action Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:14:52.062527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T16:14:51.658018Z digest=sha256:6a44dd6073ba8f6c6b5ab73cc159fed5990a9488d48f151bee82d234fadfe73f

Observation 1d2cb8fd-f9c3-4f45-9572-544ca6f370f6 · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning, 2023.

EdgeVLA: Efficient Vision-Language-Action Models Flashattention-2: Faster attention with better parallelism and work partitioning, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:51.663390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:14:51.663390Z digest=sha256:e65d914ccfa26bc5f39c614f572be560fff48743345dd5f239dee1538c517a4d

Observation e58dc51b-1e4d-4ff7-a6ae-16b7d2198660 · outbound

This paper cites Zhao, and Chelsea Finn.

EdgeVLA: Efficient Vision-Language-Action Models Zhao, and Chelsea Finn

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:14:52.030724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T16:14:51.668497Z digest=sha256:eeb2bf6c52567797aa223ec4f7e36c0c47f8d208ef6fa4168e24bc53109fe29b

Observation 017e0517-d1c4-42a9-828b-8ecd93c5d542 · outbound

This paper cites Flexattention: The flexibility of pytorch with the performance of flashattention, 2024.

EdgeVLA: Efficient Vision-Language-Action Models Flexattention: The flexibility of pytorch with the performance of flashattention, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:14:52.006801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T16:14:51.673240Z digest=sha256:91155514a1cc1bef57c5bbf348f1105fa6d742a6e4da79ea017e61b4556b04f5

Observation ba47cdc3-8731-48ee-8c30-bbe99711d169 · outbound

This paper cites Prismatic vlms: Investigating the design space of visually-conditioned language models, 2024.

EdgeVLA: Efficient Vision-Language-Action Models Prismatic vlms: Investigating the design space of visually-conditioned language models, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:14:51.978880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T16:14:51.677745Z digest=sha256:76614d414574f47b990c2254ca4b31b92b3d72d883f87de8ebe9c024338382dd

Observation 7cb397e5-b48b-4eb7-82f4-501a6295911a · outbound

This paper cites Openvla: An open-source vision-language-action model.

EdgeVLA: Efficient Vision-Language-Action Models Openvla: An open-source vision-language-action model

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:14:51.960675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T16:14:51.681757Z digest=sha256:258937a498239f3ccea9780af14a0c954c993f925257a16c84b5724fab1da030

Observation f1bd67e5-551e-4282-8c46-95159e0ca72b · outbound

This paper cites Jin Kim, Nur Muhammad Mahi Shafiullah, and Lerrel Pinto.

EdgeVLA: Efficient Vision-Language-Action Models Jin Kim, Nur Muhammad Mahi Shafiullah, and Lerrel Pinto

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:14:51.943638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T16:14:51.686592Z digest=sha256:9183e0e7c1acab89da22613d417a1135bbd112d085f1f803d2c35a2430168400

Observation 724075ea-6e53-4700-bb1e-5d6264548882 · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection, 2023.

EdgeVLA: Efficient Vision-Language-Action Models Video-llava: Learning united visual representation by alignment before projection, 2023

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:51.691041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:14:51.691041Z digest=sha256:d336251e5ae8e71ec036a5d8f5719dafb23c30549138a7db317efc8c74515228

Observation 5e746a9c-9553-41a7-a8eb-28ae70598b29 · outbound

This paper cites Dinov2: Learning robust visual features without supervision, 2024.

EdgeVLA: Efficient Vision-Language-Action Models Dinov2: Learning robust visual features without supervision, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:14:51.916406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T16:14:51.694992Z digest=sha256:94bb9add73419b72d1fc02264b1e51d95e96d1427c6e185ae7918df9f0246b77

Observation b33b3167-bea9-4fad-9bea-eae5ff169748 · outbound

This paper cites an unresolved cited work.

EdgeVLA: Efficient Vision-Language-Action Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:14:51.895106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T16:14:51.699660Z digest=sha256:6a77ddcd8752b3f9a21532de333c0a97d711f0a14d633f4012c2c363ea8f543e

Observation bb728f67-722d-45f1-a62c-cd881427c286 · outbound

This paper cites an unresolved cited work.

EdgeVLA: Efficient Vision-Language-Action Models Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:14:51.878082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T16:14:51.704415Z digest=sha256:0fd8c52416832f4478e805e4449415badefe7ee89bf1aa5bd5161244fcca4cac

Observation 18e69dac-09a5-407c-b9f2-59e05049e314 · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale, 2024.

EdgeVLA: Efficient Vision-Language-Action Models Bridgedata v2: A dataset for robot learning at scale, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:14:51.861172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T16:14:51.708513Z digest=sha256:9addd1b872e36909ab78fe94b8d0e03a9a13d7e1416ccc3beb3daf9d9f751de4

Observation 26f78f88-05c1-42de-913f-ab7c8125e026 · outbound

This paper cites Bitnet: Scaling 1-bit transformers for large language models, 2023.

EdgeVLA: Efficient Vision-Language-Action Models Bitnet: Scaling 1-bit transformers for large language models, 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:51.712660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:14:51.712660Z digest=sha256:74ac5747978f457449e9c0490b89e56a547b3a1b9eb11c4e1e597212ed86fda9

Observation 259752fe-5b07-4191-99a8-1d85cd21750f · outbound

This paper cites Qwen2 technical report, 2024.

EdgeVLA: Efficient Vision-Language-Action Models Qwen2 technical report, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:51.716782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:14:51.716782Z digest=sha256:ccefe7f7ecf9b5761b037a7d41a72bf43a8808d7d72bc11720ff90f057e08916

Observation b5798ae5-359a-4d01-b392-842f0dd672d9 · outbound

This paper cites Homerobot: Open-vocabulary mobile manipulation, 2024.

EdgeVLA: Efficient Vision-Language-Action Models Homerobot: Open-vocabulary mobile manipulation, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:14:51.814516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T16:14:51.721214Z digest=sha256:1045c07ef7aea6dd9b3aa2e7a479fb7df90929f9fc80805f27ba43bf296eed78

Observation 7d9cf56f-60fe-4b63-a7e9-e3e93e31a559 · outbound

This paper cites Sigmoid loss for language image pre-training, 2023.

EdgeVLA: Efficient Vision-Language-Action Models Sigmoid loss for language image pre-training, 2023

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:51.725819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:14:51.725819Z digest=sha256:88896711ae30576e7d7dcb95d54bafb7ef408be4195d1918fe5d5942bc94a976

Observation 56a71ae0-c734-48bb-9384-25db405ab0a4 · outbound

This paper cites Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn.

EdgeVLA: Efficient Vision-Language-Action Models Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:14:51.774308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T16:14:51.731026Z digest=sha256:25debdba366f2e2990f9e48ffcda040fad42f27f6026371fa0685ba641b0ae66

Pith citing papers

Observation da9aa6c4-60e5-45c7-9ff8-4f003dadc063 · inbound

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning cites this paper.

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning EdgeVLA: Efficient Vision-Language-Action Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:36.552950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:36.552950Z digest=sha256:97c323798c66286f2c492bda84a594fdbd68fb1dfd2a2007b669d86c2aa28e3f

Observation 54c57a68-59f9-4697-a2a0-c616ee2cd437 · inbound

ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models cites this paper.

ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models EdgeVLA: Efficient Vision-Language-Action Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:09:09.496820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-17T06:07:42.311608Z digest=sha256:4871d70eadba83252a322244d3f69f403dd5bfa6adc64cb5d912b4b515440000

Observation f7bbf31f-8fc3-4012-ace2-eedf8b7042b9 · inbound

TIDAL: Temporally Interleaved Diffusion and Action Loop for High-Frequency VLA Control cites this paper.

TIDAL: Temporally Interleaved Diffusion and Action Loop for High-Frequency VLA Control EdgeVLA: Efficient Vision-Language-Action Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T09:07:41.041488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:07:41.041488Z digest=sha256:b4de9f06304e5da5bc82d7982e1ee290283312ea1f24d07bb1f69fda444d7c33

Observation d8f0805e-1214-469c-9032-b9c6bbcbe125 · inbound

KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models cites this paper.

KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models EdgeVLA: Efficient Vision-Language-Action Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:40:14.582101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T18:39:58.095112Z digest=sha256:56e14d5aac00af763ab71b6f5f2ddd8196ba3e448e0b8624b9059300af33f044

Observation 6bf5f007-1226-4aec-bbca-a9fb2570e77d · inbound

HeiSD: Hybrid Speculative Decoding for Embodied Vision-Language-Action Models with Kinematic Awareness cites this paper.

HeiSD: Hybrid Speculative Decoding for Embodied Vision-Language-Action Models with Kinematic Awareness EdgeVLA: Efficient Vision-Language-Action Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:19:54.424179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T09:15:50.963123Z digest=sha256:7b1bb9a8f7de42a44bc340df76dc39cbd862fe72a2357e4bb17d4b4de7029179

Observation 002d6968-f493-4674-985d-6952177d4333 · inbound

FASTER: Rethinking Real-Time Flow VLAs cites this paper.

FASTER: Rethinking Real-Time Flow VLAs EdgeVLA: Efficient Vision-Language-Action Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:05:15.140852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T08:02:13.188363Z digest=sha256:4f026a9c463a688622352c3277b34640b4b3cad3fbca18a49e448d895ddb1b88

Observation 50cac3e0-43bd-4d76-9562-69ca0c71d509 · inbound

FASTER: Rethinking Real-Time Flow VLAs cites this paper.

FASTER: Rethinking Real-Time Flow VLAs EdgeVLA: Efficient Vision-Language-Action Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:50:01.129618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-21T10:48:56.280105Z digest=sha256:e69a03a89e87322588cfcfd7e8173e08c6b9e14eede35c0bfe2b6562ecfdfaa8

Observation 763f57ef-d2b5-480a-bad2-867d8b36c257 · inbound

A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model cites this paper.

A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model EdgeVLA: Efficient Vision-Language-Action Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:50:51.536005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T19:31:23.255452Z digest=sha256:230968e264afd496ef6054ed20b2e37fc647175c158ecd9f6706db8ab7d7c5c7

Observation 9c12160e-3380-4ccd-b26b-127e60ada1c1 · inbound

PolicyTrim: Boosting Intrinsic Policy Efficiency of Vision-Language-Action Models cites this paper.

PolicyTrim: Boosting Intrinsic Policy Efficiency of Vision-Language-Action Models EdgeVLA: Efficient Vision-Language-Action Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:09:43.335102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T10:27:00.283283Z digest=sha256:eca595cd0a6ceb5350cbad00fe6abfbf8fbf5c4eb516b0cb72bf5546ecd2bf19

Observation 38a9c4d0-2bbf-41ba-83b6-b07b40aa5dc0 · inbound

FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving cites this paper.

FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving EdgeVLA: Efficient Vision-Language-Action Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:09.241377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:26:09.241377Z digest=sha256:4e1123830e6bf96939a1862c01942ea1452bc9b583abf939b158b4fd2089b6c7