Pith. sign in

Paper Citation Record · LEDGER

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models

As of 7 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 1 inbound Pith citation observation for arXiv:2604.19728.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.19728 v1

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T02:10:04.003151Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T07:59:34.398439Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

78 of 78 outbound references displayed

  • verified exact36
  • verified fuzzy32
  • unresolved6
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ea607541-9413-499d-85b3-8815c89bfc14 · outbound

This paper cites an unresolved cited work.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Unresolved cited work

Reference 1

Resolution
verified exact
doi, observed 2026-05-10T02:11:57.131161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:5a5681ab59362e0c7d8575b63198705a0756f872d87b3cb73f5270113937c3d5

Observation 0fbe3b42-b226-4d74-abe7-042978b5e16d · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:52:01.553388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:8c8bca1e5ca2e4fd5e60a0d9115f024b12edebeab1bb35e3ea1abcc0cee59610

Observation 28f079a4-8f65-475c-8533-f39ce61f595a · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:16:03.855195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:7814fef4041fd6a5735029feb062c5d5293f00119c46c9f6a1cc8c2166c9acc6

Observation 9dd1efbf-0e68-4447-9f05-4bb2b5418fb3 · outbound

This paper cites Qwen3-VL Technical Report.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Qwen3-VL Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:16:03.996550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:e006f30e4e38f3e627e5caab5c055c8774e61d0534aae0fc71e640be79201594

Observation a9e21c89-a13c-4a77-b12d-b316867609e5 · outbound

This paper cites Foundation Models for Robotics: Vision-Language-Action (VLA).

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Foundation Models for Robotics: Vision-Language-Action (VLA)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.332957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:794e008a12bc65cf93a697ab9aac7387eec9096861f1d55befd247ff61db43ba

Observation 76213249-7ab4-4db6-af80-3ddd55fe4b7c · outbound

This paper cites Significance Tests for 2×2 Tables.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Significance Tests for 2×2 Tables

Reference 6

Resolution
verified exact
doi, observed 2026-05-10T02:11:57.127105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:132fb2e9401468e94344a0006f8061a1c94bc0cd9a2ce8f285d9adf6e41b0f57

Observation d821026d-71c4-4cde-812f-b6ba517dae05 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models PaliGemma: A versatile 3B VLM for transfer

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:16:04.091270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:6200507e509a72d7d21647831eec802563454284d951e9735d5f817d1fd3e4f5

Observation 4c3989b4-c1a5-483f-8861-3037353f988d · outbound

This paper cites PIQA: Reasoning about Physical Commonsense in Natural Language.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models PIQA: Reasoning about Physical Commonsense in Natural Language

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.282531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:7612b228f2cd18c264582b600c95dd9ee21551fa2af8c9b358ca09f798113d17

Observation 9c6622af-bcd9-4cba-8892-a952e108816a · outbound

This paper cites 2020.url:https://github.com/webdataset/webdataset.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models 2020.url:https://github.com/webdataset/webdataset

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.272116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:5bcd62e0a5ecca7650b833f378bce593e0879ea1df22d504f7544a1cbfcd0431

Observation 13acab17-d69e-4a21-afef-c53e52d157cd · outbound

This paper cites https://github.com/huggingface/lerobot.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models https://github.com/huggingface/lerobot

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.305456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:5043958f75329d1003b4dc78c077376554c58717befd735c132f68e6545480da

Observation 0d3419f0-c575-48a1-bb60-db89462d6476 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.742717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:f3413b33eb8ff069a9e1901467c35ff7b666bce7926b5ef18e14a76a0d211830

Observation 3f477b18-1d24-451d-b1f6-d46f466ef165 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.335743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:b76740f5e54940534357242e7840a997b3c994dbe7226333c76c371cc40e9912

Observation ec706d37-b11e-48cc-92b1-ae883020a880 · outbound

This paper cites Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.280063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:d1f4f4ce189215b0a266098f2cdb40eb2ff6bfa50efcb7f4b8897de4cdbd3789

Observation ccd39185-3a9f-49b6-b2f0-6de84b2f7f42 · outbound

This paper cites B ool Q : Exploring the Surprising Difficulty of Natural Yes/No Questions.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models B ool Q : Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 14

Resolution
verified exact
doi, observed 2026-05-10T02:11:57.133216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:64ca2de0355df351d8c45593bad8f52a235e2921056087a8397e0910572591ac

Observation bc4ecee4-8a3f-413d-9c3f-2ddd8dfa9420 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:16:03.848833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:57550faa0f1f6c00e826a7d31b144d9f66a6310a9611ef2c41edb9be45b3d6f3

Observation 3675c2c8-2123-4777-828e-b04f2c3feb5a · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:23:25.503833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:d684fd66e800bdfaa9644696c2a07681f6a391bd1635c3bfc6f27681034a59e6

Observation 4362aebc-e716-42b8-bb85-377c90947597 · outbound

This paper cites StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:16:03.908151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:2d86d99d6fa42c8efdfdd002b02df737b106afba314359b2e51cb0f4b811aedd

Observation 5c6753e4-3d10-4b80-87d7-53a16454e1bb · outbound

This paper cites an unresolved cited work.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-23T00:32:19.284935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:3ba7b2a4e17b9f77d9e0d6555a1e567b11542a16911e9450c48caca0300b6c64

Observation 870cc89e-5d11-445b-8d3f-a67e005b1af5 · outbound

This paper cites an unresolved cited work.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-23T00:32:19.239770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:6ad3624d78e75e980549a7bf57a350a86c14510ae066f46d292ad932e9f897b9

Observation ce099ebb-fd41-459c-8e46-c331399344e3 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:16:03.866165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:aaa61a736e072748d6794cf5118afe801f633b70b35d744cbbca23c7b5f543b0

Observation 2bf5c679-0649-43de-bb65-cb6034c54482 · outbound

This paper cites Computing Extremely Accurate Quantiles Using t-Digests.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Computing Extremely Accurate Quantiles Using t-Digests

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:16:04.062403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:5877ebf6eda59458cf03b46a64147ef6218d2ddbeaaf0a63d3b77bae11f66e93

Observation 991e7bda-d5d4-44f6-bc59-db48965075ed · outbound

This paper cites https://github.com/EGalahad/vla-scratch.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models https://github.com/EGalahad/vla-scratch

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.287651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:07c9cf8405a9162294aab4fdaf9ac8c537db69b2cc7a449267689f113d13b894

Observation aa642ddd-5521-465e-9a6c-762f10cff35f · outbound

This paper cites Datacomp: In search of the next generation of multimodal datasets.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Datacomp: In search of the next generation of multimodal datasets

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.300111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:5e80bcdd2aa929e33223d587d0277bafd778816a11c4660cbf10898d93035586

Observation 7f4e4d39-b429-479e-bfcb-1fec9291fbbc · outbound

This paper cites GitHub repository.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models GitHub repository

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.317503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:1aa23bff0462fc4099cf15811ddacd052d0064504efd1a80df872d722177fa0e

Observation b2a17e78-dd03-4e7c-bd03-79ef045c28d0 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Measuring Massive Multitask Language Understanding

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.266753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:7e5ffc230be479147cf45de2e108c3f667e2a3ccf168077a369f5e139fe6a675

Observation 267f65ac-1b43-4d6f-8b8d-68edf4aacab4 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:00:53.610624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:e268ac5fece7ddf8fdce920a42bb2c9ecfe0540231223a4c26163cea32cbafcd

Observation 6652028f-4e83-400d-befa-54fc04374327 · outbound

This paper cites Scaling Laws for Neural Language Models.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Scaling Laws for Neural Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:16:03.871739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:b4994634278203f306e3373728e8590bbc84cc911350c7f9db7dab2946529820

Observation 3123b62c-1a60-42c4-a81f-7441fc476163 · outbound

This paper cites Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.269499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:39778094d20dc72dba75acaf14592c0c6cbcae3ca2f611a923e31946899a4cc7

Observation 3d401fb1-0db9-4d33-a6c0-3e2d3cd6482d · outbound

This paper cites 2022.url:https://github.com/karpathy/nanoGPT.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models 2022.url:https://github.com/karpathy/nanoGPT

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.296755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:cec5c06c9dfcd0a25e7ce4c139dda09fd8b3b01154711be35569158d92cb1a7b

Observation ab335f5b-d315-4d77-bcdd-89df223d3f85 · outbound

This paper cites Should VLMs be Pre-trained with Image Data?.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Should VLMs be Pre-trained with Image Data?

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:16:03.883147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:f86e57024343d4db94def3bd04117665b0e40dc1f627f5ae62066fd79d5c85d6

Observation 7f2f86cc-aed9-4c2c-b487-3e1348db3107 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:16:03.964134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:ef7064cc93294df4a20fe906ecaef98291c2d39e20da221510339d01e29c1ae6

Observation df9aba8f-78a7-452d-8c4e-900914c589b9 · outbound

This paper cites Accessed: 2026-04-17.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Accessed: 2026-04-17

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.324638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:3b3e88246ea3549289444824a37464994f37cd033aaaa866c80724c8313643a4

Observation c0b14bd0-dd95-485b-b510-dcd6c1042149 · outbound

This paper cites MolmoAct: Action Reasoning Models that can Reason in Space.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models MolmoAct: Action Reasoning Models that can Reason in Space

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:35:23.409560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:999897802a2b0de1f111bfa37257cb5980781057b7c51b96bec9247173c2766a

Observation 0d7ea2ce-a2b3-452d-931f-b5da856b1830 · outbound

This paper cites Datacomp-lm: In search of the next generation of training sets for language models.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Datacomp-lm: In search of the next generation of training sets for language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.260742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:4afffc5d3a36efc3ee462dc82a3caa1f92e6d3371abe03d95be861dd3e9ac8a9

Observation d6e513e0-612d-4c25-ad64-7f4ceb783114 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:10:50.405848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:826f52293250b8eec86486ab2340c764390ba7b6b4a7178d512d70b8f03cb139

Observation ad49dc81-b19e-4516-93cc-26705d8433dd · outbound

This paper cites A systematic study of data modalities and strategies for co-training large behavior models for robot manipulation.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models A systematic study of data modalities and strategies for co-training large behavior models for robot manipulation

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:16:04.025910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:3ea91d236707561ab14b6ade281b4ed125f2a28fcad417ce2f4e571f47530039

Observation 409854ad-10ee-4259-bf1f-dfb88c4186c0 · outbound

This paper cites HoloBrain-0 technical report.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models HoloBrain-0 technical report

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:16:04.036982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:572e5eaeba8336312fbc802c8f45b76de6d63d960ea17b19be36199f370ae3c3

Observation 33dc2a24-1749-458d-9c8a-785a535f3fa7 · outbound

This paper cites Flow Matching for Generative Modeling.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Flow Matching for Generative Modeling

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:16:03.815742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:a3ea185d82372027b5806de197cbe50e28b30f5b8a39033c233cbee5c38e4203

Observation c17ec488-2d1c-4c80-857d-e6565dbb98ea · outbound

This paper cites Visual instruction tuning.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Visual instruction tuning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.290409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:1d93fed16dcdc9d0725b857ebc6347bbb77ac5fd5be89bc112ee94f80999cc56

Observation 3fceb114-49db-44b3-8a73-7bad8b49d944 · outbound

This paper cites LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:11:57.125399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:0c3f6884b5114489885d4a2524e4e0d5fecf53802c127076714b77143291ab86

Observation b2dd82f8-3a94-4a13-b64b-b5b948ca86d8 · outbound

This paper cites LLM360: Towards Fully Transparent Open-Source LLMs.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models LLM360: Towards Fully Transparent Open-Source LLMs

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T02:11:57.135983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:b004504610fe513f5b8db4942308a518844486a59dc295174f19432f18415901

Observation d23fceb5-745b-4ea1-9b2d-99a0b26eac8e · outbound

This paper cites SmolVLM: Redefining small and efficient multimodal models.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models SmolVLM: Redefining small and efficient multimodal models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:23:51.804797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:1b906a42a03fd67e5e3ac4fa6728a1039a2691a9f96f884874747bd85d2d37d9

Observation 3eba369d-038c-4eb3-804d-ba2930fae3e9 · outbound

This paper cites an unresolved cited work.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-05-23T00:32:19.293296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:741fefdde11ca74663bba64e66001a0f01c24ccb28178571ed526286b95f118f

Observation 3d892be4-552c-49dd-bb9c-f5db7b37fc97 · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 44

Resolution
verified exact
doi, observed 2026-05-10T02:11:57.140586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:b4dc0fef79f86a9919ed773773ccfe5ca4268a7cff1ea3c9a74dcd5eee0db191

Observation bd462d5c-6b26-443e-9765-faaea284d324 · outbound

This paper cites Ray: A Distributed Framework for Emerging AI Applications.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Ray: A Distributed Framework for Emerging AI Applications

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.302795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:40d858e9e258779b71ef5c1887993243c901f5f4ba05197c9bfd471286c9b8c5

Observation 8f295ee3-0460-41eb-8701-2c3d6cc9184c · outbound

This paper cites an unresolved cited work.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-05-23T00:32:19.309179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:d4bce97255f11f45c0b1a004794908f0d81f87d1d920878b330d4bd3feacaa39

Observation 1a182c20-f464-43ed-b91c-cbec9897249a · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:16:03.981588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:67974cfed040b0daee3ac1e0269a2ddb1414f12cc80fa8191aaea4b66f3cf491

Observation 0590d8bb-0790-4ca6-805c-7a099cf5eda8 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models DINOv2: Learning Robust Visual Features without Supervision

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:16:03.926360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:2da0a407682e8d4ef5118bb29350ab94247a2b7f12df4afc10a697f0b9ea352f

Observation 43ecea01-a308-4e09-8273-83e0d2222555 · outbound

This paper cites The fineweb datasets: Decanting the web for the finest text data at scale.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models The fineweb datasets: Decanting the web for the finest text data at scale

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.327451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:77b9caedda21e9c193d7ce38f6fc361165e5739399b4a15d0613aeeb6231ac0d

Observation 563b84a0-e4f1-47ed-abb9-dd1190b25286 · outbound

This paper cites https://github.com/nepfaff/drake-blender-tools.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models https://github.com/nepfaff/drake-blender-tools

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.330158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:9149beba306dd9079d50d24ef70719714a14cb69715d45528bc3a032b818d1ab

Observation 5ea3b885-cb35-48c8-a5ec-0c2d0e4f3d9c · outbound

This paper cites GitHub repository.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models GitHub repository

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.258242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:0b55615528f29ff6080ea7a779fef3189f4071fb8f22e3242a719b2a35a726fb

Observation a3a97002-b69d-44b5-b1a7-472819870261 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:16:03.975504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:d532d667cca838504eac39fcb0c29fc0a390577e257c213a2ac030a38bdc7f53

Observation 79dc5373-be84-473e-ba4a-68b154ed655c · outbound

This paper cites $\pi^{*}_{0.6}$: a VLA That Learns From Experience.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:16:04.018345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:584f6e6f54722551f8b5f24d2d947b1ffa324e0f1e8d50514b8b7343f6256f88

Observation 34e953d2-8b38-4dd4-8ff2-33339b620bd0 · outbound

This paper cites An algorithm for a letter-based representation of all-pairwise comparisons.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models An algorithm for a letter-based representation of all-pairwise comparisons

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.312220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:a40c71a3c2df09b6050111ed82f58e18414efef6ae1b9ad96960971366b2945c

Observation a1bd9aa2-04f6-416b-98ee-42c666bd7f75 · outbound

This paper cites Learning transferable visual models from natural language supervision.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Learning transferable visual models from natural language supervision

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.274623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:1f32c6121f1cf3489cc95958cdf3897b86cec52c2c4bb8437115a509712b7c3b

Observation 2518daef-95c7-47bf-9e99-3f576caa9316 · outbound

This paper cites Manning, 2024.isbn: 978-1633437166.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Manning, 2024.isbn: 978-1633437166

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.277507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:c5b82d7798dcae955757c2e972847443ca2aff2162b98f094ebc8602a9d4d9cb

Observation 82272919-f7c8-41a2-affe-8abbec3749fc · outbound

This paper cites Rasley, S.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Rasley, S

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T02:11:57.138780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:dd54b1e6cbfbbb2435efcfa2dcdc38fd150477d2b2847912846c68505b4c7c1d

Observation 7513b13e-38a7-41ea-aaa7-660737b15e44 · outbound

This paper cites WinoGrande: An Adversarial Winograd Schema Challenge at Scale.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models WinoGrande: An Adversarial Winograd Schema Challenge at Scale

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.321056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:5f357a5840153fbc815029161f400ada9c0c364f995cfe8b8793484e658d8669

Observation 1a17e312-15df-405a-93b2-2b60b661360f · outbound

This paper cites 2024.url:https: //github.com/ServiceNow/Fast-LLM.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models 2024.url:https: //github.com/ServiceNow/Fast-LLM

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.232086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:c0c46c7a8964fcc889c8f8a9dc769d38057937b88319954cba51d87777d08a6f

Observation 3eee2fd4-e2e2-49cc-a433-812ce41d7eda · outbound

This paper cites Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.228975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:b52acec1530e953ca78584c02c4031b5504cc4964ec438583879ec797d320b67

Observation 3564188d-df00-4ff2-b7e2-5672777733c0 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:16:04.002704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:47cc93bf3595da4abf9d13c7e2281304c824314f555f0802a0739fdf4fdfc101

Observation 01dc065e-13dc-44b0-9930-15e5508da823 · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:22:37.721356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:cebce2debf818041d156df368c7e1fbfb76337b1e605b2bfe64ddd6af0dcc4fe

Observation 106c0a77-2643-4bb3-9f47-3df54c9f5fc7 · outbound

This paper cites DINOv3.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models DINOv3

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:16:03.793832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:c6f215d9cac58b755f0830b963242ebd6b2400ce5bd466619d422ae150a0de44

Observation 4848d4c8-bc38-4298-ae1f-53eb97487984 · outbound

This paper cites Is Your Imitation Learning Policy Better Than Mine? Policy Comparison with Near-Optimal Stopping.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Is Your Imitation Learning Policy Better Than Mine? Policy Comparison with Near-Optimal Stopping

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.251827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:72ada2d815c1f8881adcb9be70180caeae831c93ba52fc1d002c152659c6dbf8

Observation 723359a7-04b9-4800-af29-a28cd10e934e · outbound

This paper cites Big little lies: A compendium and simulation of p-hacking strategies.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Big little lies: A compendium and simulation of p-hacking strategies

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.242948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:7f262cbdbb63a7b277815e615ab219ffc384f8148161dccafe166af01915d0f4

Observation 6dbafd0d-1e2d-4d47-a33d-683189868e09 · outbound

This paper cites A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:32:56.991484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:ea156ad1a9f8e63b753c1b3b29ecd9cfb9e3664cc32186cbbddf8d2bff1aee4f

Observation 0811941d-d34e-4a50-892b-ed10804ce04d · outbound

This paper cites Toyota Research Institute.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Toyota Research Institute

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.315024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:fb3ce4646bcc7dfca0c7688908394a4b0c5d7e4bffa22ce422171fe57d241667

Observation a982e4b2-d84c-4450-b3dc-126bfb722028 · outbound

This paper cites 2 OLMo 2 Furious.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models 2 OLMo 2 Furious

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:50:29.153357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:76830dcd5abe041065efa48f8de3592730e9fc0bc594cf054da904a84a9eecc4

Observation f275805c-db75-4332-80ae-f2236c1b998b · outbound

This paper cites 2019.url:https://drake.mit.edu.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models 2019.url:https://drake.mit.edu

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.246049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:a9ceb7db0de2886afbd8e219d44df94041bff8c6fabf3846b5b03a81b76c29ea

Observation 1d59fd8e-a717-47b5-ac00-d5c499006182 · outbound

This paper cites GitHub repository.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models GitHub repository

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.249014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:a1d8f8cad972611a2817e7d7e4fc8de38f551026e94a92be0ec410d1f4395817

Observation 6e4cb887-27aa-4e56-834d-92344c03c98c · outbound

This paper cites an unresolved cited work.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-05-23T00:32:19.234476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:ea8f7d82b7e213f5a4c396b1c39e70ec61300b490529a434067b46251a0c7efe

Observation c7769bf4-f229-437e-b6d9-241d3a75ee81 · outbound

This paper cites Dexbotic: Open-source vision-language-action toolbox.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Dexbotic: Open-source vision-language-action toolbox

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:16:03.838811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:39d370378fa985d054bf0337f470539ec7c1cf0ebb2cb916df880064503245cc

Observation c7c378db-cbb9-470b-a50a-f8512867cd83 · outbound

This paper cites World Action Models are Zero-shot Policies.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models World Action Models are Zero-shot Policies

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:18:16.286837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:d662ae68407fc0d84b903acba3f6778f532132faa327cf1dde82aae0768d5f96

Observation 894d1ed0-5869-4069-9a89-ecf8d900cdd3 · outbound

This paper cites URL https:// doi.org/10.18653/v1/p19-1472.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models URL https:// doi.org/10.18653/v1/p19-1472

Reference 74

Resolution
metadata mismatch
doi, observed 2026-05-10T02:11:57.129366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:8cb4889f734d95a4a847169ebc2187e264d589a2a3bc1bd0dbf21eb24e212e74

Observation ac30ec90-586d-4df4-9f7a-2ce04e090ef0 · outbound

This paper cites Sigmoid loss for language image pre-training.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Sigmoid loss for language image pre-training

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.338788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:be1830d56df97388837a265d7ecf6870f2cd0da2b1f2f345fab0034f23a50978

Observation 76621654-bdf7-43d0-8496-03f14b427c02 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 76

Resolution
malformed identifier
raw_fallback, observed 2026-05-23T00:32:19.237497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:f0689975dbc81529ce2e552a250bc88696fbf0298c680202eb6b9f9519892686

Observation 8201eb28-43bf-456b-b6b1-5d3c83b3e5a0 · outbound

This paper cites On the Continuity of Rotation Representations in Neural Networks.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models On the Continuity of Rotation Representations in Neural Networks

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T00:32:19.255535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:2704caecd34975724148e8c65906c5afecf88928678a5fa9e39ef58283978896

Observation 60ad49a7-19f9-4b68-acc7-3070df3d44f7 · outbound

This paper cites an unresolved cited work.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-05-23T00:32:19.263772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:d08f592ff6299fae8df600eedb4fdbc71fd9e3e3047e0ad77bbe6785e3196673

Pith citing papers

Observation b7927a69-cb4f-48f4-899d-5e17cc498f62 · inbound

EVA-Client: A Unified Data Collection, Inference, and Deployment Framework for Embodied Policies on Real Robots cites this paper.

EVA-Client: A Unified Data Collection, Inference, and Deployment Framework for Embodied Policies on Real Robots VLA Foundry: A Unified Framework for Training Vision-Language-Action Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T07:59:34.398439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:59:34.398439Z digest=sha256:63472b86cdd658cb595d545789eb6b5385bd25f92afaa653ddb8709d81ca2d33