Pith. sign in

Paper Citation Record · LEDGER

From Foundation to Application: Improving VLA Models in Practice

As of 18 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 11 inbound Pith citation observations for arXiv:2607.06403.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06403 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-08T07:06:13.473493Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:39:40.520150Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T14:52:36.159020Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact31
  • verified fuzzy14
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5cfc9951-4d25-4d42-a0ba-7eb2e0803ad2 · outbound

This paper cites Cosmos 3: Omnimodal World Models for Physical AI.

From Foundation to Application: Improving VLA Models in Practice Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.485926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:662066bd1ad156da1de6bb982b379f475e9ef6124cd88a87d72c3d6e083caf3d

Observation ad679c6b-c74c-46c5-8811-c36109dc0cf5 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

From Foundation to Application: Improving VLA Models in Practice V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.493399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:2ef33bb0383271fc5b7189faef5029758abb8b31f845c5b3bfae683b7b5a7ce7

Observation 003e4f82-c2c8-4c38-9672-109e2ab194a2 · outbound

This paper cites HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation.

From Foundation to Application: Improving VLA Models in Practice HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.498061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:f327f7b75962f9aa4e9117405aecc00758ea9802ed82970ffc558a19a20e1cfc

Observation f74c6e13-d353-4e0f-8184-eb646ba6a0bc · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

From Foundation to Application: Improving VLA Models in Practice GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.502118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:a96e13d904294a7c7646c9304dbd51f6fbbaf97748b20777ffc7f81f17c1d4bb

Observation 86cab0f7-a598-4d4c-b632-021ef85ad981 · outbound

This paper cites an unresolved cited work.

From Foundation to Application: Improving VLA Models in Practice Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-07-08T07:24:43.712960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:09a597e5362a25cf02145e3df2b9981b4f27d674fe07b93c60e935e4da0c1a18

Observation dda130c9-d3f6-40bd-a14a-ac9bfca4c787 · outbound

This paper cites InProceedings of Robotics: Science and Systems.

From Foundation to Application: Improving VLA Models in Practice InProceedings of Robotics: Science and Systems

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.704491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:0984d4e344ffe24e70304da43668c76972db1fe307f0b8efc1a9f7a428318896

Observation 91d09b23-e93c-4f91-a230-f1db8a83125a · outbound

This paper cites arXiv preprint arXiv:2602.12684 (2026).

From Foundation to Application: Improving VLA Models in Practice arXiv preprint arXiv:2602.12684 (2026)

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.505934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:127aaebc6421b7e13c99868c04bb43808bbf0863cac9b3fb115e4d45aaa3dfa0

Observation 7d8a7ae8-e647-4482-9483-5308709382ad · outbound

This paper cites GR-3 Technical Report.

From Foundation to Application: Improving VLA Models in Practice GR-3 Technical Report

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.482564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:f41f4d5f6a65c28feeaecc7f5fb3e2363681f5669cb6d93192be1c0e1d820679

Observation 84bce366-b3c6-4b47-aa2e-27f5c335d9db · outbound

This paper cites Lawam: Latent world action models for efficient dynamics-aware robot policies.arXiv preprint arXiv:2606.15768, 2026.

From Foundation to Application: Improving VLA Models in Practice Lawam: Latent world action models for efficient dynamics-aware robot policies.arXiv preprint arXiv:2606.15768, 2026

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.468230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:a0e5ebe09b289bf3ab3e00eeca7165b80886621fe903342953d977358ac01ecd

Observation 49ec92c7-707e-4b92-9041-be7498122d83 · outbound

This paper cites ABot-M0.5: Unified Mobility-and-Manipulation World Action Model.

From Foundation to Application: Improving VLA Models in Practice ABot-M0.5: Unified Mobility-and-Manipulation World Action Model

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.472368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:bbe3d3504968ce4d317ed3caa9042b4c5fc81903f51922d02d3a90e1cf86036f

Observation bc4f3dae-fafa-47c1-ba37-c76f365535c6 · outbound

This paper cites Dexworldmodel: Causal latent world modeling towards automated learning of embodied tasks, 2026.

From Foundation to Application: Improving VLA Models in Practice Dexworldmodel: Causal latent world modeling towards automated learning of embodied tasks, 2026

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.675975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:55866ad13ea863195a164d6e3161ab9617175da9b2554afb70e8b6465cfffe18

Observation 22113d00-53bb-49f0-bb94-23dbbf411775 · outbound

This paper cites HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies.

From Foundation to Application: Improving VLA Models in Practice HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-09T02:19:49.242114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:7f60dd858dc6079b316ffa065a41871fee1fabedb595703ac0e870eb28213dd7

Observation cc3a46fa-a757-4bb3-a01b-df17909688d3 · outbound

This paper cites Galaxea g0.5 technical report, 2026.

From Foundation to Application: Improving VLA Models in Practice Galaxea g0.5 technical report, 2026

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.710268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:1dd3794199b12d5b8533b60bad5f8534bf2deee85b605fc8a1cc60f34fa3ee85

Observation 853afccb-f2b6-4f39-8a32-ce3bea1f01f7 · outbound

This paper cites Galaxea Open-World Dataset and G0 Dual-System VLA Model.

From Foundation to Application: Improving VLA Models in Practice Galaxea Open-World Dataset and G0 Dual-System VLA Model

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.462253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:0ba0894231ba575b3502711b6fb0e718904a77bd4bb0f2c1d63f9bfb7a84baa3

Observation b0d968cf-f23c-4deb-b81d-c421f2bd1c1c · outbound

This paper cites RLDX-1 Technical Report.

From Foundation to Application: Improving VLA Models in Practice RLDX-1 Technical Report

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.478990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:12d121a44387941cb2047c236a40101808721670054de6929b9e7c401c7853d9

Observation 10e80ea4-bf94-4929-b266-8784b39cb6cf · outbound

This paper cites OpenVLA: An open-source vision-language-action model.

From Foundation to Application: Improving VLA Models in Practice OpenVLA: An open-source vision-language-action model

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.681788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:f533de3825ddfb98d40511ac12c7f44d72c9bcd33bffee34877648c2f7c82446

Observation fc61d89c-b1a4-4adc-ba11-40f28e2e2673 · outbound

This paper cites Forcevla2: Unleashing hybrid force-position control with force awareness for contact-rich manipulation.

From Foundation to Application: Improving VLA Models in Practice Forcevla2: Unleashing hybrid force-position control with force awareness for contact-rich manipulation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.454088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:b0a26524c6dd8fe444b6b0ab5c094932c1830a753b473eac23a6ff2d899b5505

Observation 414594b6-6e00-4fa7-923a-6c634c02b9fe · outbound

This paper cites HoloBrain-0 technical report.

From Foundation to Application: Improving VLA Models in Practice HoloBrain-0 technical report

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.438202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:5ab7d954846ae6d4c80910297df98dbed8a7186ec52e45b3021b312a88133add

Observation 7a9673b0-6999-4e6f-9221-507e0d6c44ba · outbound

This paper cites DeepSeek-V3 Technical Report.

From Foundation to Application: Improving VLA Models in Practice DeepSeek-V3 Technical Report

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.441802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:a4b681f747aff222a8301275f1644d84890e9192273e4a35f6d394a624e1098e

Observation c898bc55-29e6-4ff2-b04c-63542be34257 · outbound

This paper cites Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos.

From Foundation to Application: Improving VLA Models in Practice Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.445289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:76d989d2367708f7e466a0e9cca937762d8d570b1f0cf1c20f0418d64e0f5720

Observation 6c238041-d146-4740-8dc5-4dd34b9f904b · outbound

This paper cites Being-h0.

From Foundation to Application: Improving VLA Models in Practice Being-h0

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T07:14:45.449397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:d2518ffe18a3eee1b05bc460f06ffa5ecebc8213f013e0d5367d8e11eba68d05

Observation 91bcbe7b-0b66-45da-a1f9-82fd7c9fa25c · outbound

This paper cites Being-H0.7: A Latent World-Action Model from Egocentric Videos.

From Foundation to Application: Improving VLA Models in Practice Being-H0.7: A Latent World-Action Model from Egocentric Videos

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.425343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:1a98e29e00e620f4ab586df9dd1ae02558edf4cbdf686e76ef6d2dd15d518744

Observation 9c7c5ceb-597e-468c-b36b-ef4f15c62d99 · outbound

This paper cites LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion.

From Foundation to Application: Improving VLA Models in Practice LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T07:14:45.429695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:1e1520e39632c7d9236b9fe483a203851a59d8f06fd5aede8c533f8222d8854d

Observation 09d20cf5-02cc-40c1-b310-e721e2ea7bbf · outbound

This paper cites LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment.

From Foundation to Application: Improving VLA Models in Practice LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.433532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:661555aa6e392f1ce5ff49fdf9e8b798729db4606c6dcc7c8a4d3eacefc87100

Observation a7e6a948-fdb6-4cbc-9887-7506ede97bf2 · outbound

This paper cites Qwen3.6-27B: Flagship-level coding in a 27B dense model, April 2026.

From Foundation to Application: Improving VLA Models in Practice Qwen3.6-27B: Flagship-level coding in a 27B dense model, April 2026

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.693592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:7967f633bf903818f6bc846bec7243a1a6f3ac5dd6748c6a8029c36c090a0171

Observation 200de719-b032-4001-99d9-62d85cf3582a · outbound

This paper cites an unresolved cited work.

From Foundation to Application: Improving VLA Models in Practice Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-07-08T07:24:43.695786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:429dc0369f067311e6de52dce63d7321f5384de29d8e3314dcf49de18fb7776f

Observation a50f5f59-ac4b-4439-9906-daba83719e83 · outbound

This paper cites World guidance: World modeling in condition space for action generation.

From Foundation to Application: Improving VLA Models in Practice World guidance: World modeling in condition space for action generation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.687497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:f30ad3dd2a80a071ebddc4b1f325f2ed585d06024a17c59aef7e8849b28a9f24

Observation 1d7926b9-6a1e-44e7-b461-e6af4cac3c85 · outbound

This paper cites Masked depth modeling for spatial perception.

From Foundation to Application: Improving VLA Models in Practice Masked depth modeling for spatial perception

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.678764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:e4e3d858d8c871e6c965d47b037edd8a44e49cb898fe4136327d2e2c6292c03e

Observation d86c13d0-8e17-4dfe-ac6a-dab771786a42 · outbound

This paper cites Towards human-like manipulation through RL-augmented teleoperation and mixture-of-dexterous-experts VLA.arXiv preprint arXiv:2603.08122, 2026.

From Foundation to Application: Improving VLA Models in Practice Towards human-like manipulation through RL-augmented teleoperation and mixture-of-dexterous-experts VLA.arXiv preprint arXiv:2603.08122, 2026

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.408586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:ab0b8bfd6592b3a49a05a23a011e710aaa4056d100c7cd6ae39f62d71e1df69b

Observation f87599c4-14d9-4a46-97d3-26cbd1f2227d · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

From Foundation to Application: Improving VLA Models in Practice Gemini Robotics: Bringing AI into the Physical World

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.412977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:ed94af1ee5046e9a39bf26fcc917a8b2a9266ab8b7afa2578c1f1ef4a6b49c72

Observation e1815cef-e8cf-43ef-8c62-97586520cf08 · outbound

This paper cites Gen-1: Scaling embodied foundation models to mastery.Generalist AI Blog, 2026.

From Foundation to Application: Improving VLA Models in Practice Gen-1: Scaling embodied foundation models to mastery.Generalist AI Blog, 2026

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.701906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:10a084d77e1a867f107a43eface2750caa308668fb07edbb1f9e04096e2adac8

Observation 6fc84b0e-6424-425a-b3e2-68d61766908b · outbound

This paper cites GR00T N1.6: An improved open foundation model for generalist humanoid robots.

From Foundation to Application: Improving VLA Models in Practice GR00T N1.6: An improved open foundation model for generalist humanoid robots

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.684694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:c36973af2f16a299323b1e2fc07f35807d0659ac4115bc249c379bdcad4495b4

Observation 9a260665-5e6e-4215-bd3a-d4184da1253e · outbound

This paper cites Qwen-robotmanip technical report: Alignment unlocks scale for robotic manipulation foundation models.

From Foundation to Application: Improving VLA Models in Practice Qwen-robotmanip technical report: Alignment unlocks scale for robotic manipulation foundation models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.707305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:d69beb4c4411ad912238b13d350d8b3dd447cfccf9f4b046d8ca91c7fde565f6

Observation e3815cc7-0ec4-4778-a075-b92499ecd99c · outbound

This paper cites Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments.

From Foundation to Application: Improving VLA Models in Practice Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.417284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:9787f2260304168edbc735b8c441cb0feaf2eb4463076dd0a43a65df1a3c9cd0

Observation dfc55af1-d17c-4f73-bb2c-ae5e12688dc7 · outbound

This paper cites Vision-centric activation and coordination for multimodal large language models.

From Foundation to Application: Improving VLA Models in Practice Vision-centric activation and coordination for multimodal large language models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.398070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:2b6b66666e9a856a1e8ca02cc62c5c507bdb7e1c0dd23c6aae45c8d25c95e0c9

Observation e14e13a8-c8c0-411d-8d07-72518703617a · outbound

This paper cites The Great March 100: 100 detail-oriented tasks for evaluating embodied ai agents.

From Foundation to Application: Improving VLA Models in Practice The Great March 100: 100 detail-oriented tasks for evaluating embodied ai agents

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.698622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:0efe3642115d687a5194ea66ba663f1df5715127760c447704005c06dfe14164

Observation a33f30d5-2454-4287-bcc9-6475f0bd0426 · outbound

This paper cites Videorope: What makes for good video rotary position embedding? InInt.

From Foundation to Application: Improving VLA Models in Practice Videorope: What makes for good video rotary position embedding? InInt

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.690508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:541058e8abd0d28a5b0c7271e20336f5e15d906e510d244e0f64ecedad37ffb5

Observation fad77dea-ea50-4f54-a5e8-29f3e5b9bfa3 · outbound

This paper cites A Pragmatic VLA Foundation Model.

From Foundation to Application: Improving VLA Models in Practice A Pragmatic VLA Foundation Model

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.387940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:64daca65e7a100ea9b92d041797bada06348f1b1d26e9876a1f19fc590eca456

Observation 826ef41b-3fe4-4b80-82aa-a6989304975a · outbound

This paper cites Hy-embodied-0.5-x: An enhanced embodied foundation model for real-world agents.

From Foundation to Application: Improving VLA Models in Practice Hy-embodied-0.5-x: An enhanced embodied foundation model for real-world agents

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.718522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:2093055882b1cc81746f4c003a5b1017bb1e6dae62f2381aa72e3c1e1296c7df

Observation b24c01e9-c49b-49d3-9a91-7304df6f9778 · outbound

This paper cites OmniStream: Mastering perception, reconstruction and action in continuous streams.

From Foundation to Application: Improving VLA Models in Practice OmniStream: Mastering perception, reconstruction and action in continuous streams

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.393708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:1571b03b34383f4ba25c5b9920a875173ed64c74bd69281ab12a6d8a3cb43ee5

Observation 0151f29a-b4ac-4cd7-841c-c12b40ffbc18 · outbound

This paper cites Magma: A foundation model for multimodal AI agents.

From Foundation to Application: Improving VLA Models in Practice Magma: A foundation model for multimodal AI agents

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.715649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:c00cca2b4f5010bd9145b2b84e2d79c3e9f7d58b054eeb0a98764e5955d01e58

Observation 233158c1-bc37-43e6-9c03-04ddb6cefe9f · outbound

This paper cites ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning.

From Foundation to Application: Improving VLA Models in Practice ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.381619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:a1a252baa2d2c3817d94285bcee0dc6f35cf7d203e8965ac8d6eded2398f0f41

Observation 77d13554-2337-4648-a28b-1f0e9f6599ba · outbound

This paper cites StarVLA-$\alpha$: Reducing Complexity in Vision-Language-Action Systems.

From Foundation to Application: Improving VLA Models in Practice StarVLA-$\alpha$: Reducing Complexity in Vision-Language-Action Systems

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.385142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:f46d3d4d80ad3c0208d09fbd230853fb599d49838b9cdfbe421dafba77532923

Observation 738073c5-f75c-40ba-aed7-3ee744efdefc · outbound

This paper cites Samoe- vla: A scene adaptive mixture-of-experts vision-language-action model for autonomous driving.

From Foundation to Application: Improving VLA Models in Practice Samoe- vla: A scene adaptive mixture-of-experts vision-language-action model for autonomous driving

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.403016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:8cd3a6df490f8b0cbf702d60e09df9710bddea7d5aea855e813c102cc73247d2

Observation 7019218b-1805-44b6-9062-ccefc4d59f97 · outbound

This paper cites Forcevla: Enhancing vla models with a force-aware moe for contact-rich manipulation.

From Foundation to Application: Improving VLA Models in Practice Forcevla: Enhancing vla models with a force-aware moe for contact-rich manipulation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.421603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:b6056aff983f12ca480c43f298daf4162703f5c5f44585f7b911ed10980ab6d1

Observation 9d0f20ee-5153-47d9-8194-1471d3e15e7f · outbound

This paper cites Wall-OSS-0.5 Technical Report.

From Foundation to Application: Improving VLA Models in Practice Wall-OSS-0.5 Technical Report

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.374157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:efa7af26193fadedd8ca94e64119502c8f965da8a9ffd33d3c67b875002caf94

Observation b61eabe8-f1df-46c2-bb7a-665fb8187bde · outbound

This paper cites Igniting vlms toward the embodied space.

From Foundation to Application: Improving VLA Models in Practice Igniting vlms toward the embodied space

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.377496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:04fd7b8206423fa3baf2d58f31631877fae103c4011f02d4e4bf1adde5d316e5

Observation b610e1b9-44f4-4f9c-867c-890dbb7e5f3b · outbound

This paper cites AtomicVLA: Unlocking the potential of atomic skill learning in robots.arXiv preprint arXiv:2603.07648, 2026.

From Foundation to Application: Improving VLA Models in Practice AtomicVLA: Unlocking the potential of atomic skill learning in robots.arXiv preprint arXiv:2603.07648, 2026

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.370620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:c216b4b261270dd827bb8ac5bb583f6c0c7e2d5243f8749fe2b37f9a9d3f21bb

Observation 5f4a89b6-e6f7-4190-8252-542136a69e98 · outbound

This paper cites GEM: Generative Supervision Helps Embodied Intelligence.

From Foundation to Application: Improving VLA Models in Practice GEM: Generative Supervision Helps Embodied Intelligence

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.366955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:2106e33df7672c31e2b9861ac4f6f62ea1a82d4c8f51bf39545beb6e3c720b36

Pith citing papers

Observation 98631932-0bc3-4b2c-b9cf-0519035d1d01 · inbound

Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation cites this paper.

Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation From Foundation to Application: Improving VLA Models in Practice

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T16:20:01.326319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:20:01.326319Z digest=sha256:ac0917093bac952a038456047119ba2ea66384748ee778bd26e39b744adba5b6

Observation b9201e3b-071c-424e-9891-3f7e9cbdeee6 · inbound

$N_0$-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation cites this paper.

$N_0$-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation From Foundation to Application: Improving VLA Models in Practice

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-30T12:42:18.596168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:42:18.596168Z digest=sha256:09f08d1683eb00b663448d04f5d18e8f0af766f77186bb432b8d2b3ee6f63adb

Observation a10fb45f-1def-4378-b5c1-d489d00bbf8c · inbound

Route by Kinematics, Act by Observation: Kinematics-Supervised Expert Routing in MoE-Augmented VLA cites this paper.

Route by Kinematics, Act by Observation: Kinematics-Supervised Expert Routing in MoE-Augmented VLA From Foundation to Application: Improving VLA Models in Practice

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-30T20:30:33.298679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T20:30:33.298679Z digest=sha256:77ea4108663a36cc3c1844901d4c456399e68e0183b85e0fbf6ce663db426784

Observation 1da9a5a6-7138-4126-b25e-e09a868d0d89 · inbound

OC-VLA++: Monocular Geometry-Guided Cross-View Consistency for Viewpoint-Robust Robotic Manipulation cites this paper.

OC-VLA++: Monocular Geometry-Guided Cross-View Consistency for Viewpoint-Robust Robotic Manipulation From Foundation to Application: Improving VLA Models in Practice

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T15:17:24.954806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:17:24.954806Z digest=sha256:4226fddd40565580d20a3d8147a5762a216d4105a5dcfbaf955214ff4aba62f8

Observation dc7be8e0-8ee9-462e-97ca-e942cd233604 · inbound

Learning Panorama-Aware VLA for Mobile Manipulation with Whole-Body Teleoperation cites this paper.

Learning Panorama-Aware VLA for Mobile Manipulation with Whole-Body Teleoperation From Foundation to Application: Improving VLA Models in Practice

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T10:39:54.173311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:39:54.173311Z digest=sha256:b390206e356c51a888330e758278416caf38bcc8371dfd771078704c6db7991d

Observation d3e28936-f8d4-4800-9924-c778c4249614 · inbound

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud cites this paper.

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud From Foundation to Application: Improving VLA Models in Practice

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T14:52:36.161955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T14:52:35.400227Z digest=sha256:a839d216931ed6b4ac643bc9c0f081bef4c21a6b1d1e6b811dc8f09c4dc38816

Observation 2e0ad503-8cce-4b5c-bb7f-c4ed5f8d8c9d · inbound

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud cites this paper.

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud From Foundation to Application: Improving VLA Models in Practice

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T14:52:22.234915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:52:22.234915Z digest=sha256:e641a1b4c1f45a417132960741ae768dcfe4f2c4bbcee4673699060264ca70d8

Observation 0aaf8514-feb4-413c-919f-035b15ed4ac0 · inbound

FACT: Failure-Aware Causal Training for World-Action Models cites this paper.

FACT: Failure-Aware Causal Training for World-Action Models From Foundation to Application: Improving VLA Models in Practice

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:39.641954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:39.641954Z digest=sha256:05ae2eca40fcd2e5f7f0003f5582db574b9c11f6993de0926cf7d866ae865931

Observation 50d7803a-337b-4a30-844b-00f571cee7df · inbound

Flex-$\pi$: A Multi-Stream World-Action Model with Compute Flexibility cites this paper.

Flex-$\pi$: A Multi-Stream World-Action Model with Compute Flexibility From Foundation to Application: Improving VLA Models in Practice

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T15:40:38.437224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:40:38.437224Z digest=sha256:1f30a11d63fff2b26895b52575534c616ab330c9f7843839afa8429d32c9f134

Observation 553899b5-2105-499a-9572-6a57a3fa57aa · inbound

Flex-$\pi$: A Multi-Stream World-Action Model with Compute Flexibility cites this paper.

Flex-$\pi$: A Multi-Stream World-Action Model with Compute Flexibility From Foundation to Application: Improving VLA Models in Practice

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T14:17:54.845505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:17:54.845505Z digest=sha256:5dad23349a69e53a6d4962391e2469f6fd5b5bc50c36dc03841fc965eb84bbb3

Observation 479b66dc-ecf4-4bec-b458-164d4580e070 · inbound

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models cites this paper.

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models From Foundation to Application: Improving VLA Models in Practice

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.520150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.520150Z digest=sha256:8af2f945a96ad5f929cf5f5ac34ad6364d476c5822d594fdc820b6be14757b71