Pith. sign in

Paper Citation Record · LEDGER

From Foundation to Application: Improving VLA Models in Practice

As of 28 July 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2607.06403.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06403 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-08T07:06:13.473493Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-28T06:31:03.373048+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact31
  • verified fuzzy16
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5cfc9951-4d25-4d42-a0ba-7eb2e0803ad2 · outbound

This paper cites Cosmos 3: Omnimodal World Models for Physical AI.

From Foundation to Application: Improving VLA Models in Practice Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.485926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:45438db6b875bbd5f3134ed69f961e5d3720693f73369e1520f30fff29456d42

Observation ad679c6b-c74c-46c5-8811-c36109dc0cf5 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

From Foundation to Application: Improving VLA Models in Practice V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.493399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:b07902589651414bd5eaaf6adc96f5ee3dbd7bc5d31d3e85e7adc8477ae01dca

Observation 003e4f82-c2c8-4c38-9672-109e2ab194a2 · outbound

This paper cites HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation.

From Foundation to Application: Improving VLA Models in Practice HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.498061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:27e8dc40a3e4cc06b3a8dbcb6ece9e42455055da6815ffcda38445c15a5a45e4

Observation f74c6e13-d353-4e0f-8184-eb646ba6a0bc · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

From Foundation to Application: Improving VLA Models in Practice GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.502118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:12818d2ae8e28f91a3a09dc4f2f7111916c629070cfcf27eed6a376cb631d573

Observation 86cab0f7-a598-4d4c-b632-021ef85ad981 · outbound

This paper cites an unresolved cited work.

From Foundation to Application: Improving VLA Models in Practice Unresolved cited work

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.712960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:8d3f75d50feb39a970c51ed5f52fd1acbb474ce1fa90742d7a4c1eac4244bc13

Observation dda130c9-d3f6-40bd-a14a-ac9bfca4c787 · outbound

This paper cites InProceedings of Robotics: Science and Systems.

From Foundation to Application: Improving VLA Models in Practice InProceedings of Robotics: Science and Systems

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.704491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:4c245f6aae3ebc0352587fbdc2d3c2abbf2fd7e0d363bf7efd1e4cefc796ed0a

Observation 91d09b23-e93c-4f91-a230-f1db8a83125a · outbound

This paper cites arXiv preprint arXiv:2602.12684 (2026).

From Foundation to Application: Improving VLA Models in Practice arXiv preprint arXiv:2602.12684 (2026)

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.505934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:cfd959f34d74a872e626f36109f7e089e39d1d62042fac0d33682ce65f9b76ca

Observation 7d8a7ae8-e647-4482-9483-5308709382ad · outbound

This paper cites GR-3 Technical Report.

From Foundation to Application: Improving VLA Models in Practice GR-3 Technical Report

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.482564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:fa3f0d266f93f4674e10c41ca63e16b0ea94662dcb7154dff1ac5e9fd6ea71c5

Observation 84bce366-b3c6-4b47-aa2e-27f5c335d9db · outbound

This paper cites Lawam: Latent world action models for efficient dynamics-aware robot policies.arXiv preprint arXiv:2606.15768, 2026.

From Foundation to Application: Improving VLA Models in Practice Lawam: Latent world action models for efficient dynamics-aware robot policies.arXiv preprint arXiv:2606.15768, 2026

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.468230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:c7b9178a7c5c8ca2d220a0972d4ff1af0d107f65989289dbfbb4609a255c26e3

Observation 49ec92c7-707e-4b92-9041-be7498122d83 · outbound

This paper cites ABot-M0.5: Unified Mobility-and-Manipulation World Action Model.

From Foundation to Application: Improving VLA Models in Practice ABot-M0.5: Unified Mobility-and-Manipulation World Action Model

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.472368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:5cf17f4837bfd716721cd93d7532abbd54847cdd61faa36b9d5ff7fb5e38a9f0

Observation bc4f3dae-fafa-47c1-ba37-c76f365535c6 · outbound

This paper cites Dexworldmodel: Causal latent world modeling towards automated learning of embodied tasks, 2026.

From Foundation to Application: Improving VLA Models in Practice Dexworldmodel: Causal latent world modeling towards automated learning of embodied tasks, 2026

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.675975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:79ae66c8b08fe87d3ea74b2f03af4dc15ded212d3e5625001d7fcc828b498248

Observation 22113d00-53bb-49f0-bb94-23dbbf411775 · outbound

This paper cites HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies.

From Foundation to Application: Improving VLA Models in Practice HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-09T02:19:49.242114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:d2bd90679d39e0e2463bb86be9838fb1a56f8f836a712579f806bb49f0ddf67f

Observation cc3a46fa-a757-4bb3-a01b-df17909688d3 · outbound

This paper cites Galaxea g0.5 technical report, 2026.

From Foundation to Application: Improving VLA Models in Practice Galaxea g0.5 technical report, 2026

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.710268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:1099d01e258f523128a957177269ce4358af3fed7892b6651e727c4900d9898a

Observation 853afccb-f2b6-4f39-8a32-ce3bea1f01f7 · outbound

This paper cites Galaxea Open-World Dataset and G0 Dual-System VLA Model.

From Foundation to Application: Improving VLA Models in Practice Galaxea Open-World Dataset and G0 Dual-System VLA Model

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.462253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:2cc6bdcc9541e89e30ccda9d49175627300563c7136621e440b40394fe75214d

Observation b0d968cf-f23c-4deb-b81d-c421f2bd1c1c · outbound

This paper cites RLDX-1 Technical Report.

From Foundation to Application: Improving VLA Models in Practice RLDX-1 Technical Report

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.478990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:eb96ed13e9101bcc81fc9d83cb2a70835e8d7999cc2269ee0d6a510881fba9e8

Observation 10e80ea4-bf94-4929-b266-8784b39cb6cf · outbound

This paper cites OpenVLA: An open-source vision-language-action model.

From Foundation to Application: Improving VLA Models in Practice OpenVLA: An open-source vision-language-action model

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.681788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:de2c83b99aff809ca98c85483012d2c5968b2aa4cce4a264bc8c31798cc21d08

Observation fc61d89c-b1a4-4adc-ba11-40f28e2e2673 · outbound

This paper cites Forcevla2: Unleashing hybrid force-position control with force awareness for contact-rich manipulation.

From Foundation to Application: Improving VLA Models in Practice Forcevla2: Unleashing hybrid force-position control with force awareness for contact-rich manipulation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.454088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:ccef5ce54267ae00fe0c44804cd0457117f072e49d1fffd6f463888a2221afb7

Observation 414594b6-6e00-4fa7-923a-6c634c02b9fe · outbound

This paper cites HoloBrain-0 technical report.

From Foundation to Application: Improving VLA Models in Practice HoloBrain-0 technical report

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.438202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:f7b04fd97dd84f640fbe5adcb7600f70c27f56f55381cbc45f609439b9851d7f

Observation 7a9673b0-6999-4e6f-9221-507e0d6c44ba · outbound

This paper cites DeepSeek-V3 Technical Report.

From Foundation to Application: Improving VLA Models in Practice DeepSeek-V3 Technical Report

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.441802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:9425b1250a28c3e821a836f76e7a9197c585aada5001fefd738eea147979cb9d

Observation c898bc55-29e6-4ff2-b04c-63542be34257 · outbound

This paper cites Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos.

From Foundation to Application: Improving VLA Models in Practice Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.445289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:814eb0576220a40a474c1b7b007e4ba141c6685147cd6f217415fd8b1bcbb8fe

Observation 6c238041-d146-4740-8dc5-4dd34b9f904b · outbound

This paper cites Being-h0.

From Foundation to Application: Improving VLA Models in Practice Being-h0

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T07:14:45.449397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:26100c7b8ab708b813bdce8a058cb0feacf5931425bf5ced986d0dcf3b018038

Observation 91bcbe7b-0b66-45da-a1f9-82fd7c9fa25c · outbound

This paper cites Being-H0.7: A Latent World-Action Model from Egocentric Videos.

From Foundation to Application: Improving VLA Models in Practice Being-H0.7: A Latent World-Action Model from Egocentric Videos

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.425343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:ec54adcbe4efde008d5dc5e70aeb442d0ac0667e60eeb994a0daecc6e5f12f36

Observation 9c7c5ceb-597e-468c-b36b-ef4f15c62d99 · outbound

This paper cites LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion.

From Foundation to Application: Improving VLA Models in Practice LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T07:14:45.429695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:3aef25317bbbb9f8471a65a62ea757b340aee67f5eec437454b88bf117c0a039

Observation 09d20cf5-02cc-40c1-b310-e721e2ea7bbf · outbound

This paper cites LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment.

From Foundation to Application: Improving VLA Models in Practice LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.433532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:43e126b40fd9f5fef494d0ba9ccfff0548043ef3b2ea7608cef88dea25665a39

Observation a7e6a948-fdb6-4cbc-9887-7506ede97bf2 · outbound

This paper cites Qwen3.6-27B: Flagship-level coding in a 27B dense model, April 2026.

From Foundation to Application: Improving VLA Models in Practice Qwen3.6-27B: Flagship-level coding in a 27B dense model, April 2026

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.693592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:fab02b311a1ecbb3ebe44bd8d779fd2f1b7e9933f1ff64535a44aa5a9ab72797

Observation 200de719-b032-4001-99d9-62d85cf3582a · outbound

This paper cites an unresolved cited work.

From Foundation to Application: Improving VLA Models in Practice Unresolved cited work

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.695786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:68c00bf5339e70ffd6dd86cf25267d8999ca22a37f1bff034ac2aacfdfc3e973

Observation a50f5f59-ac4b-4439-9906-daba83719e83 · outbound

This paper cites World guidance: World modeling in condition space for action generation.

From Foundation to Application: Improving VLA Models in Practice World guidance: World modeling in condition space for action generation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.687497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:fcd9cf38dc46e8e36c8fa4e7629dfdf4e26a89ce586463d0949e12e4c57207ad

Observation 1d7926b9-6a1e-44e7-b461-e6af4cac3c85 · outbound

This paper cites Masked depth modeling for spatial perception.

From Foundation to Application: Improving VLA Models in Practice Masked depth modeling for spatial perception

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.678764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:cf9600f01e76db31ed33ae79476c508913c8b5110fd551a59329b7b90c172106

Observation d86c13d0-8e17-4dfe-ac6a-dab771786a42 · outbound

This paper cites Towards human-like manipulation through RL-augmented teleoperation and mixture-of-dexterous-experts VLA.arXiv preprint arXiv:2603.08122, 2026.

From Foundation to Application: Improving VLA Models in Practice Towards human-like manipulation through RL-augmented teleoperation and mixture-of-dexterous-experts VLA.arXiv preprint arXiv:2603.08122, 2026

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.408586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:0b3e507451c4b2496ac65883378cec57514eed6fca4e41913c1a6cf80a95a24e

Observation f87599c4-14d9-4a46-97d3-26cbd1f2227d · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

From Foundation to Application: Improving VLA Models in Practice Gemini Robotics: Bringing AI into the Physical World

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.412977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:6691ec2a8340da2277283c1291af91934f3d87eeed105064968c8b7bf1904760

Observation e1815cef-e8cf-43ef-8c62-97586520cf08 · outbound

This paper cites Gen-1: Scaling embodied foundation models to mastery.Generalist AI Blog, 2026.

From Foundation to Application: Improving VLA Models in Practice Gen-1: Scaling embodied foundation models to mastery.Generalist AI Blog, 2026

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.701906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:4db8a64105d50e8f786234b76c9cc29ff57b77fc2f911cbb5b320828bb909a3a

Observation 6fc84b0e-6424-425a-b3e2-68d61766908b · outbound

This paper cites GR00T N1.6: An improved open foundation model for generalist humanoid robots.

From Foundation to Application: Improving VLA Models in Practice GR00T N1.6: An improved open foundation model for generalist humanoid robots

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.684694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:756ebeb58a464e48a6e7b7d60a7765ba363327fc6e19d717e55df4a64d36385c

Observation 9a260665-5e6e-4215-bd3a-d4184da1253e · outbound

This paper cites Qwen-robotmanip technical report: Alignment unlocks scale for robotic manipulation foundation models.

From Foundation to Application: Improving VLA Models in Practice Qwen-robotmanip technical report: Alignment unlocks scale for robotic manipulation foundation models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.707305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:1687d5d403c73b87be2dc3bc3c22edcef457d96ad9f20601e4421ae606c3388a

Observation e3815cc7-0ec4-4778-a075-b92499ecd99c · outbound

This paper cites Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments.

From Foundation to Application: Improving VLA Models in Practice Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.417284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:91a5d7bcee3c764b3104f18ad4dec57bc3e3f52cb78635aa46ce0abab0d87018

Observation dfc55af1-d17c-4f73-bb2c-ae5e12688dc7 · outbound

This paper cites Vision-centric activation and coordination for multimodal large language models.

From Foundation to Application: Improving VLA Models in Practice Vision-centric activation and coordination for multimodal large language models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.398070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:eb57cd4c2c472ef0d0a98993a1c37bc3184e54cb6107a42094a2e8fa6209322b

Observation e14e13a8-c8c0-411d-8d07-72518703617a · outbound

This paper cites The Great March 100: 100 detail-oriented tasks for evaluating embodied ai agents.

From Foundation to Application: Improving VLA Models in Practice The Great March 100: 100 detail-oriented tasks for evaluating embodied ai agents

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.698622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:de039b45e843c624aebe152c8c60df0d6aff1538b92cc0cd78828386f662f08d

Observation a33f30d5-2454-4287-bcc9-6475f0bd0426 · outbound

This paper cites Videorope: What makes for good video rotary position embedding? InInt.

From Foundation to Application: Improving VLA Models in Practice Videorope: What makes for good video rotary position embedding? InInt

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.690508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:e628a3435ef6d1ea2b689099351a190d2994a8ed083f4c1c8ed6e1b7b5eb7670

Observation fad77dea-ea50-4f54-a5e8-29f3e5b9bfa3 · outbound

This paper cites A Pragmatic VLA Foundation Model.

From Foundation to Application: Improving VLA Models in Practice A Pragmatic VLA Foundation Model

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.387940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:dc23cf0dd7f08e5d4e99822c36a6430cb43c45ccb98bc657702ba7b91945c973

Observation 826ef41b-3fe4-4b80-82aa-a6989304975a · outbound

This paper cites Hy-embodied-0.5-x: An enhanced embodied foundation model for real-world agents.

From Foundation to Application: Improving VLA Models in Practice Hy-embodied-0.5-x: An enhanced embodied foundation model for real-world agents

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.718522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:6b871fdf62a1ec25aa10a4e8ed56f80f2402b6f616088bab43c01830a423f3bc

Observation b24c01e9-c49b-49d3-9a91-7304df6f9778 · outbound

This paper cites OmniStream: Mastering perception, reconstruction and action in continuous streams.

From Foundation to Application: Improving VLA Models in Practice OmniStream: Mastering perception, reconstruction and action in continuous streams

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.393708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:6efc448f2fcdf65e9b1d9ae6253a1c1c13be7c9e6eb2f384cd95057e06029d14

Observation 0151f29a-b4ac-4cd7-841c-c12b40ffbc18 · outbound

This paper cites Magma: A foundation model for multimodal AI agents.

From Foundation to Application: Improving VLA Models in Practice Magma: A foundation model for multimodal AI agents

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.715649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:c558eb71f71eb5114dbad751adb261b516e13fb7756ee2f25212033564d5b4fe

Observation 233158c1-bc37-43e6-9c03-04ddb6cefe9f · outbound

This paper cites ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning.

From Foundation to Application: Improving VLA Models in Practice ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.381619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:588e77ece860ba2c25ca764a2e943ada6f23fd295d8a7eed8bd0274ad8cf9dd1

Observation 77d13554-2337-4648-a28b-1f0e9f6599ba · outbound

This paper cites StarVLA-$\alpha$: Reducing Complexity in Vision-Language-Action Systems.

From Foundation to Application: Improving VLA Models in Practice StarVLA-$\alpha$: Reducing Complexity in Vision-Language-Action Systems

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.385142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:79c3d289f5c3377fc791f25ca7194a5826f6896514a7ce3755339dd651cdfa22

Observation 738073c5-f75c-40ba-aed7-3ee744efdefc · outbound

This paper cites Samoe- vla: A scene adaptive mixture-of-experts vision-language-action model for autonomous driving.

From Foundation to Application: Improving VLA Models in Practice Samoe- vla: A scene adaptive mixture-of-experts vision-language-action model for autonomous driving

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.403016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:0e6a948925c84288b4e349fb7463bcf7ac6c2dc16ec9bc64893aa8c2ad11afe5

Observation 7019218b-1805-44b6-9062-ccefc4d59f97 · outbound

This paper cites Forcevla: Enhancing vla models with a force-aware moe for contact-rich manipulation.

From Foundation to Application: Improving VLA Models in Practice Forcevla: Enhancing vla models with a force-aware moe for contact-rich manipulation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.421603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:5a39283cd0f75d6c3ecac670f093cfff86fbc86e9fc2c3c82521aaaac6188faa

Observation 9d0f20ee-5153-47d9-8194-1471d3e15e7f · outbound

This paper cites Wall-OSS-0.5 Technical Report.

From Foundation to Application: Improving VLA Models in Practice Wall-OSS-0.5 Technical Report

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.374157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:9c7f8e839af02cb004c9714216b8b678e29a2283773b21d859afddd2418bdd36

Observation b61eabe8-f1df-46c2-bb7a-665fb8187bde · outbound

This paper cites Igniting vlms toward the embodied space.

From Foundation to Application: Improving VLA Models in Practice Igniting vlms toward the embodied space

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.377496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:3cb02510de1396a4e7f5359df2e94b9ce1acac1987f9a68c9ce202c0133d85c9

Observation b610e1b9-44f4-4f9c-867c-890dbb7e5f3b · outbound

This paper cites AtomicVLA: Unlocking the potential of atomic skill learning in robots.arXiv preprint arXiv:2603.07648, 2026.

From Foundation to Application: Improving VLA Models in Practice AtomicVLA: Unlocking the potential of atomic skill learning in robots.arXiv preprint arXiv:2603.07648, 2026

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.370620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:f41f98829085d34b15dc9288129c83efc43c6768c03fed4b0aa25b0db720258e

Observation 5f4a89b6-e6f7-4190-8252-542136a69e98 · outbound

This paper cites GEM: Generative Supervision Helps Embodied Intelligence.

From Foundation to Application: Improving VLA Models in Practice GEM: Generative Supervision Helps Embodied Intelligence

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.366955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:6ea727d76cf9b832173d7b77d3e9b2c1cbd4551cafa5b6049de206b4119778c6

Pith citing papers

No inbound Pith citation observations are available.