Pith. sign in

Paper Citation Record · LEDGER

VLANeXt: Recipes for Building Strong VLA Models

As of 20 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 7 inbound Pith citation observations for arXiv:2602.18532.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.18532 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-21T12:58:30.777235Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T11:57:40.853069Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T14:09:53.718432Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact34
  • verified fuzzy1
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b2529d4c-c5b3-4386-af60-45f20360ac71 · outbound

This paper cites Qwen3-VL Technical Report.

VLANeXt: Recipes for Building Strong VLA Models Qwen3-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.843841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:620f32cc98cb26c4ecb492640b9cdb77f18c90c5b15333042789ba362cd224d1

Observation 7a76bb13-8fbc-4c1f-a856-83b19701ba6b · outbound

This paper cites 3d cavla: Leveraging depth and 3d context to generalize vision language action models for unseen tasks.

VLANeXt: Recipes for Building Strong VLA Models 3d cavla: Leveraging depth and 3d context to generalize vision language action models for unseen tasks

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.938258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:2320065ecdd2faf42acba1d702d4e9034784df9e6753b956e9aa8a9b6b76b23e

Observation ea65ca17-e5fe-4943-9cd2-aad2b1be43dc · outbound

This paper cites Motus: A Unified Latent Action World Model.

VLANeXt: Recipes for Building Strong VLA Models Motus: A Unified Latent Action World Model

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.853621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:4ffa42430df923b4deaf4304b9e7d7c44759725afd2ad4bba0fc45598ccb1784

Observation 91a89579-96fe-4a66-bd99-578a2e942d9b · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

VLANeXt: Recipes for Building Strong VLA Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.829471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:768de3b8f42501c586ba488a760cc49e9a8fbe1335dc4cb58aea5bab914ee4b7

Observation 6a830836-c062-455f-9806-e218bc442433 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

VLANeXt: Recipes for Building Strong VLA Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.800641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:2493dc618046f3170775c1b2d0a121b094a7cd8869c71cb0d73a8943370ae4ff

Observation 3403115c-1040-4070-a9d6-984910f7010d · outbound

This paper cites RynnVLA-002: A Unified Vision-Language-Action and World Model.

VLANeXt: Recipes for Building Strong VLA Models RynnVLA-002: A Unified Vision-Language-Action and World Model

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:36.254972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:399238ea42f84273dad87cb1f427e3a7fe74e7e12b66f033e2c4e1fc9e089f64

Observation 07a9d8cc-c5b5-4d83-a626-80ea2d0ad0fa · outbound

This paper cites Combatvla: An efficient vision-language-action model for combat tasks in 3d action role-playing games.arXiv preprint arXiv:2503.09527, 2025a.

VLANeXt: Recipes for Building Strong VLA Models Combatvla: An efficient vision-language-action model for combat tasks in 3d action role-playing games.arXiv preprint arXiv:2503.09527, 2025a

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.917410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:c506d65e3fa4c672e52cf5e9dc9c5b836e647d7d8a3853e7dabb607d56c3cf0c

Observation 3a721f1f-b528-43cd-9b49-b910113cc8c1 · outbound

This paper cites Emu3.5: Native Multimodal Models are World Learners.

VLANeXt: Recipes for Building Strong VLA Models Emu3.5: Native Multimodal Models are World Learners

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.819893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:14e6731580419299f9efefa6f19552cf340fa3ecf76436e024e8d196f96fdb0f

Observation 2843b0aa-9727-4034-a106-5d3bff1f88ab · outbound

This paper cites Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration.

VLANeXt: Recipes for Building Strong VLA Models Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.825098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:8c3287ae82f129532e0cf6df821389dee2438fc072d0dd8571d96e969889a31e

Observation 8b3cf350-377f-4801-b242-20b4adf402eb · outbound

This paper cites Srpo: Self-referential policy optimization for vision-language-action models.

VLANeXt: Recipes for Building Strong VLA Models Srpo: Self-referential policy optimization for vision-language-action models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.943727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:34523e8aff0284efaca760e0ed5f84a66d634a93dc2df605d8c1ad8b3960ca44

Observation f82c2aad-f36e-4573-be43-5453fa722962 · outbound

This paper cites Vla-0: Building state-of-the-art vlas with zero modification.

VLANeXt: Recipes for Building Strong VLA Models Vla-0: Building state-of-the-art vlas with zero modification

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.948808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:c7587eeb87fbc66ec362d2cfb2ac5b9fa0854c75aacf89bfbb0e2bedf48f5534

Observation b207c2ce-82e7-45ea-878d-caf95d688861 · outbound

This paper cites The Llama 3 Herd of Models.

VLANeXt: Recipes for Building Strong VLA Models The Llama 3 Herd of Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.877521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:835198a83e503dc9a329b22ca0679f07f303df3af55aa0d65298a16f1c5a31df

Observation eaa56b9f-aad6-4800-98a0-4a53e04b47e6 · outbound

This paper cites Vla-reasoner: Empowering vision-language-action models with reasoning via online monte carlo tree search.

VLANeXt: Recipes for Building Strong VLA Models Vla-reasoner: Empowering vision-language-action models with reasoning via online monte carlo tree search

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.902948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:dc85dc04d2f8083051483c121b2a0cd96771f63e0dfbe69882a49715a049191a

Observation f0cc7abe-4ecd-4a84-8899-020d74d74d93 · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

VLANeXt: Recipes for Building Strong VLA Models Training Large Language Models to Reason in a Continuous Latent Space

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.872549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:3901159d481ef38209a34a31a16f41d6390d88f1febed304fd9aabb80b927dff

Observation 14e156da-bc7c-4d62-92c7-27c5c93c74ea · outbound

This paper cites ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning.

VLANeXt: Recipes for Building Strong VLA Models ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.922795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:615e0ca7a56283d0caf0d4652cfa27d8f275bb60d00f9feee3a7e60915bf2496

Observation 562aa024-58df-464a-823f-066859c496f7 · outbound

This paper cites NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks.

VLANeXt: Recipes for Building Strong VLA Models NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T13:00:09.958097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:ed3c1572b91034cef18bfc689f98df2da31f2820d079dbc25f300acffdcaea8c

Observation fb9fa478-13d0-4ec9-b9ca-ca2f7c7e2e85 · outbound

This paper cites $\pi^{*}_{0.6}$: a VLA That Learns From Experience.

VLANeXt: Recipes for Building Strong VLA Models $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.863649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:d2f395a3086d5d9375b26613d87b07b85cb3a4b1c6a6fba6dc5480adb5fd0019

Observation 7feaa7fa-d48d-45e3-9943-bdf0ef59ee9e · outbound

This paper cites Emergence of human to robot transfer in vision-language-action models.arXiv preprint arXiv:2512.22414.

VLANeXt: Recipes for Building Strong VLA Models Emergence of human to robot transfer in vision-language-action models.arXiv preprint arXiv:2512.22414

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.953953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:1dfbd8ca7f10d04d2598bf9e02c68b9444b1accac4eb4dcc47d3ede2d9daad0b

Observation a2995e19-b286-431b-8389-ed660c291555 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

VLANeXt: Recipes for Building Strong VLA Models Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T13:00:09.983615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:18b115721bd1a36ca1d8490e5b43435ca1676fb9d13f9172e5e94ab04a99a9a1

Observation 553266e8-9b48-40fd-9460-5071d629a452 · outbound

This paper cites Adapt Your Body: Mitigating Proprioception Shifts in Imitation Learning.

VLANeXt: Recipes for Building Strong VLA Models Adapt Your Body: Mitigating Proprioception Shifts in Imitation Learning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.927300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:665bc179c963ae4f195d1dfe383eb990a68eeb788b233545346fbbec2b7e66ab

Observation 8624a883-38a4-4a8c-a3dc-56efec558575 · outbound

This paper cites MolmoAct: Action Reasoning Models that can Reason in Space.

VLANeXt: Recipes for Building Strong VLA Models MolmoAct: Action Reasoning Models that can Reason in Space

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.933214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:55bbc8cf2c6fa5196195b1a5d128e79d0f3c941a44c3dd2e11fde6dfbb6296e8

Observation 340d6e3b-10d0-4dff-936f-89ffa78a0dec · outbound

This paper cites CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling, October 2025.

VLANeXt: Recipes for Building Strong VLA Models CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling, October 2025

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.788929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:bf0a1a5076da0586a1b18e4c912bcd30a38246ddacb5b78594cc52288495217a

Observation a004a340-b210-4fa3-a532-5fd319d6f2b5 · outbound

This paper cites Mm-act: Learn from multimodal parallel generation to act.arXiv preprint arXiv:2512.00975.

VLANeXt: Recipes for Building Strong VLA Models Mm-act: Learn from multimodal parallel generation to act.arXiv preprint arXiv:2512.00975

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.897990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:4c4197187f6a6d7d0ae1bd43c36d7a6b95ab64a2de445ad798ef12d030762abd

Observation bb374873-0a04-4ebc-b858-0cd97984f424 · outbound

This paper cites VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning.

VLANeXt: Recipes for Building Strong VLA Models VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.810733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:7a8266484c9cdf00218a539271bb20298538cd7c5bfd553e3e073e6d0e080c3f

Observation 10ea45f0-60d8-4f14-b5c3-0a5129e41dd5 · outbound

This paper cites F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions.

VLANeXt: Recipes for Building Strong VLA Models F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.839448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:d1b6a5b08e4b4ad9ff15789f04eb2b9c16112f9f89b0947f0b405fcd8b1a9e6b

Observation 81c2b160-50b1-46f6-adb6-5307249002d5 · outbound

This paper cites A Survey on Vision-Language-Action Models for Embodied AI.

VLANeXt: Recipes for Building Strong VLA Models A Survey on Vision-Language-Action Models for Embodied AI

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.966662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:01651f6b3ff624c3c50c871faab32c9e657487d67f685c5f62609ef164aa7291

Observation 7f42c612-968c-4900-abd2-b075f39f4f38 · outbound

This paper cites Transfer between Modalities with MetaQueries.

VLANeXt: Recipes for Building Strong VLA Models Transfer between Modalities with MetaQueries

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T13:00:09.815233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:2441eb800ed7e7ba2ccd7d266c140e7a004000a012f901137b4b44ca734e2201

Observation c2198c2b-e888-4d1f-b497-b45c68da367f · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

VLANeXt: Recipes for Building Strong VLA Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.805192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:e98544b77729a8e031ccf261df79cdc79d66b68b92ae55296ba925c31bc6a72e

Observation 5e275428-fc9b-4194-94da-00872201c09d · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

VLANeXt: Recipes for Building Strong VLA Models SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.974891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:f9780ff515a1bcc0ca712ac7c45a809cefc90160efdd26c3ee31a9d80ba0156c

Observation 3e623b8b-dfec-4782-bb0b-0e1bece95321 · outbound

This paper cites FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies.

VLANeXt: Recipes for Building Strong VLA Models FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.834452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:423690185b4e8519984598500928e48a1ac8590d02e97c4c98476ddb4685d174

Observation e4622c13-61b2-431f-a2fb-aa6d7df9ab69 · outbound

This paper cites MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation.

VLANeXt: Recipes for Building Strong VLA Models MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.882340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:fb7458311a7c87c54059261c44fac29f8544b2da0c776c03ae95d96ceb6d3e5a

Observation f83a9739-9b9f-4adf-ae36-042adf8aab82 · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

VLANeXt: Recipes for Building Strong VLA Models SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.892667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:ba900dc2d8ca828e006548605054640ba468a4f9391f34ccf968172e650e2a34

Observation fee226ad-4cf7-4ec4-8e05-8746dc4aae4d · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

VLANeXt: Recipes for Building Strong VLA Models Gemini Robotics: Bringing AI into the Physical World

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T13:00:09.848903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:cefb642112ba25f51f6e1c7db8fc875fe652492ad7b8fe329d496085659d45d0

Observation bb42da93-9506-4e30-bc3c-1fa1ea1ba714 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

VLANeXt: Recipes for Building Strong VLA Models SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T13:00:09.970523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:fa35d111f87928b9e76e27acd2a3ea8705d8d7c535c93f53aa0b493ccea535fb

Observation 2d520977-76dd-4e28-b380-d883e3d437b3 · outbound

This paper cites End-to-end Listen, Look, Speak and Act.

VLANeXt: Recipes for Building Strong VLA Models End-to-end Listen, Look, Speak and Act

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.979464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:d5a1ebb64d7fa038ae10aacbad17480c03e3a62bd63172ec18e470d8f9f5156f

Observation 71c88fad-a9b5-4a7f-b2be-7ada6c9808c1 · outbound

This paper cites World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training.

VLANeXt: Recipes for Building Strong VLA Models World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.868211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:5d32d6013c15710f9c0e24a2ac8949990383270d46bc678ef6a5733d491b56ea

Observation b231477a-b9f9-488e-a402-c4690cd96367 · outbound

This paper cites 4d-vla: Spatiotemporal vision-language-action pretraining with cross-scene calibration.arXiv preprint arXiv:2506.22242, 2025a.

VLANeXt: Recipes for Building Strong VLA Models 4d-vla: Spatiotemporal vision-language-action pretraining with cross-scene calibration.arXiv preprint arXiv:2506.22242, 2025a

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.795112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:72a3bd5b3ae335f5685f0630ac92263a4049c4cffbfa17d762dd62da20504531

Observation bf431467-507b-4f2d-8d10-0212c85a8db3 · outbound

This paper cites Dreamvla: a vision-language-action model dreamed with comprehen- sive world knowledge.

VLANeXt: Recipes for Building Strong VLA Models Dreamvla: a vision-language-action model dreamed with comprehen- sive world knowledge

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.962649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:e4a00008c98d32dbdb08a6b1500e35adad93e59e7d1bea0d2f3ece3aacfc5f95

Observation 563fcdf1-8af8-4374-a347-18814a6f7153 · outbound

This paper cites Flowvla: Visual chain of thought-based motion reason- ing for vision-language-action models.arXiv preprint arXiv:2508.18269.

VLANeXt: Recipes for Building Strong VLA Models Flowvla: Visual chain of thought-based motion reason- ing for vision-language-action models.arXiv preprint arXiv:2508.18269

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:00:09.858979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:7c34367e975dcb10503fadbb50dd80b2c5e1f3cc808bc45d1a094d39185b472f

Observation 65c5a765-12b4-45d2-82a8-53d1723e8c90 · outbound

This paper cites More Experimental Results A.1.

VLANeXt: Recipes for Building Strong VLA Models More Experimental Results A.1

Reference 40

Resolution
malformed identifier
raw_fallback, observed 2026-05-21T13:00:10.452298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:c11c3a990180c216a576857233a05079e4151837d6e590d4c71d97d13a953c65

Observation bc88e369-620b-4fa8-a040-c67ca39c56a0 · outbound

This paper cites primordial soup.

VLANeXt: Recipes for Building Strong VLA Models primordial soup

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:00:10.455299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:1c5e1cfcb81710019afd190039d14532d3436f596f66ca6b72d1e1db3b886960

Pith citing papers

Observation c136c42d-0514-483d-8f53-d9fc1928c427 · inbound

VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts cites this paper.

VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts VLANeXt: Recipes for Building Strong VLA Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:04:11.439936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-08T09:11:21.715023Z digest=sha256:a05c9ef71b1b377b85dd6a5d9bf382be97deaf1931f2d239f224363ff95b3673

Observation 54fbf5b6-c48d-4a25-815b-fedfe200d0ce · inbound

VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts cites this paper.

VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts VLANeXt: Recipes for Building Strong VLA Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:04:11.439936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T01:05:07.509291Z digest=sha256:13b601f896b631a61b6310ac7db4630d30d4c904aef2e9fb788ec90ee16fa2bb

Observation 102eb8f3-cd5d-49f3-a700-c087586b9b1f · inbound

Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR cites this paper.

Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR VLANeXt: Recipes for Building Strong VLA Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:04:11.439936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T07:14:31.613251Z digest=sha256:3c40dd70a930b618db6646be7e7830e4d44a6ed3c97776253bf7065cd341e23b

Observation 20be91d9-bff0-49eb-a1b7-848a0b41da9a · inbound

Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring cites this paper.

Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring VLANeXt: Recipes for Building Strong VLA Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-28T22:42:46.568474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:39:15.211339Z digest=sha256:0adc766dd70be90ccfd1ed16b5ddf8716cdc7ecaeeef1a2412a66b226e9f200b

Observation 94d4f7ab-fae3-4fa6-b308-181e2d207902 · inbound

PearlVLA: Progressive Embodied Action-Plan Refinement in Latent Space cites this paper.

PearlVLA: Progressive Embodied Action-Plan Refinement in Latent Space VLANeXt: Recipes for Building Strong VLA Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:58:58.543012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T00:57:58.312567Z digest=sha256:74fc1e645e57eb1fc18bc65b5482f5e308d67ca172cd6491f3280f401ea102f0

Observation 5e853214-c132-47ed-87d7-1375423f15f7 · inbound

RouterVLA: Turning Smoke Tests into Supervision for Heterogeneous VLA Selection cites this paper.

RouterVLA: Turning Smoke Tests into Supervision for Heterogeneous VLA Selection VLANeXt: Recipes for Building Strong VLA Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-04T14:09:53.719676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T04:25:18.870217Z digest=sha256:549da4cd58507ed596b2c60ab5f7202fe27f9ae719d4835dc7dbec637d1c2f7c

Observation f0f812ab-e057-4896-a429-b6a98d1b7a4c · inbound

Look Where It Matters: Adaptive Visual Refinement for Vision-Language-Action Models cites this paper.

Look Where It Matters: Adaptive Visual Refinement for Vision-Language-Action Models VLANeXt: Recipes for Building Strong VLA Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T11:57:40.853069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:57:40.853069Z digest=sha256:58c92206a57a4820ffca9ca06eb5850903f05892be86447285cc624459986169