Pith. sign in

Paper Citation Record · LEDGER

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions

As of 20 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 3 inbound Pith citation observations for arXiv:2505.11214.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11214 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:01:28.461260Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:56:22.911558Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T20:28:16.078786Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c2d23df2-6129-4393-ac16-292debc961be · outbound

This paper cites Training language models to follow instructions with human feedback.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Training language models to follow instructions with human feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:28.286575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:01:28.286575Z digest=sha256:4aecd4333e20d7289bde24e981c382f9c87e648b5ab7a7a97ba177db957d5c00

Observation 86cc614d-4ae6-4aea-a45b-8b0db30e8e83 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions LLaMA: Open and Efficient Foundation Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:28.292368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:01:28.292368Z digest=sha256:b1ab48d6e5f9b514d4a0ef1739b1f8a6e83a62041ba75c57d9fbf0e25268d640

Observation 9837886e-a052-44b9-884f-cfd8efcfc023 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Gemini: A Family of Highly Capable Multimodal Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:28.297376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:01:28.297376Z digest=sha256:b7fbd26d5d15579c50c454380a28e0f3338cd88047e97038d2c75d6d746d4c4e

Observation 51b3bb3a-32b5-4421-9783-798e0704ef5a · outbound

This paper cites Visual Instruction Tuning.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Visual Instruction Tuning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:29.145766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T21:01:28.303202Z digest=sha256:ea3bb8f94dab4aed358d023fb6d82b5535af057ec84a26f9388234b3e87eda35

Observation 75a71d25-1387-4dfc-9d95-8b5bb733ce80 · outbound

This paper cites BLIP -2: Bootstrapping Language - Image Pre -training with Frozen Image Encoders and Large Language Models.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions BLIP -2: Bootstrapping Language - Image Pre -training with Frozen Image Encoders and Large Language Models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:29.130979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T21:01:28.308402Z digest=sha256:bd160daa8a920d5d8f3ceadfb03de8fefebd9a267f19b62be13aed3f66f3e7d4

Observation 67836b1e-f353-4f6f-b6ee-27d5906336a7 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:28.313787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:01:28.313787Z digest=sha256:c19524918938457790cd362e854ef1cef8910aec9a41ff5b69fa1f14fe00a646

Observation 895c3502-6c74-4c1e-96ed-243edc68a864 · outbound

This paper cites Prismatic VLMs : Investigating the Design Space of Visually - Conditioned Language Models.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Prismatic VLMs : Investigating the Design Space of Visually - Conditioned Language Models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:29.115683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T21:01:28.319334Z digest=sha256:19740d7cceb1709bb8837d188cbb92c18ba97b812b87eab95c626efb9d78299e

Observation 5f536b50-b7c1-4291-a734-6928c560abe1 · outbound

This paper cites LL a MA -adapter: Efficient fine-tuning of large language models with zero-initialized attention.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions LL a MA -adapter: Efficient fine-tuning of large language models with zero-initialized attention

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:29.099786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T21:01:28.323996Z digest=sha256:028c2e4081ed6a7b87672e51d6bc5081da6c136e078bb63a15f6ea99979b2b05

Observation a2129233-259e-4803-8872-3cfcd773837a · outbound

This paper cites Sanketi, Grecia Salazar, Michael S.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Sanketi, Grecia Salazar, Michael S

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:29.084252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T21:01:28.328677Z digest=sha256:eb60636541e6264c46dc2a07b3a98451fecec9ec8d67c2f037124088816d963e

Observation b2297c5f-afdc-4902-97ae-1f19e6aa7f4c · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:28.334351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:01:28.334351Z digest=sha256:0f4547c45cc07bb2583b0ebb88a778f18486cf25ae6388c6a82916725d3e0bb0

Observation 62778bcf-76fa-4b56-b216-7f24a0b45018 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions OpenVLA: An Open-Source Vision-Language-Action Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:28.339447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:01:28.339447Z digest=sha256:bd4e7f4e6b3e9d3f1abc9591e2a4f4f0fd83c911bca15523e10ef71a4d6cc248

Observation 42e2d3e5-449b-4db8-835b-5c6554cb432c · outbound

This paper cites RT -1: Robotics Transformer for Real - World Control at Scale.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions RT -1: Robotics Transformer for Real - World Control at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:28.344595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:01:28.344595Z digest=sha256:8afedcda9e3408668cefaea8dcfad52c864faf47fbb29698f6cfded0f6bdd833

Observation d59a5527-a746-42fc-a6ec-037925766b08 · outbound

This paper cites Language Conditioned Imitation Learning Over Unstructured Data.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Language Conditioned Imitation Learning Over Unstructured Data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:28.348934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:01:28.348934Z digest=sha256:2d61aa7ee39557213960f7362309a3e1c62e3be3a3811ee4bc4df85e9619c5d9

Observation f9e5cd03-7cfa-4a63-bf53-636ac935f2f2 · outbound

This paper cites Vision- Language Foundation Models as Effective Robot Imitators.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Vision- Language Foundation Models as Effective Robot Imitators

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:29.058706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T21:01:28.353405Z digest=sha256:40920c7f1519d12333e3d70857ee2e16045396998084721883e58a51be06709f

Observation 74e367ab-0ddd-42bc-a9ba-b650cf217d50 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:28.357831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:01:28.357831Z digest=sha256:a4d5c5252c20fc13edeb26160ec2a2bcdce482c9c3fbc04e795722481a9ddbe2

Observation 35913785-6f73-4528-b94a-e1c861a9f75f · outbound

This paper cites RDT -1b: a diffusion foundation model for bimanual manipulation.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions RDT -1b: a diffusion foundation model for bimanual manipulation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:29.042062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T21:01:28.362736Z digest=sha256:8cda3075bb66dd8d2e7aae89b803e0673561715fa13ca770cafb52c3285dbd01

Observation 6d0cbb99-6dac-4f1f-a239-d73381e74d3c · outbound

This paper cites What Matters in Building Vision-Language-Action Models for Generalist Robots.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:28.367229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:01:28.367229Z digest=sha256:2852f3874a70cbe35ec36c050c904e1992b9053cbd7deb7a05edde89587d0cf5

Observation 0fee3452-f08e-4d51-9183-496992f53f19 · outbound

This paper cites TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:28.371944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:01:28.371944Z digest=sha256:737062ad31d3f6ee3432c9c977358171bf8a92f94148da5e9850d0d942b87919

Observation 21a7a0a9-4637-46e2-9892-ee2983bb4cba · outbound

This paper cites Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:28.377189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:01:28.377189Z digest=sha256:ad742b3dedba7a2213adc98ee06d0f3b53c46ffea5d727366fc8efd139f818cd

Observation 705c1280-a39b-4398-a971-5b54c75c5ba9 · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:29.025582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T21:01:28.382672Z digest=sha256:c67169ddc0b9cb8d1eadffb80ba680a2a7513bbb03558fa927ba3725bb754f45

Observation 663d3148-ceb7-48e1-8f76-4a48bfcdec63 · outbound

This paper cites Diffusion Policy : Visuomotor Policy Learning via Action Diffusion.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Diffusion Policy : Visuomotor Policy Learning via Action Diffusion

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:28.387297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:01:28.387297Z digest=sha256:be0fdcf2d005d83be24aaecb125c396f068c377ae0076277bcd323ffbf1715ba

Observation 91034c97-c35a-454e-82bb-a6b3caef9dd0 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:28.392029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:01:28.392029Z digest=sha256:d9a5e6e54bcb8a83dad81decd35d41793dc3de06165104a4370bb4a3e95cd339

Observation e38d7c17-ca60-473a-ab00-33601141c25b · outbound

This paper cites 3D - VLA : A 3D Vision - Language - Action Generative World Model.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions 3D - VLA : A 3D Vision - Language - Action Generative World Model

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:29.008180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T21:01:28.396975Z digest=sha256:64919877917a451d5bf4735635346bc2bd9a12c28f8c93167f2ffe240e393c98

Observation 504d44d2-dd21-4996-ad3c-eb1fd027f004 · outbound

This paper cites Dream to manipulate: Compositional world models empowering robot imitation learning with imagination.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Dream to manipulate: Compositional world models empowering robot imitation learning with imagination

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:28.992785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T21:01:28.401842Z digest=sha256:664327b029ad782800c43aacbfa64ed6a4649c83a02d86a47cbb5e1ee670d907

Observation d92e8050-a612-454c-88e3-0b9330c8fa52 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:28.406539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:01:28.406539Z digest=sha256:d7b43ef35ad1bac3071692532eae23db788b0d6f6e36c1534b8fa3a1d53aab93

Observation 140de4da-d8d0-4176-9040-ba8267deab69 · outbound

This paper cites RT-H: Action Hierarchies Using Language.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions RT-H: Action Hierarchies Using Language

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:28.411354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:01:28.411354Z digest=sha256:e83efe25dc3e044f0977fab56d4e5e01a9e2b6330520fa49be079ff060c6e3d0

Observation 6f229904-50ff-4ba9-a486-c6d94f54e345 · outbound

This paper cites Robotic control via embodied chain-of-thought reasoning.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Robotic control via embodied chain-of-thought reasoning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:28.976982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T21:01:28.416183Z digest=sha256:349691bb0d8531f2edd33a86509af5066ac2a07c7bc7a90709c2f8e9b8ff47ff

Observation 02fe49ba-6f9b-4a6c-a2af-a7bb1b3a62d2 · outbound

This paper cites DayDreamer : World Models for Physical Robot Learning.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions DayDreamer : World Models for Physical Robot Learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:28.961902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T21:01:28.421139Z digest=sha256:894a48670321088450c39185b7bbce67d5db95aab1fa15e5ee2fbded73452143

Observation 537c13e6-b982-4f62-b34a-a0fd35f8243b · outbound

This paper cites VIMA : Robot Manipulation with Multimodal Prompts.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions VIMA : Robot Manipulation with Multimodal Prompts

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:28.946233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T21:01:28.425661Z digest=sha256:a8de3b66f855a2f27efe511ea60ee2846d2e90c814e30f14323124b874572d78

Observation 7e67f5d9-f7d3-4871-a214-97dd0b4c2fc2 · outbound

This paper cites Interleave- VLA : Enhancing Robot Manipulation with Interleaved Image - Text Instructions , May 2025.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Interleave- VLA : Enhancing Robot Manipulation with Interleaved Image - Text Instructions , May 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:28.430149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:01:28.430149Z digest=sha256:f8d37b46a9feef869ce4984c1c47258a104f967899c7331af8d21c8d1d2515e8

Observation e1d33075-eab2-4d16-8742-9deac8a80318 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:28.435690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:01:28.435690Z digest=sha256:e15211ad374b64be77f27610c1d1035483382b10b40a0d95fee23cbe1d7a3024

Observation 464a12f3-9d26-428f-8762-c3b071891814 · outbound

This paper cites Sigmoid Loss for Language Image Pre - Training.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Sigmoid Loss for Language Image Pre - Training

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:28.440856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:01:28.440856Z digest=sha256:900b67c0e837c402da812f0fc04d3710790dd5c30c11bf5dd4df2d63c9b19f7f

Observation f6fb430f-7cde-4909-b37e-648acbc336c9 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions An image is worth 16x16 words: Transformers for image recognition at scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:28.445519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:01:28.445519Z digest=sha256:f3b362fad3694d6bf13e7670bee9789a98feac8fa62f0032df309e44ef5d82c5

Observation 2355a54a-c018-4b16-8638-10958b256fa3 · outbound

This paper cites Qwen Technical Report.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Qwen Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:28.450294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:01:28.450294Z digest=sha256:b6ded92815a81613337fe58a5f574d656d19fb955699aa5f2637c83253178bd5

Observation 0e364c4d-9564-4082-bdab-640e6b79ba5c · outbound

This paper cites Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:28.455317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:01:28.455317Z digest=sha256:9e673a9015803f76bd307e9d7b0274182c8d52af4d6b0ee8f55a3fda554cd6b3

Observation d05ca3a6-9c4e-41f1-b1f9-5f3e2a272c95 · outbound

This paper cites CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:28.461260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:01:28.461260Z digest=sha256:ab340e6691b7b5052477af509fa348817d065fb68c5e166102af185707e44895

Pith citing papers

Observation a2730981-4504-4ffa-9c78-9c5e3e9f836d · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:28:16.081523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:fbf74dc5ad3afebc58a5ac42b35c41185d3ba0ea5cca19f38163d9f3f4bc9580

Observation e9948e22-77e3-4aba-b0a1-d6a0661eb9dd · inbound

CLAW: A Vision-Language-Action Framework for Weight-Aware Robotic Grasping cites this paper.

CLAW: A Vision-Language-Action Framework for Weight-Aware Robotic Grasping Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T15:56:22.911558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:56:22.911558Z digest=sha256:06b8a076c6a3673ece8c97615ebeaac36824e8e02093e4dbf0da1c0c11161669

Observation 904c3149-d11a-41bc-9d72-902213218bb8 · inbound

UAOR: Uncertainty-aware Observation Reinjection for Vision-Language-Action Models cites this paper.

UAOR: Uncertainty-aware Observation Reinjection for Vision-Language-Action Models Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T22:07:48.479937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:07:48.479937Z digest=sha256:3121e0f5371733240b6ff4359b3d174db9cbf2b6a99f001e2705361b76329652