Pith. sign in

Paper Citation Record · LEDGER

LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2406.11815.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.11815 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:32:17.452155Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T15:18:33.791076Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0637749e-4422-4456-9a3f-32347f4afa38 · inbound

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs cites this paper.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.634324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.634324Z digest=sha256:80fa667599df6975dc514d6dc9c80c06a60b636107edf83bcc37ba4fd581cbaa

Observation 51d6c4cf-99bf-44fa-981d-a445625cb05c · inbound

TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies cites this paper.

TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:27:22.981258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-15T18:27:22.760982Z digest=sha256:0740df6de580e0b8e6b8074b701b3f6b426a129407fcb899193e9500f60e6aef

Observation 4b6c3e5e-0f5b-46d8-b868-0a212ac9f727 · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 186

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.619344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.619344Z digest=sha256:7ec824e95aff51e23645e60d77799d492bfcf92cf77ead1cf6a9d80cd5a98094

Observation 660aedeb-4eb8-470a-9b07-344a75dbed69 · inbound

Robotic Programmer: Video Instructed Policy Code Generation for Robotic Manipulation cites this paper.

Robotic Programmer: Video Instructed Policy Code Generation for Robotic Manipulation LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:41:29.028549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:41:29.028549Z digest=sha256:36f24236bb30f762a0c6d519a67074721269191e1bc23f7c724043f79789f4eb

Observation 5dd60720-343e-46a7-b758-2be01aacc80d · inbound

Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs cites this paper.

Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T11:32:17.452155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:32:17.452155Z digest=sha256:3b9541fb1eed21bf918841d717a4b191822b0abfc59c9404f6573e06a4728cc7

Observation c6264bba-281b-41fd-b8b8-bb7f49518a39 · inbound

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization cites this paper.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.919649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:4bf1c6c936640de14be198328979f60f8b6c8164ddfad0d80fc83c05a6c86624

Observation 8c9a8876-44f2-4354-8a3f-7e9fc893b05f · inbound

RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration cites this paper.

RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:35.317604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:35.317604Z digest=sha256:0445fc5c7011629c434aae175b91adfecf575db52d758a7fb91f4f5a31aad711

Observation 013774cc-b1e7-436a-9cc5-ed7e4d8a81a3 · inbound

Pixel Motion as Universal Representation for Robot Control cites this paper.

Pixel Motion as Universal Representation for Robot Control LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T22:12:54.984719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:12:54.984719Z digest=sha256:f6280b52a8da90ccc8ab0b20ab2692365f16675fbfb4235620635ba4b18d8d84

Observation c2df2264-8199-4720-ba16-74c35d79db47 · inbound

ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models cites this paper.

ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:24.077240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:24.077240Z digest=sha256:7f4ac270cf57f23066b74a9be94abb2a25295db0f8bff7425ccad819768c1b06

Observation 9e66d2b2-3332-410e-b1d7-9f79c874c03c · inbound

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization cites this paper.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:16.013005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:16.013005Z digest=sha256:e44ba2cd39830a43ca023364177fe86a5933fe5d31903993165143785854bbea

Observation 4d5d0a9d-c5bd-49e2-a846-cf6f124625f9 · inbound

CEED-VLA: Consistency Vision-Language-Action Model with Early-Exit Decoding cites this paper.

CEED-VLA: Consistency Vision-Language-Action Model with Early-Exit Decoding LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:03:56.636770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:03:56.636770Z digest=sha256:cdbbd0b853466bbfe153b1316fd2f6b24b7ce6b5383032ce6b17b2a68b553f9b

Observation 0f393c74-1e1a-4441-9de2-6e647414a394 · inbound

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning cites this paper.

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.021432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T03:18:14.655384Z digest=sha256:9545ea7e52441b3b9e4ef287329b40ab17c8823d4fb72f3045f7b8bb65b40634

Observation fe60effd-845e-43db-8772-d8a66a868c31 · inbound

RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping cites this paper.

RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T10:32:25.140850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:32:25.140850Z digest=sha256:a6e95700187f2237ad6cdf4725fd50f919aaac3e33ec823ac382f96788ffc7b8

Observation dc6d773c-9e4a-4678-94a5-6d73163ec34e · inbound

ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver cites this paper.

ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T20:32:45.390488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:32:45.390488Z digest=sha256:012853c76bfff95161847ca20d03b047ebfc2b14f5ba6236a348f65be1d70048

Observation c674e204-38c9-4720-9e81-7012a735d468 · inbound

Ego-centric Predictive Model Conditioned on Hand Trajectories cites this paper.

Ego-centric Predictive Model Conditioned on Hand Trajectories LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T15:29:25.505592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:29:25.505592Z digest=sha256:2b20d3569aea1c648be38a862c04803ab9823134a487f6eb4984c893a5b44837

Observation c72e8b14-500b-43d8-b202-a7f9555e8d68 · inbound

Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics cites this paper.

Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-03T16:27:34.062828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:27:34.062828Z digest=sha256:dd2e49b2f05c33274c1fcb16d0db892ed1abf878bcbfe9ce3a841bb1733f8676

Observation 6beaaddc-ce38-4903-9f5a-de8fff1354f4 · inbound

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation cites this paper.

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:08:01.994473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T15:05:21.907878Z digest=sha256:2e3aeb9a0c7700f702ff1bb4b5ce9418e5ffcdc48b359abf49bf8944a00c130e

Observation e2dbe256-8bf9-4d7e-8486-d813c3031b00 · inbound

Decompose and Recompose: Reasoning New Skills from Existing Abilities for Cross-Task Robotic Manipulation cites this paper.

Decompose and Recompose: Reasoning New Skills from Existing Abilities for Cross-Task Robotic Manipulation LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:56:06.795208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-09T14:30:00.890117Z digest=sha256:dd3bc64cd04399605eff553eaf2a746456c7a0b26ef6400afc8e63ec176875c6

Observation 451547fd-6822-49fb-b9c4-ff8ecbc68ca4 · inbound

Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning cites this paper.

Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.630877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T07:47:52.739735Z digest=sha256:38e1dd3187e2c3a6f855186ba0cd8bab7548306305b4704f3bfba0267d15875d

Observation e064ffb5-c8a1-4853-8dcd-e25a536c8d55 · inbound

Grasp-Then-Plan with Failure Attribution: A Closed Two-Stage Framework for Precise and Generalizable Robotic Manipulation cites this paper.

Grasp-Then-Plan with Failure Attribution: A Closed Two-Stage Framework for Precise and Generalizable Robotic Manipulation LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:56:34.473901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T09:37:58.434897Z digest=sha256:47feffecb972a5760ddcf349c4afa85dceafe4c675fed326c2e1109dbf535631

Observation 53be3b28-1bf8-4a88-9266-7cde92c907ed · inbound

Trajectory-Level Redirection Attacks on Vision-Language-Action Models cites this paper.

Trajectory-Level Redirection Attacks on Vision-Language-Action Models LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:33.793178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T06:33:53.013076Z digest=sha256:8a55aa56f92cf87ec7ddd8723b547f38d306018a1681a9bb9c7790fa58aea674

Observation e21013a9-57c6-4ceb-8be4-35d30bdc961b · inbound

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation cites this paper.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:42.768376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:42.768376Z digest=sha256:62da659b79114a3693c355a4d06ae9723cf2ef84275ad7926d219dbced58afc4

Observation 5335d7da-d38d-4ad2-b593-3db0b03e1cd0 · inbound

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation cites this paper.

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T14:39:40.274531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:39:40.274531Z digest=sha256:809fc757d3d611624fc58022567378484076c941a9bc80fc7b0665b9b78f3dc2