Pith. sign in

Paper Citation Record · LEDGER

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies

As of 23 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 1 inbound Pith citation observation for arXiv:2605.27284.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.27284 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T17:19:04.477325Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T14:13:43.160700Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact25
  • verified fuzzy0
  • unresolved27
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 77170b98-4ed4-43b6-906b-8cbedf936175 · outbound

This paper cites ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:23:44.921783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:dfc4126c07d6b9f2886d07d8f704c95972e740ee116af6a65682673339699bf1

Observation 10a67a13-bdf8-445b-b54a-24e0e0b28c21 · outbound

This paper cites A Pragmatic VLA Foundation Model.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies A Pragmatic VLA Foundation Model

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:23:44.921456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:ea729ee234011bb42d39a0fa2c8b9ece75ff12e341efcea1bcdc56b930817da2

Observation fac17dfd-4895-4af0-b7e7-986465c4ffea · outbound

This paper cites Nvidia isaac gr00t.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Nvidia isaac gr00t

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:d248968cf9d1bba0cd5cbb58d4366e96f6e1cdb7a2fc4334e32cda85a5563914

Observation be5c5a54-86e7-4d41-9f8b-7772ffd671ae · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:5dfcc3d854bd0d63c1a2a6f8e2f04ed75404c32e4f5f829c9ebe3b1b85f9671d

Observation 3a975193-7074-45d3-8558-e5f2fbc37bd3 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:23:44.919185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:ec819a5e0c62e46cd4f3caa5e9851d2ded8a19d7c4e92c1b4f73497dc8c9b053

Observation 48be4c0f-c411-4ecf-af9c-39d32773f460 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:23:44.926072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:0c39aab804201d2f8069707bec80d9fd7ab8939db4d3cca852136136ebe105fd

Observation a37e55ad-f482-454b-93ff-37e929cbe22d · outbound

This paper cites RoboVQA: Multimodal Long-Horizon Reasoning for Robotics.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies RoboVQA: Multimodal Long-Horizon Reasoning for Robotics

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T17:23:44.942418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:fb5b732b87fd5e7ba76cb59577dedbfcfe7194a756e6035b6b4eae242b3e5f1c

Observation e405e4fd-0188-42b7-9f12-26f5a907802a · outbound

This paper cites Robobench: A comprehensive evaluation benchmark for multimodal large language models as embodied brain, 2025.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Robobench: A comprehensive evaluation benchmark for multimodal large language models as embodied brain, 2025

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:eae72c5fe950907d18610bcdbc86df3ad653c396cb70cfbe0b07be6178bf9c61

Observation 5e99ec83-6d05-43fd-97e4-d30d0a10d217 · outbound

This paper cites Handyvqa: A video qa benchmark for fine-grained hand-object interaction dynamics.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Handyvqa: A video qa benchmark for fine-grained hand-object interaction dynamics

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:23:44.950904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:1c9aafcb543aa418a0e95ea94ecb903aeba60b4f5f2708bb49e4f7331e806394

Observation 51a173f8-9cc3-48b2-acce-31f5e3587a7b · outbound

This paper cites Qwen3.5: Towards native multimodal agents, February 2026.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Qwen3.5: Towards native multimodal agents, February 2026

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:90cd28aa221812f8ffa171b4fa333cf19ede1f250c03995cca6fa467750ca375

Observation 2427e160-c65a-4a94-a6b3-205544a521bd · outbound

This paper cites BridgeData V2: A Dataset for Robot Learning at Scale.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies BridgeData V2: A Dataset for Robot Learning at Scale

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:23:44.954034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:764a608147c455dfb648bc4987ec52f5e423f3a1ac466b21464fa63f6e390f94

Observation bcc11873-90cf-4d29-8f84-5bf953f107aa · outbound

This paper cites Bc-z: Zero-shot task generalization with robotic imitation learning,.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Bc-z: Zero-shot task generalization with robotic imitation learning,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:04ff02fc72fd542557340d27dc546be986671fcb2fbd254da2fa64c0d2414112

Observation c240c5ab-de63-4096-bbfd-754a48dcc04b · outbound

This paper cites BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:23:44.943146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:c1ce1df61d95d82929c0cf7c9f400f5120d233faff7c59dbc0f82b31946c82d8

Observation 03f9d295-e0bd-466e-a141-a0ffbb246c38 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies RT-1: Robotics Transformer for Real-World Control at Scale

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:23:44.932205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:c11d0a12197e5e6aecad483593bad8bb3bb4f8acab072a7cc40305c017ba331d

Observation 9211b989-36e3-452d-ac47-febbf02a7331 · outbound

This paper cites Galaxea Open-World Dataset and G0 Dual-System VLA Model.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Galaxea Open-World Dataset and G0 Dual-System VLA Model

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:23:44.940463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:b16bcfe699f577f7ce6e62ef7ceaf75e0e9030ef999535c0f68fd081e47ba47c

Observation 87882627-16bc-4000-8218-d2bb3983579d · outbound

This paper cites URL http://dx.doi.org/10.15607/RSS.2025.XXI.152.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies URL http://dx.doi.org/10.15607/RSS.2025.XXI.152

Reference 16

Resolution
metadata mismatch
doi, observed 2026-06-29T17:23:44.401136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:dec0b27053003d84f83d4c1ab48152ca9f676ad38fec1bb8fcab6ba2c63640e1

Observation 9e89b392-c8c2-444d-8ec5-c6607d7046c3 · outbound

This paper cites Robomind 2.0: A multimodal, bimanual mobile manipulation dataset for generalizable embodied intelligence.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Robomind 2.0: A multimodal, bimanual mobile manipulation dataset for generalizable embodied intelligence

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:23:44.947959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:74f546b7c3f7177471bfd6c3df0e6c9bb1b0b7f9e5b8970f25285b61881157f8

Observation 4848680e-69f9-4d71-8f6d-93e49743127d · outbound

This paper cites RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:23:44.911537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:22b6c575e6279e6da04c773988f8b139af44146b1666eae6b2dbf09d231cf847

Observation 28488e8b-e62f-4778-a6c2-45c888654db1 · outbound

This paper cites Rh20t: A comprehensive robotic dataset for learning diverse skills in one-shot,.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Rh20t: A comprehensive robotic dataset for learning diverse skills in one-shot,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:56f4abb4910e4d6f8789ef5377dfa27e702446257685ab56ab4e79f66ce1ebfc

Observation 5b6b1bf4-fe99-4fe9-8c7b-01841233f225 · outbound

This paper cites RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:23:44.916841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:5a82ed16b413c5b7a125d9477df876c7128e022236858da8cf01d3b015c0cb61

Observation 85361763-a493-469c-b7a3-acf63262d240 · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:23:44.892194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:df3d542ffa80c9a7b4388047325f49085346f7f13df41835b16f63ee44b22e77

Observation 85ba6212-81ac-442b-ae70-be8981b9601d · outbound

This paper cites StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:23:44.894826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:ec4a5c38762ad3268d51c89b047bc3d77b7a4b2f3c23e7e74339153d5fdc5555

Observation 3a404a20-ecdf-45ac-b582-ef18514d00c4 · outbound

This paper cites RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins (early version).

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins (early version)

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:23:44.897400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:b0dfeb5401c646e6bfb692027aa51d5677bdc6f76a607dd32f7bb64bd4051d42

Observation 62567983-9e3c-4c23-ac22-6cb251a99e45 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:23:44.913881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:333048874b923f0ef99bcf1548d6e258cf280d5b592d3251cc947a7696431893

Observation 2275bb3f-fa75-4b62-b5f1-b15249c9b085 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies OpenVLA: An Open-Source Vision-Language-Action Model

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:23:44.904879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:27a1ba5b68107bc24c39cb27a3ee522d998cab8e2e0e882d74b99fe2b720e1c8

Observation 733a7607-0446-4682-a0af-7c6f6f7f99d1 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:23:44.923746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:205ec2adad99f37ef6f5381f5f78f6f9c5521ea2a5aeceac7cf8b48b91b97942

Observation d6ff64a6-a339-428d-9bee-e2f39b9a11af · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Octo: An Open-Source Generalist Robot Policy

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:23:44.919434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:916c6024c291ea227786d9cbfbe4e2cc6b87aa07b7277dbb804e2fcf0e7879b2

Observation 5af2092c-fc3d-49f7-b280-3d8679e7a9bc · outbound

This paper cites RoboInter: A holistic intermediate representation suite towards robotic manipulation.arXiv preprint arXiv:2602.09973, 2026.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies RoboInter: A holistic intermediate representation suite towards robotic manipulation.arXiv preprint arXiv:2602.09973, 2026

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:23:44.953547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:73aa52ae83660e44626240914057265b4cb0023385c64958aabf308f97b8ab9f

Observation 5a7e1444-131a-4cc0-96ec-56dbdc4576f7 · outbound

This paper cites STEER: Flexible Robotic Manipulation via Dense Language Grounding.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies STEER: Flexible Robotic Manipulation via Dense Language Grounding

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:23:44.937340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:810b85a4010fca20f211d574752df9eb342842ba070a88f86f16da8d16f9bf4a

Observation 9eed3036-5fa9-433b-a62a-17c7bddfe3eb · outbound

This paper cites PartInstruct: Part-level Instruction Following for Fine-grained Robot Manipulation.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies PartInstruct: Part-level Instruction Following for Fine-grained Robot Manipulation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:23:44.946075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:a73f999fd7e5fcaf1110c2cc7a77b7f297481e6f467670454cd042817b48cd81

Observation 87a4314c-5c8e-4816-9c26-e207e941d957 · outbound

This paper cites Qwen3-VL Technical Report.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Qwen3-VL Technical Report

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:23:44.955913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:316ea85a9834ea593465477fd572126d9a723be942740fa6c0cca30d09f444e0

Observation 2b405af1-95bc-4f6b-861c-341cc5346d29 · outbound

This paper cites Qwen3.5-Omni technical report, 2026.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Qwen3.5-Omni technical report, 2026

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:3117dbf127ea77bbbabb914b0b95f286f299d8a62cad16a13325c2d60c6af9ef

Observation 9f0d7d01-72f0-4063-90af-4daa7239d4d5 · outbound

This paper cites Wolf: Dense Video Captioning with a World Summarization Framework.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Wolf: Dense Video Captioning with a World Summarization Framework

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:23:44.937908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:52258c55365f376764379299dc55ab00d6dc459405d2e432b772f4183860f0db

Observation b5c370f8-4a46-49eb-a44d-496b181cd740 · outbound

This paper cites Robotic Skill Acquisition via Instruction Augmentation with Vision-Language Models.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Robotic Skill Acquisition via Instruction Augmentation with Vision-Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:23:44.958498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:cc585923fad79af4a757a7039cbf0f1c931d3c190f1a20099b4e2151f213ca1e

Observation 8252332d-2122-4e1b-8a4e-3e7497ff1a97 · outbound

This paper cites RoboAnnotatorX: A comprehensive and universal annotation framework for accurate understanding of long-horizon robot demonstration.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies RoboAnnotatorX: A comprehensive and universal annotation framework for accurate understanding of long-horizon robot demonstration

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:cbae7afb6c041af1bf9273aa346f7535bffbe75c4e58522fe3eb03fc396279ed

Observation 6810bc69-ab46-4fd6-a821-07127e9cc7f8 · outbound

This paper cites 16 A Appendix Contents A.1 FineVLA-Tool Details.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies 16 A Appendix Contents A.1 FineVLA-Tool Details

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:b89b5247f1683c24600ed85820efca4da54f2b211db1b3edd877c4ffb8268abf

Observation be9629e0-c078-4c29-89f5-d04717e1289d · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:c1036fb68f3aecc4bfc448678ae40b9a0102b6049f3ff2f58694efefb8a424da

Observation 86deb8c7-0014-4790-9d1e-b8e6b51d6ec8 · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:c7ef1276d92bfe1a453a457a6a52a8dda0f9f6d7639306face62be860b852faf

Observation ca6f998f-4188-4f7c-b143-77148b98d4dc · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:7dcdce373ca5e5ecb2bb4575a275879e340fabc636dde1af9dfff8b58b999dda

Observation 8c0b8748-e75c-4c41-b8bd-0bfb07103f16 · outbound

This paper cites Step1":.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Step1":

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:f782def0ee258e750bfacbe0a652f4931a52855c242fb0c64076a971bf8bf3a6

Observation 014cf17a-fb95-4522-a750-cce714d5e4b5 · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:efa271515d349358a88f237c897b8f7fbfa6aff528368206fb88e3a245955b90

Observation 4aa96c8b-c84b-4d20-942b-c91ce50bead7 · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:350dcacf2dacf9bc8b7c981fb979cd6ad4058a49385de2896229ce4ef92422fb

Observation 7382d66d-9a14-4473-939d-f8d17f933fb7 · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:f38e7b5ddcc0fdd79dfc06f18f365b9f9ac00880bc21701f608eeb36f99b27e5

Observation 885371fc-fef3-4f87-99fa-3446aeb9707b · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:f6659ce04e6096976ef85874693ba01bc8f5d7542c514f1d7764baf620c64e40

Observation ca2b34de-6f1e-4ac9-9644-eb9455c7d2e7 · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:b01f8de2474b13d8e9031106939023ef57ce985980325dc1f0f214e739bae231

Observation d42f5209-f7b3-45f8-b387-922d7490c546 · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:5fcd333447f53f985a56d273e16740a4c7951e86ddb0fd515cb467be6e472dbb

Observation 7d772460-0b99-434e-a551-d4d41b0d846f · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:932d9bed6f5a9560852492122648ec7c1191f85c5ee5e2f2baa94cefc79dab93

Observation 46443247-aee5-412e-b485-76deb9383c4b · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:32c4be50ec15fc960eee7bc417af75f43753d396f6799c3f539839d5e213f113

Observation 00246846-1535-4510-8c95-456146c25938 · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:a2b618fe429f92de1a9d6c2629dcbcb1a52ae5b4319cbe1899865461def13eb4

Observation 6f8bf47c-0219-4721-917f-5d4a4a2338c6 · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:ff10a1c8263d485127e322d21668b819f81c5db555aa6cf31f651abb7571bea2

Observation d82bea80-3a71-40df-a2a9-d9b275e8f34e · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:1a3ec192100a8f2336cbc89209a409c36f546beee89009536c1591b95f2db3f9

Observation 9011210b-eb83-44fc-8d11-d49920987b57 · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:55a42015e52553c86d9ec011002ea791571977df7a30f592aaf05c5c4d940384

Observation 11b58061-772b-4d96-8494-1666cac46438 · outbound

This paper cites all/none of the above.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies all/none of the above

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:1b4c8092931879b5cb4def6d57ecfe7781492ff4ea92bb14272c190aa095b1a8

Observation 1c685995-ba6c-4b5c-b313-85326ad0f960 · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:f9d16010239b8b2f3c681dca536fa9bc82d83a1788ec423a4bcacd85b78ae23e

Observation 42313c12-6e06-46ff-b05b-31b9614d66bc · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 55

Resolution
malformed identifier
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:15473899453ea89d5d3bcb9860bdf476dd93ac4876e2598b07cdbcb343d9bca6

Pith citing papers

Observation 7c868e7c-0a96-43be-8bcc-ce4d9ef74c5f · inbound

WorldScape Policy 2.0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory cites this paper.

WorldScape Policy 2.0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T14:13:43.160700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:13:43.160700Z digest=sha256:6a2ee82e5f5d590af37d071d47810675f5b65e3b480c016ac81375f391f7813b