Pith. sign in

Paper Citation Record · LEDGER

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies

As of 23 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 1 inbound Pith citation observation for arXiv:2605.27284.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.27284 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T17:19:04.477325Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T14:13:43.160700Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact25
  • verified fuzzy0
  • unresolved27
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 77170b98-4ed4-43b6-906b-8cbedf936175 · outbound

This paper cites ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:23:44.921783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:bed7b9ae503f4b0fe28653dbdd29b868022c81e46484073b6fe710bc1dce4a2b

Observation 10a67a13-bdf8-445b-b54a-24e0e0b28c21 · outbound

This paper cites A Pragmatic VLA Foundation Model.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies A Pragmatic VLA Foundation Model

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:23:44.921456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:637b40e2d3f15b60273d5bb85f88421bec1b58e4361b5f751f2ad58686e48dd1

Observation fac17dfd-4895-4af0-b7e7-986465c4ffea · outbound

This paper cites Nvidia isaac gr00t.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Nvidia isaac gr00t

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:d248968cf9d1bba0cd5cbb58d4366e96f6e1cdb7a2fc4334e32cda85a5563914

Observation be5c5a54-86e7-4d41-9f8b-7772ffd671ae · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:5dfcc3d854bd0d63c1a2a6f8e2f04ed75404c32e4f5f829c9ebe3b1b85f9671d

Observation 3a975193-7074-45d3-8558-e5f2fbc37bd3 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:23:44.919185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:20c73c49763bd676c3b4d2de7bbe1fab096d6b3f5fcfe4d8b79ae106314a21ea

Observation 48be4c0f-c411-4ecf-af9c-39d32773f460 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:23:44.926072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:161ccc00b1e07200153adec012c7ae6c2d0f3d30022c5f93f33d1ee80711840e

Observation a37e55ad-f482-454b-93ff-37e929cbe22d · outbound

This paper cites RoboVQA: Multimodal Long-Horizon Reasoning for Robotics.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies RoboVQA: Multimodal Long-Horizon Reasoning for Robotics

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T17:23:44.942418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:729d99f5bd7f3d9282a705212a646d61477a29619c82339704a8a2a8024e5daa

Observation e405e4fd-0188-42b7-9f12-26f5a907802a · outbound

This paper cites Robobench: A comprehensive evaluation benchmark for multimodal large language models as embodied brain, 2025.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Robobench: A comprehensive evaluation benchmark for multimodal large language models as embodied brain, 2025

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:eae72c5fe950907d18610bcdbc86df3ad653c396cb70cfbe0b07be6178bf9c61

Observation 5e99ec83-6d05-43fd-97e4-d30d0a10d217 · outbound

This paper cites Handyvqa: A video qa benchmark for fine-grained hand-object interaction dynamics.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Handyvqa: A video qa benchmark for fine-grained hand-object interaction dynamics

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:23:44.950904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:d91e2aa82c01052aec82dcf1e1c8b26abf53284b580bf19bd54ad83a0dcbc12d

Observation 51a173f8-9cc3-48b2-acce-31f5e3587a7b · outbound

This paper cites Qwen3.5: Towards native multimodal agents, February 2026.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Qwen3.5: Towards native multimodal agents, February 2026

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:90cd28aa221812f8ffa171b4fa333cf19ede1f250c03995cca6fa467750ca375

Observation 2427e160-c65a-4a94-a6b3-205544a521bd · outbound

This paper cites BridgeData V2: A Dataset for Robot Learning at Scale.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies BridgeData V2: A Dataset for Robot Learning at Scale

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:23:44.954034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:791822c2bb4f5c3aed6cc567a261df6b18489c143492b874ff1b2528c9d25537

Observation bcc11873-90cf-4d29-8f84-5bf953f107aa · outbound

This paper cites Bc-z: Zero-shot task generalization with robotic imitation learning,.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Bc-z: Zero-shot task generalization with robotic imitation learning,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:04ff02fc72fd542557340d27dc546be986671fcb2fbd254da2fa64c0d2414112

Observation c240c5ab-de63-4096-bbfd-754a48dcc04b · outbound

This paper cites BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:23:44.943146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:12eb8a04861378f256fcc924b3130ceb3afaae09d4f9938ac9537467ddb15954

Observation 03f9d295-e0bd-466e-a141-a0ffbb246c38 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies RT-1: Robotics Transformer for Real-World Control at Scale

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:23:44.932205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:c170d07ed5c92a0e62513f367c909ba0653565f612ebc6887a99c0d1b0d2f7f7

Observation 9211b989-36e3-452d-ac47-febbf02a7331 · outbound

This paper cites Galaxea Open-World Dataset and G0 Dual-System VLA Model.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Galaxea Open-World Dataset and G0 Dual-System VLA Model

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:23:44.940463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:612426084dbb371a95964fb6946414f56ab647f38f3adbd444f54c8183c62f6c

Observation 87882627-16bc-4000-8218-d2bb3983579d · outbound

This paper cites URL http://dx.doi.org/10.15607/RSS.2025.XXI.152.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies URL http://dx.doi.org/10.15607/RSS.2025.XXI.152

Reference 16

Resolution
metadata mismatch
doi, observed 2026-06-29T17:23:44.401136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:eb511454f67212e4838aeb9e42b663d71b0c610de9d33b40694664a37dbedb51

Observation 9e89b392-c8c2-444d-8ec5-c6607d7046c3 · outbound

This paper cites Robomind 2.0: A multimodal, bimanual mobile manipulation dataset for generalizable embodied intelligence.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Robomind 2.0: A multimodal, bimanual mobile manipulation dataset for generalizable embodied intelligence

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:23:44.947959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:0c901c1de5a3e4be1ec4672f9e94ce073e3a97d1c5469c559061f8baf72b9785

Observation 4848680e-69f9-4d71-8f6d-93e49743127d · outbound

This paper cites RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:23:44.911537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:c0012bdc3d144b93c384d1927f67715d0981ca2a25fa16d9f0a3cf9f2bcd4fbb

Observation 28488e8b-e62f-4778-a6c2-45c888654db1 · outbound

This paper cites Rh20t: A comprehensive robotic dataset for learning diverse skills in one-shot,.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Rh20t: A comprehensive robotic dataset for learning diverse skills in one-shot,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:56f4abb4910e4d6f8789ef5377dfa27e702446257685ab56ab4e79f66ce1ebfc

Observation 5b6b1bf4-fe99-4fe9-8c7b-01841233f225 · outbound

This paper cites RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:23:44.916841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:81fa5b2b6d95e8d6177cc6175a226ff79a52162ba118355fb944c59fae0f8382

Observation 85361763-a493-469c-b7a3-acf63262d240 · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:23:44.892194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:9903444cad6ffd71b9880e09bd48eaae694edbe7451818ff9dee2ac0dca49118

Observation 85ba6212-81ac-442b-ae70-be8981b9601d · outbound

This paper cites StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:23:44.894826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:d30e793ecc4a0e1572327563b8cb11e76ba7062ffa0c997e082c31a854a082b4

Observation 3a404a20-ecdf-45ac-b582-ef18514d00c4 · outbound

This paper cites RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins (early version).

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins (early version)

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:23:44.897400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:bf7b1900d7d3ac78caad3d24ce315d55dee2ed22dfdc65c381e28c692a0d1e7a

Observation 62567983-9e3c-4c23-ac22-6cb251a99e45 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:23:44.913881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:447e7374dbf250453d01f72f554c0a081d8cec2adedef1da6a0b2f28623dd459

Observation 2275bb3f-fa75-4b62-b5f1-b15249c9b085 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies OpenVLA: An Open-Source Vision-Language-Action Model

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:23:44.904879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:4533d00207091afe291479af468efc988bef7898c4f6d5f9f197f39a5e07edb7

Observation 733a7607-0446-4682-a0af-7c6f6f7f99d1 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:23:44.923746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:9d09d18b190d67ff071ddd96c5cf69df18a5e6523ce46c1d2293ba5181935e64

Observation d6ff64a6-a339-428d-9bee-e2f39b9a11af · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Octo: An Open-Source Generalist Robot Policy

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:23:44.919434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:f3b137ed72aa928f23e6b459a9cfd774bb1f64032029fbc5f10c954d1e785ecc

Observation 5af2092c-fc3d-49f7-b280-3d8679e7a9bc · outbound

This paper cites RoboInter: A holistic intermediate representation suite towards robotic manipulation.arXiv preprint arXiv:2602.09973, 2026.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies RoboInter: A holistic intermediate representation suite towards robotic manipulation.arXiv preprint arXiv:2602.09973, 2026

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:23:44.953547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:c314e5b47f16858bedc3d6bf2989fe9197d1d106e53a8c8ebeceaf4b5bed9305

Observation 5a7e1444-131a-4cc0-96ec-56dbdc4576f7 · outbound

This paper cites STEER: Flexible Robotic Manipulation via Dense Language Grounding.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies STEER: Flexible Robotic Manipulation via Dense Language Grounding

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:23:44.937340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:e01743708155b7786580b0726fb069dd555b4b7715f5198ea5eb7f0a211d86d6

Observation 9eed3036-5fa9-433b-a62a-17c7bddfe3eb · outbound

This paper cites PartInstruct: Part-level Instruction Following for Fine-grained Robot Manipulation.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies PartInstruct: Part-level Instruction Following for Fine-grained Robot Manipulation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:23:44.946075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:7f29ca867d51cb2eeb017376ed9ffdbfdf85342b621c7fe65c2c9acb6bca7d5c

Observation 87a4314c-5c8e-4816-9c26-e207e941d957 · outbound

This paper cites Qwen3-VL Technical Report.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Qwen3-VL Technical Report

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:23:44.955913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:e104986e884aa27d1dee34af5a82838e28cd60bc971b2778f5683d7e95b4a34e

Observation 2b405af1-95bc-4f6b-861c-341cc5346d29 · outbound

This paper cites Qwen3.5-Omni technical report, 2026.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Qwen3.5-Omni technical report, 2026

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:3117dbf127ea77bbbabb914b0b95f286f299d8a62cad16a13325c2d60c6af9ef

Observation 9f0d7d01-72f0-4063-90af-4daa7239d4d5 · outbound

This paper cites Wolf: Dense Video Captioning with a World Summarization Framework.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Wolf: Dense Video Captioning with a World Summarization Framework

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:23:44.937908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:2f4079e56774ef40896e65c9697782e7c1fceaf8dcd18a54c699c3dcf648bce6

Observation b5c370f8-4a46-49eb-a44d-496b181cd740 · outbound

This paper cites Robotic Skill Acquisition via Instruction Augmentation with Vision-Language Models.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Robotic Skill Acquisition via Instruction Augmentation with Vision-Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:23:44.958498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:951778348412c56fd7da209ba38c15b899fcdc0aa2a84e44392f2bd99d6351b5

Observation 8252332d-2122-4e1b-8a4e-3e7497ff1a97 · outbound

This paper cites RoboAnnotatorX: A comprehensive and universal annotation framework for accurate understanding of long-horizon robot demonstration.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies RoboAnnotatorX: A comprehensive and universal annotation framework for accurate understanding of long-horizon robot demonstration

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:cbae7afb6c041af1bf9273aa346f7535bffbe75c4e58522fe3eb03fc396279ed

Observation 6810bc69-ab46-4fd6-a821-07127e9cc7f8 · outbound

This paper cites 16 A Appendix Contents A.1 FineVLA-Tool Details.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies 16 A Appendix Contents A.1 FineVLA-Tool Details

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:b89b5247f1683c24600ed85820efca4da54f2b211db1b3edd877c4ffb8268abf

Observation be9629e0-c078-4c29-89f5-d04717e1289d · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:c1036fb68f3aecc4bfc448678ae40b9a0102b6049f3ff2f58694efefb8a424da

Observation 86deb8c7-0014-4790-9d1e-b8e6b51d6ec8 · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:c7ef1276d92bfe1a453a457a6a52a8dda0f9f6d7639306face62be860b852faf

Observation ca6f998f-4188-4f7c-b143-77148b98d4dc · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:7dcdce373ca5e5ecb2bb4575a275879e340fabc636dde1af9dfff8b58b999dda

Observation 8c0b8748-e75c-4c41-b8bd-0bfb07103f16 · outbound

This paper cites Step1":.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Step1":

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:f782def0ee258e750bfacbe0a652f4931a52855c242fb0c64076a971bf8bf3a6

Observation 014cf17a-fb95-4522-a750-cce714d5e4b5 · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:efa271515d349358a88f237c897b8f7fbfa6aff528368206fb88e3a245955b90

Observation 4aa96c8b-c84b-4d20-942b-c91ce50bead7 · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:350dcacf2dacf9bc8b7c981fb979cd6ad4058a49385de2896229ce4ef92422fb

Observation 7382d66d-9a14-4473-939d-f8d17f933fb7 · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:f38e7b5ddcc0fdd79dfc06f18f365b9f9ac00880bc21701f608eeb36f99b27e5

Observation 885371fc-fef3-4f87-99fa-3446aeb9707b · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:f6659ce04e6096976ef85874693ba01bc8f5d7542c514f1d7764baf620c64e40

Observation ca2b34de-6f1e-4ac9-9644-eb9455c7d2e7 · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:b01f8de2474b13d8e9031106939023ef57ce985980325dc1f0f214e739bae231

Observation d42f5209-f7b3-45f8-b387-922d7490c546 · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:5fcd333447f53f985a56d273e16740a4c7951e86ddb0fd515cb467be6e472dbb

Observation 7d772460-0b99-434e-a551-d4d41b0d846f · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:932d9bed6f5a9560852492122648ec7c1191f85c5ee5e2f2baa94cefc79dab93

Observation 46443247-aee5-412e-b485-76deb9383c4b · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:32c4be50ec15fc960eee7bc417af75f43753d396f6799c3f539839d5e213f113

Observation 00246846-1535-4510-8c95-456146c25938 · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:a2b618fe429f92de1a9d6c2629dcbcb1a52ae5b4319cbe1899865461def13eb4

Observation 6f8bf47c-0219-4721-917f-5d4a4a2338c6 · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:ff10a1c8263d485127e322d21668b819f81c5db555aa6cf31f651abb7571bea2

Observation d82bea80-3a71-40df-a2a9-d9b275e8f34e · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:1a3ec192100a8f2336cbc89209a409c36f546beee89009536c1591b95f2db3f9

Observation 9011210b-eb83-44fc-8d11-d49920987b57 · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:55a42015e52553c86d9ec011002ea791571977df7a30f592aaf05c5c4d940384

Observation 11b58061-772b-4d96-8494-1666cac46438 · outbound

This paper cites all/none of the above.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies all/none of the above

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:1b4c8092931879b5cb4def6d57ecfe7781492ff4ea92bb14272c190aa095b1a8

Observation 1c685995-ba6c-4b5c-b313-85326ad0f960 · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:f9d16010239b8b2f3c681dca536fa9bc82d83a1788ec423a4bcacd85b78ae23e

Observation 42313c12-6e06-46ff-b05b-31b9614d66bc · outbound

This paper cites an unresolved cited work.

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Unresolved cited work

Reference 55

Resolution
malformed identifier
no resolver link, observed 2026-06-29T17:19:04.477325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:19:04.477325Z digest=sha256:15473899453ea89d5d3bcb9860bdf476dd93ac4876e2598b07cdbcb343d9bca6

Pith citing papers

Observation 7c868e7c-0a96-43be-8bcc-ce4d9ef74c5f · inbound

WorldScape Policy 2.0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory cites this paper.

WorldScape Policy 2.0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T14:13:43.160700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:13:43.160700Z digest=sha256:6a2ee82e5f5d590af37d071d47810675f5b65e3b480c016ac81375f391f7813b