Pith. sign in

Paper Citation Record · LEDGER

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization

As of 4 August 2026, this Paper Citation Record lists 100 of 118 outbound references and 1 inbound Pith citation observation for arXiv:2605.12369.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.12369 v2

Coverage vector

measured 100 of 118 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T22:11:21.596611Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T06:57:41.245418Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-06-29T07:23:13.398570Z

Reference resolution

100 of 118 outbound references displayed

  • verified exact58
  • verified fuzzy37
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dc2a593c-80c5-40ad-b0bc-599497373c1e · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.768212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:1e8afe738bf34aae3b9b3334a6a5a6cc34996eb8e4a2b268f5a4012d21cec6da

Observation 907794c5-0769-4fb1-8f13-fee97b8c8c09 · outbound

This paper cites Qwen3-VL Technical Report.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Qwen3-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.752028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:a2655ac383573c9a6eb13e864aeb6f7db8711142034b021144271ce1a2394672

Observation f65455d5-38c7-49b0-99e0-6171b4ea0117 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization PaliGemma: A versatile 3B VLM for transfer

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.744841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:11480a0e1ee123208aeca5a4bfeb5f27d529d5f304035315d4faad4f9da76dd1

Observation abda6cf9-8281-44b4-936c-7d0965dcaf16 · outbound

This paper cites 3d cavla: Leveraging depth and 3d context to generalize vision language action models for unseen tasks.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization 3d cavla: Leveraging depth and 3d context to generalize vision language action models for unseen tasks

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.721570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:2a9c3100b6023c76078c6b198463a10348d2bac1b4ed4973fb915753610212b0

Observation c382d61c-c90a-406c-877b-012fe65853ff · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.747659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:ba3cdec59bff5767719c048df615293fd5e676c9a972cc7d07e973d859b69ae2

Observation 36a9286f-4481-4d43-a7cf-b7108fd59050 · outbound

This paper cites In9th Annual Conference on Robot Learning.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization In9th Annual Conference on Robot Learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.762299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:33f5e02f7a53201bf4b58f16446fb3bfc478b5c6e85379483e60219f09d81a07

Observation 5aca3279-be2c-4130-8f72-277c3af3d875 · outbound

This paper cites InRSS.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization InRSS

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.765784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:c1699cdfe29952c5354479b4f7b52bc966d26814731af53a264494e2470073c3

Observation bf6cb1ad-f015-40b1-9f32-f64204dfd787 · outbound

This paper cites Real-Time Execution of Action Chunking Flow Policies.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Real-Time Execution of Action Chunking Flow Policies

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.761147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:4477f8b75abe4057056f8112c5770d6765593c3a8f63c14e13cb059fc1893c82

Observation 66cd4033-5063-4a39-8746-44f5c99b5beb · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization RT-1: Robotics Transformer for Real-World Control at Scale

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.659268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:f8c79d739a025a04690a0fc37cc61014d337296153ba048fb12add6bf468d0d3

Observation 0205eb4e-a35b-4fbc-ab84-9bf1061acbd5 · outbound

This paper cites UniVLA: Learning to Act Anywhere with Task-centric Latent Actions.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization UniVLA: Learning to Act Anywhere with Task-centric Latent Actions

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.727871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:04ddabe80f8f87d7b60e5a627a1a2519af70551ad0357eb7911182b9ff9d337e

Observation 4fcdae89-cf79-434d-91d1-eb06871ca981 · outbound

This paper cites WorldVLA: Towards Autoregressive Action World Model.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization WorldVLA: Towards Autoregressive Action World Model

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.757983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:4e9c8e8a563e4ac619c985c02c666e8da49d79e3f22213e777493602d645ae73

Observation fe976080-1848-4161-8805-2ac4693fa229 · outbound

This paper cites STORM: Slot-based Task-aware Object-centric Representation for robotic Manipulation.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization STORM: Slot-based Task-aware Object-centric Representation for robotic Manipulation

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.734179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:04ce2073a44ce5a2f04e663b738b65d987cf09762fa3efa1b9d8ca90589b21a3

Observation 3e0a94e2-c7fd-49de-b8d3-70363e7a1287 · outbound

This paper cites Unified diffusion vla: Vision-language-action model via joint discrete denoising diffusion process.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Unified diffusion vla: Vision-language-action model via joint discrete denoising diffusion process

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.803303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:48fc6f2a3b3acfafc42e754d07602645335007843a1dbe78a6e77ff71a1a1e19

Observation c1c9e977-2617-446a-9622-4c579efbbdff · outbound

This paper cites TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-24T01:23:05.692011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:372ab34b1e681e63c936616c6e1fcb90c16f64641fde16a2da329760967e5e51

Observation f7ac345a-fff2-4351-9276-ad895e2992b7 · outbound

This paper cites RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.680888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:979ba0085672de339ead5e375635184267d4073b2b3710ee2950b3d52cb73aad

Observation 0eab96a9-07cf-4109-9690-748843fe3d67 · outbound

This paper cites Moe-dp: An moe-enhanced diffusion policy for robust long-horizon robotic manipulation with skill decomposition and failure recovery.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Moe-dp: An moe-enhanced diffusion policy for robust long-horizon robotic manipulation with skill decomposition and failure recovery

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.755305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:e26a376916f2116877fcb8e0410f457e700ab868a19a582a9bdf882d117f386a

Observation 20fa5fa8-2a7d-4e44-9604-17af9a37af51 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, 44(10-11):1684–1704.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, 44(10-11):1684–1704

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.770074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:4b038e10fb70bce1704c814508a5aaa2d8e71bb192d2ade55d0fa2cbed2056ce

Observation 80bc0e30-54ac-474b-a87c-e3285ae2f55d · outbound

This paper cites AimBot: A Simple Auxiliary Visual Cue to Enhance Spatial Awareness of Visuomotor Policies.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization AimBot: A Simple Auxiliary Visual Cue to Enhance Spatial Awareness of Visuomotor Policies

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.771008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:d296ea595371685621c90825bbb3dac02134efc161aa6c30e685e06a5d5269f3

Observation daa4452f-bd03-4453-a140-ff5f26d2f238 · outbound

This paper cites RoboNet: Large-Scale Multi-Robot Learning.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization RoboNet: Large-Scale Multi-Robot Learning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.702470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:2169c26a3b13289793ae138511375cdb7eb89e309ba8501260094b6fb3c51ead

Observation 1f3b432a-a490-4b19-9950-73688370d0e8 · outbound

This paper cites Causal confusion in imitation learning.Advances in neural information processing systems, 32.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Causal confusion in imitation learning.Advances in neural information processing systems, 32

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.766514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:f9dd61cbc28b384c8a070f45122a6e8e71b249ed31e2a458b26afd1fc2c0a7a0

Observation e57183a7-40e2-4b37-981a-ad9601eefce0 · outbound

This paper cites StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.683779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:290e2f9eab670f0489bbc7d66f868d41fc09603c96678fea78b07baabc6d4785

Observation 6a89c519-2d95-49e3-8222-22cb4b3e4e34 · outbound

This paper cites Palm-e: An embodied multimodal language model.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Palm-e: An embodied multimodal language model

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.758776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:4b0b61ac23966dcaabd8dd715eb0edf0bbadc5b1990e0ac007159fa861742b2a

Observation 62b3164c-147c-4f20-b68d-4fb83836f7eb · outbound

This paper cites BridgeData V2: A Dataset for Robot Learning at Scale.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization BridgeData V2: A Dataset for Robot Learning at Scale

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.668471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:9940ddf8f144147a4fb3e12cc843cf824a1963f7a8825ebc218f0770d5729e0f

Observation 8874bb89-180b-43d9-a346-e09e726c0e52 · outbound

This paper cites Interleave-vla: Enhancing robot manipulation with interleaved image- text instructions.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Interleave-vla: Enhancing robot manipulation with interleaved image- text instructions

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.754829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:9fedb94abcbe6a5f230f0c951aa00eed05d70e2af116bb7da6c2702bfe1699d6

Observation 98673013-3de8-477c-aed5-d633874f3ff1 · outbound

This paper cites Peafowl: Perception-enhanced multi-view vision-language-action for bimanual manip- ulation.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Peafowl: Perception-enhanced multi-view vision-language-action for bimanual manip- ulation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.808781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:c3ed8d7a28b145d28381be5251adf35b7e329f765e6c682f662a50bae8866891

Observation 01391be9-7701-4855-9166-e2133a39f239 · outbound

This paper cites Learning skills from action-free videos.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Learning skills from action-free videos

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.649296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:f3d73935c80f54ccbc002303ba914d003ff45bd18e18054e8c6b2523ec0acf14

Observation 5e721c90-c723-4252-bba0-12935dd8fd43 · outbound

This paper cites LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.639681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:1a7fac72e1a020da5231fa002727e2e02c5a1ebe00723f8faebbcdb17247e5b7

Observation e45e44cd-981e-4a39-85c6-90e86a0ccfd3 · outbound

This paper cites Imagenet-trained cnns are biased towards texture; increasing shape bias improves accuracy and robustness.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Imagenet-trained cnns are biased towards texture; increasing shape bias improves accuracy and robustness

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.752205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:22dc0424701ce33a5860220fb0a4c02cffee2e3b4875f2b48471d6a163295610

Observation 5685341c-247a-4456-8774-86bd645646a9 · outbound

This paper cites Shortcut learning in deep neural networks.Nature Machine Intelligence, 2 (11):665–673.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Shortcut learning in deep neural networks.Nature Machine Intelligence, 2 (11):665–673

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.756557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:45ba4eacd38ffe2f2940b2db58e7a0f79218f911bf0b4dbda7b38e08dafd452a

Observation 2a215a6f-e48c-4746-9891-8fc8fee02c31 · outbound

This paper cites Octo: An open- source generalist robot policy.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Octo: An open- source generalist robot policy

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.746001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:64ca2a6c6efcccd7da7a8c01d0e40477955164796d7cb2fd286316b3c10a9560

Observation ffddb4f1-8409-4a8a-8282-ef865607fb61 · outbound

This paper cites Point Policy: Unifying Observations and Actions with Key Points for Robot Manipulation.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Point Policy: Unifying Observations and Actions with Key Points for Robot Manipulation

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.695797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:faf6e23078a1e644dc12a00297da964b5a1ea691cfdcfa05424c52dbdd33e466

Observation 8a2be12c-a90e-4163-b932-3ace62d835b1 · outbound

This paper cites In: 2025 IEEE International Conference on Robotics and Automation (ICRA), pp.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization In: 2025 IEEE International Conference on Robotics and Automation (ICRA), pp

Reference 33

Resolution
metadata mismatch
doi, observed 2026-06-30T22:15:06.230574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:792a80c6b317a0477e8e87afff1ebd413863a42d956a6b41d2b754985b858171

Observation 9793bf6b-f90b-4a80-8771-821d4e32338f · outbound

This paper cites Lora: Low-rank adaptation of large lan- guage models.ICLR, 1(2):3.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Lora: Low-rank adaptation of large lan- guage models.ICLR, 1(2):3

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.753061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:60388ec3fa11c72bd71e56dc3d6f2164234d258b2b17f5fa68d4dd6b75366277

Observation 41cabbca-8318-45c0-9878-268868b632ad · outbound

This paper cites arXiv preprint arXiv:2601.11266 (2026).

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization arXiv preprint arXiv:2601.11266 (2026)

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.698964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:4985fe50c57e2253d6ff3a978004899dd5eca9dbc6df2601c4d8705f0acdf633

Observation 385c36c4-b23e-4695-929c-3085d1fbd09c · outbound

This paper cites Rekep: Spatio-temporal rea- soning of relational keypoint constraints for robotic manipulation.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Rekep: Spatio-temporal rea- soning of relational keypoint constraints for robotic manipulation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.756741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:629caab0e3932ca4b11a485dd58a16ec4ed1fb965bfb815c1e0f0f68d8d8cd56

Observation 41c0716e-5d90-487b-b39c-932016bf342f · outbound

This paper cites arXiv preprint arXiv:2601.03782 (2026).

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization arXiv preprint arXiv:2601.03782 (2026)

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.811947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:05c6584b6f213dfc3a697f39d93764b97524b87e63aa09cb1ba5a3083ffcf32a

Observation 4afbb126-945a-4ee8-afd7-f2e474b63c54 · outbound

This paper cites NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.711751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:fb23aa5dc73895c1a65e6b6073f2c62133f9f3c4dee80ff99fc9ef26b28b5917

Observation f27f4d53-0503-475b-9945-4b8f4c8ed231 · outbound

This paper cites Rlbench: The robot learning bench- mark & learning environment.IEEE Robotics and Automation Letters, 5(2):3019–3026.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Rlbench: The robot learning bench- mark & learning environment.IEEE Robotics and Automation Letters, 5(2):3019–3026

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.768283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:1c41a7385bd7fc9aa0defef6c9b0a4c4a1f2987d639b1a6b45b737a29eb0ac78

Observation 91c22306-6ef1-4d04-aabb-822e51183444 · outbound

This paper cites Galaxea Open-World Dataset and G0 Dual-System VLA Model.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Galaxea Open-World Dataset and G0 Dual-System VLA Model

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.820592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:76917fcbae8c664b52f19a1e3437d4ef00fc0f09f61577e18f2e198441415f30

Observation 813fccc6-3acb-4a32-8fe8-ece4f6c704e5 · outbound

This paper cites AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.779848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:f339118cf5f07ecf93923d54c611de2ec6ff93733a536e993183188780047d33

Observation 23151c02-dbba-4a7e-a20a-c68fe51985fb · outbound

This paper cites VIMA: General Robot Manipulation with Multimodal Prompts.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization VIMA: General Robot Manipulation with Multimodal Prompts

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T14:15:46.737758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:ed4429925d946a1efc6603989fe6032822019c3ae4167d94a114cf07b87b4d50

Observation 964b6e31-659f-406c-b6a4-6dce8e440a36 · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.785105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:73c70d21a88bc0189046d238691a60413efb6b03351f59b4ba3a22a4c401b7ee

Observation 1194d431-83ae-4828-8a93-0d3773786eca · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.774050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:f84430ce6fc1330df3d7e27955676b6777b632db2bc672274bbbea1d2911ac51

Observation cb9c0ab5-33aa-4577-ad49-d5bfd8952f85 · outbound

This paper cites Openvla: An open-source vision-language- action model.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Openvla: An open-source vision-language- action model

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.767652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:d12585e16e27511eee425577b642a352a2726a467e1750f4f544130817495ebc

Observation 9a4e08dd-7356-4480-9b19-d54ee8b7d561 · outbound

This paper cites Trace- gen: World modeling in 3d trace space enables learn- ing from cross-embodiment videos.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Trace- gen: World modeling in 3d trace space enables learn- ing from cross-embodiment videos

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.705494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:5595b0774ee2c0478c24f0ce9c18fc46febfb91b6f134ae47c2670e10d8d43ee

Observation f7e82ceb-08de-4048-895e-5aa71e945c92 · outbound

This paper cites Spatial forcing: Implicit spatial repre- sentation alignment for vision-language-action model.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Spatial forcing: Implicit spatial repre- sentation alignment for vision-language-action model

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.735897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:80d5cbb207b77194a0d2cd803a0eb258068b70a2aa1bad8302b72045e283d422

Observation 8fc40c16-c0f9-4902-b206-f2cfbc94713c · outbound

This paper cites H2r: A human-to-robot data augmentation for robot pre- training from videos.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization H2r: A human-to-robot data augmentation for robot pre- training from videos

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.693021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:286958d4dda078aa0143250cfece4dd3c06c6de5c5681a6573b592354cc41f69

Observation c588b211-cab8-4135-aaa8-63d180fc080a · outbound

This paper cites Language-guided object-centric diffusion policy for generalizable and collision-aware manipulation.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Language-guided object-centric diffusion policy for generalizable and collision-aware manipulation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.737565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:a655aebb9594a61de61e38e5c5a4ffb43b812bccef7cca7fafddc93d48003d49

Observation f2e55135-e957-4b1b-8334-404550414357 · outbound

This paper cites Coa-vla: Improving vision-language-action models via visual-text chain-of- affordance.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Coa-vla: Improving vision-language-action models via visual-text chain-of- affordance

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.739543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:9d48e28c540e459d0979ca300da04e49586948aa1e47f4c3e25b85145068d9bc

Observation e5a625ed-7125-469c-b74b-6676af214999 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.727904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:8266ce10f46b504c363097de1bf924a7f3c5d650d61688b828e5cc36c0ad7bbc

Observation dc46ab99-7999-47bc-8701-123f54108435 · outbound

This paper cites Bridgevla: Input-output alignment for efficient 3d manipulation learning with vision-language models.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Bridgevla: Input-output alignment for efficient 3d manipulation learning with vision-language models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.791195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:90889aa56e6841598bf7b4ed13f16e680918801ff15479fe003a3c2fabf893b3

Observation 11ae13b6-56c3-4c1a-871a-68da35c742d4 · outbound

This paper cites Posa-vla: Enhancing action generation via pose-conditioned anchor attention.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Posa-vla: Enhancing action generation via pose-conditioned anchor attention

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.764609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:b3fbc3ac5f59a6308b2f2e694ee51bb079c503a2014d30aaaca1666cee13a7e4

Observation 3b04815d-efec-4187-a60d-6ead98fb6da2 · outbound

This paper cites Skilldiffuser: Interpretable skill planning for latent diffusion-based manipulation.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Skilldiffuser: Interpretable skill planning for latent diffusion-based manipulation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.723782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:0071683c434d805bdfeb58dd0223850717f92c3c8f5ecc1f722ba0f6e118c31b

Observation ed43a84d-712b-4023-b276-d850142d72b1 · outbound

This paper cites Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.642975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:3c4192e600f01c95ca76f770b99516ac595ac32145f63745c34f599ddcf8e34a

Observation ca914e65-27ce-432c-bed6-778cf7540a44 · outbound

This paper cites ManiCM: Real-time 3D Diffusion Policy via Consistency Model for Robotic Manipulation.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization ManiCM: Real-time 3D Diffusion Policy via Consistency Model for Robotic Manipulation

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.690012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:d70337ddd51724980ea820f9292e3849cd9d01ceb38646776a152799cb9448d8

Observation ded24e10-4128-4875-80d8-7abad477e59d · outbound

This paper cites Depth Anything 3: Recovering the Visual Space from Any Views.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Depth Anything 3: Recovering the Visual Space from Any Views

Reference 57

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T14:15:46.687056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:83533763c7b481766fb992dadac317b2ef02b91ee90c57b7e34f487797cccd9b

Observation 8a29bcc9-0e99-4b03-914a-a10e5f48866c · outbound

This paper cites Constraint-preserving data generation for one- shot visuomotor policy generalization.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Constraint-preserving data generation for one- shot visuomotor policy generalization

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.719403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:54153e29d8ad07ddded04a3d6fcaa0260a6586f7a8765098d335f2c24d747aca

Observation 0259c7f2-4635-43f3-bae9-0feeff25ac83 · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning.Ad- vances in Neural Information Processing Systems, 36: 44776–44791.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Libero: Benchmarking knowledge transfer for lifelong robot learning.Ad- vances in Neural Information Processing Systems, 36: 44776–44791

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.725829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:95c1351790e618faf01335c82f06d360d35b442d01fe98e43610317b9d3a2472

Observation 4cae5174-5271-46f1-a013-9ff83c5e5f5e · outbound

This paper cites Rdt-1b: a diffusion foundation model for bimanual manipulation.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Rdt-1b: a diffusion foundation model for bimanual manipulation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.729660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:5a449ab59fcc5dbe2ae4e674b1a806e078d4af0d6908338f7c36f125b9139b20

Observation 49a110b2-f129-42e4-ba87-ffc0bdd773d9 · outbound

This paper cites Hierarchical diffu- sion policy for kinematics-aware multi-task robotic ma- nipulation.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Hierarchical diffu- sion policy for kinematics-aware multi-task robotic ma- nipulation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.741856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:175f1c0cc8efaf75d139ff06a7459f768dbd8f84c1631534ff4574e713b9c7d8

Observation d2fe6d8f-68c0-4e08-82b5-60bc7443f24e · outbound

This paper cites arXiv preprint arXiv:2510.26742 (2025).

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization arXiv preprint arXiv:2510.26742 (2025)

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.665706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:7866282eb6b4a4d402f7d37e83920fc0c241e508483d7a1baf2702e90339f448

Observation 8b13eaf6-778b-4801-a7ad-cfce6ebfcbe0 · outbound

This paper cites Roboturk: A crowdsourcing platform for robotic skill learning through imitation.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Roboturk: A crowdsourcing platform for robotic skill learning through imitation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.711525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:b2a17538253a8b8829c17af71accbd3be488ade2cba534a90929212eaa014c7e

Observation 339471e0-9a20-4f9a-960c-50f5a63e568a · outbound

This paper cites Calvin: A benchmark for language- conditioned policy learning for long-horizon robot ma- nipulation tasks.IEEE Robotics and Automation Let- ters, 7(3):7327–7334.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Calvin: A benchmark for language- conditioned policy learning for long-horizon robot ma- nipulation tasks.IEEE Robotics and Automation Let- ters, 7(3):7327–7334

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.719162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:2836fdd9a3cd83d489ea7d1aaa729cb1602dd5c7d8c48a4c9bb0c4d075b93062

Observation 9be1d31f-8128-44da-9b4f-b5941e5a0b82 · outbound

This paper cites R3M: A Universal Visual Representation for Robot Manipulation.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization R3M: A Universal Visual Representation for Robot Manipulation

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.646323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:8e72b1079434e756b6baa6425cbda4a309701152800109e208cc1f9aca4c8761

Observation b9cae7b2-75fc-445a-9f48-faeea0ba6504 · outbound

This paper cites Vo-dp: Semantic-geometric adaptive diffusion policy for vision- only robotic manipulation.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Vo-dp: Semantic-geometric adaptive diffusion policy for vision- only robotic manipulation

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.636813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:32a68b20e6270cd7c24f9c3093b7881307f13122f572a87471c47b1850433f25

Observation 08ccf904-cfac-4c24-872c-2f44a7dbd42f · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collabo- ration 0.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collabo- ration 0

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.715382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:b8647929a72a33259b61c80c6d9c475a2ba851f5f69237deb95b8655c444867c

Observation 5ea24653-d07f-4a96-a24e-b9b3554ebe16 · outbound

This paper cites Omnimanip: Towards general robotic manipulation via object-centric interac- tion primitives as spatial constraints.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Omnimanip: Towards general robotic manipulation via object-centric interac- tion primitives as spatial constraints

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.763919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:8a66ea27f2056a1ae1f9e4c615d65fb587c3c0b606ee64fae17e5c9cfa61bbf8

Observation 34755786-5e2d-44f1-a6e3-5db2611a8ba6 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.718648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:6d1aa096b7466459a072c8788d7ee0953657d73fc58f3d5640ed7f4e82164846

Observation 3e860926-4124-47f6-9d5f-98c7f60f4118 · outbound

This paper cites GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.708138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:e38f533c565cd2c8e9a5dcf54b40a7732dfb65f3e83dd8bd2d26bf9b81e4c129

Observation 6bea0502-d20c-421f-b337-877abb3cecf2 · outbound

This paper cites Spatialvla: Exploring spatial representations for visual-language-action model.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Spatialvla: Exploring spatial representations for visual-language-action model

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.703474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:e7c5957ea229c073596fbc0b9fd4c73884c905a78699ecf1a3f176d36117ec14

Observation 94d9e78a-f3b4-4ea6-ac23-7ca1baafa546 · outbound

This paper cites SAM 2: Segment anything in images and videos.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization SAM 2: Segment anything in images and videos

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.707922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:35ca93cbada2fd627513e3607819b50d5a0463f0bcbac8923fa0fc9d1fa63d05

Observation 51dbcf3e-eac4-4d55-99a9-fbd9bc24c6c9 · outbound

This paper cites Grounded sam: Assembling open-world models for di- verse visual tasks.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Grounded sam: Assembling open-world models for di- verse visual tasks

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.723342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:268326dfe65c4c8f08761c2303b20fc4d930b503203f7f61de241ed2dbe551d1

Observation f0261b67-3f12-40b1-95f3-0cd89980cdf9 · outbound

This paper cites CLARE: Continual Learning for Vision-Language-Action Models via Autonomous Adapter Routing and Expansion.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization CLARE: Continual Learning for Vision-Language-Action Models via Autonomous Adapter Routing and Expansion

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.788366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:64264bdf2e5b925d07c7b302d7f669a90ffc2b04e0bf6b709484b610780902e8

Observation afb30d02-6fdf-4379-bac1-0ad5836eb951 · outbound

This paper cites Expertise need not monopolize: Action-specialized mixture of experts for vision-language-action learning.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Expertise need not monopolize: Action-specialized mixture of experts for vision-language-action learning

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.806175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:fb4fcc7a3e21e5569f2629b60bcffd9557d2d2066a966c301c05466245f7073f

Observation 02fd77d4-1ee8-499e-9500-af46074635e6 · outbound

This paper cites GeoVLA: Empowering 3D Representations in Vision-Language-Action Models.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.741680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:2ae707b68249fd60d60d34de0bcc20f4940598ea7809fc523defbe903e940dfe

Observation f744ab31-d85d-42e9-b602-e1a92a3b1570 · outbound

This paper cites Interactive Post-Training for Vision-Language-Action Models.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Interactive Post-Training for Vision-Language-Action Models

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.782239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:db69efb06193a43ba746b958e292c1b636b8bfe782c099ef66a9283ec10f11cd

Observation 424f01bf-4196-4d42-9636-179e6ebaa185 · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Gemini Robotics: Bringing AI into the Physical World

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.730823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:bbefa9ac8843406bb293c7d9ac5c39fd15ef459d29b5200f1cc47d40e6529977

Observation 856ce5a5-da23-4465-af5c-050e10ea4607 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.814311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:6755fad082c3b384a16f63c04139a5ffa44b4f0cc6174122143c967f31abcfd7

Observation 8f495e66-b840-4698-ae55-51a91a413d0e · outbound

This paper cites Attention is all you need.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Attention is all you need

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.697454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:fde1040b59bab606e5b3a7732dcc4c93936416f163dd331eaf2ed7d756fc5e5b

Observation 59576cda-2c6b-49f8-8301-01283c491aef · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Bridgedata v2: A dataset for robot learning at scale

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.721376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:1f960a5359c4cddb1e1fc61c76f5067e4cd8467ec0c31259001cf91be934ea0d

Observation e56ff89c-c3b8-4864-9d4c-a5edd6a83c77 · outbound

This paper cites Aerial Tensile Perching and Disentangling Mechanism for Long-Term Environmental Monitoring.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Aerial Tensile Perching and Disentangling Mechanism for Long-Term Environmental Monitoring

Reference 82

Resolution
metadata mismatch
doi, observed 2026-06-30T22:15:06.243272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:8d6a3880682e938cc93537beb510ba29a83ebe188c45000efaf0eb9d60d6f270

Observation 0869ef57-6046-4d22-b3fd-92bac388203d · outbound

This paper cites 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.817348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:67681094ac7bdf1407e37935eacf90c44c6194498e530397967d6d7d9689b91d

Observation 0e35531d-9595-42b1-88bd-421b28ba7fb2 · outbound

This paper cites Vla-adapter: An effective paradigm for tiny-scale vision-language-action model.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Vla-adapter: An effective paradigm for tiny-scale vision-language-action model

Reference 84

Resolution
verified exact
doi, observed 2026-06-30T22:15:06.228055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:2c69fdc6629832fa708c78788d05aea28d5dc124ddef7e3782671855d8465bfd

Observation 6a8e3e34-85de-437e-998b-533c131989ae · outbound

This paper cites VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.671785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:25602c352bebf57af974e95761efc7770b4f5fb0a9d2ffd7512bbe742342baee

Observation b4de4511-e01f-4789-b99e-83ddd771bac2 · outbound

This paper cites dvla: Diffusion vision-language-action model with multimodal chain-of-thought.arXiv preprint arXiv:2509.25681.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization dvla: Diffusion vision-language-action model with multimodal chain-of-thought.arXiv preprint arXiv:2509.25681

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.652581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:5b27e36615a8a961b132db32ce04c1fe3ca4cd4e427d2596ac7001b7e1014765

Observation b66ab7d4-3561-4fae-a941-c696d553740e · outbound

This paper cites DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.777011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:7c75fc249dc668e50454693c192e3e1dc8e1978a69b98bc6c21123bc1def69c4

Observation 1dbd3003-1750-4132-a549-57ddf7480848 · outbound

This paper cites Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation.IEEE Robotics and Automation Letters.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation.IEEE Robotics and Automation Letters

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.701481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:efb675bfa9215fbb8ce45eac8b8be1cc9314509783a79351431b89b8be33bef0

Observation cde63f61-20ea-48e5-a2de-24e30f251151 · outbound

This paper cites Af- forddp: Generalizable diffusion policy with transferable affordance.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Af- forddp: Generalizable diffusion policy with transferable affordance

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.699383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:b2bffcf7b6518b5083dc1d1008a546dd9d1fd411f8308930ca0e9bb958d0224a

Observation 2a0facf7-99d6-4d0d-bc8a-f0bc42d61b7b · outbound

This paper cites DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.793658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:4e72fdb837bdc46df109a3a91f423668ab94bcda7c1eb4574d3097630e08270e

Observation 9d7b9c21-80f2-43ea-be59-0829da473950 · outbound

This paper cites Point what you mean: Visually grounded instruction policy.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Point what you mean: Visually grounded instruction policy

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.799926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:123546a8b535e1c430dbc022f0408e32334c84d639f6a78023c6121d05062f26

Observation 326f7b51-42f0-4f96-a899-adb782c0fcc2 · outbound

This paper cites Meta-world: A benchmark and evaluation for multi- task and meta reinforcement learning.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Meta-world: A benchmark and evaluation for multi- task and meta reinforcement learning

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.727232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:0d78ca1cf9e20937cd2b3219476069bbcf8031f4fc615e35f7b0e46bec0ad71f

Observation d6a3858a-c0a8-44ca-b3e0-d605277de5a6 · outbound

This paper cites Depthvla: Enhancing vision-language-action models with depth-aware spatial reasoning.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Depthvla: Enhancing vision-language-action models with depth-aware spatial reasoning

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.715584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:86b5e6c0c5ed912f4ba5ae27cf3f99736cc89247b7d68cc8168421b4cfb8f475

Observation a81d98b1-e7d0-40c8-92b7-7a122807adc8 · outbound

This paper cites 3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization 3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations

Reference 94

Resolution
verified exact
doi, observed 2026-06-30T22:15:06.224416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:5b9ae59d6599c2d5bebe364234f9582dcee89a0a408613a1a03be018311ebc47

Observation d050c9eb-8992-4aea-ae9b-536123d17ea7 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Adding conditional control to text-to-image diffusion models

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.733442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:fd4f749b80e5f5ffcfe677fc043fd8ad5062bba49908c55a78454c4d15722d49

Observation c3714327-8f08-43f1-aaf2-60a7276a7e08 · outbound

This paper cites Dreamvla: a vision- language-action model dreamed with comprehensive world knowledge.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Dreamvla: a vision- language-action model dreamed with comprehensive world knowledge

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.769693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:708de1a71f67499a709abc3c8d2722d079effd411a26a76ed55a0a211ce4b6c4

Observation 4265e98b-0429-462b-aca4-56206e8184c2 · outbound

This paper cites Mos-vla: A vision-language-action model with one-shot skill adaptation.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Mos-vla: A vision-language-action model with one-shot skill adaptation

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.796815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:b7c96b3a98781ad40406bfd00c87f01886cc3aeac62bd3a8961d4c9dc6dcde16

Observation 2464386f-1ff1-42d0-83e5-a098f9a40105 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 98

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.674610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:253b791d63453ee8cf71a89ee26aa9d23ab241e00d9fdcc0fb6ed3bd6f8e4dbf

Observation 5cb8089d-27e2-4331-8fe5-61276a61e83e · outbound

This paper cites arXiv preprint arXiv:2512.24673 (2025).

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization arXiv preprint arXiv:2512.24673 (2025)

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.655539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:d106d9f2edd095f6c2b2a21753cc4510dcef3316fe34a0de99880246989e14d6

Observation ea800739-84e0-446d-bba8-b640f159be34 · outbound

This paper cites 3d-vla: A 3d vision-language-action generative world model.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization 3d-vla: A 3d vision-language-action generative world model

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:43:52.760735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:c83ec3c748b0e30895cb052128372399f54c3069dc89971b2ead7251cd818b83

Observation a4a46732-8e89-4f04-a48b-0583ab9006b2 · outbound

This paper cites TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies

Reference 101

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.662542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:3e27c5f024fcb2570fb4b3beb0654f55378c76784fceec75009ff6f5370f9a7e

Pith citing papers

Observation ca51edbe-f8d4-4f05-ba83-c7afb2962e2e · inbound

ELAN4D: Embodiment-Centric 4D Supervision for Vision-Language-Action Models via Plug-and-Play Adaptation cites this paper.

ELAN4D: Embodiment-Centric 4D Supervision for Vision-Language-Action Models via Plug-and-Play Adaptation GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:23:13.400609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T06:57:41.245418Z digest=sha256:f9b84e061ed2205a283d03caa0007bbc5361704978eca8f1c79d84a01be44dc1