Pith. sign in

Paper Citation Record · LEDGER

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment

As of 11 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 1 inbound Pith citation observation for arXiv:2605.17517.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.17517 v1

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T12:49:02.418663Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:35:34.154516Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T00:35:34.361224Z

Reference resolution

77 of 77 outbound references displayed

  • verified exact18
  • verified fuzzy58
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a80f3fb3-6906-453b-962a-2a7991877373 · outbound

This paper cites Flexible robotic hand harnesses large deformations for full-coverage human-like multimodal haptic perception.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Flexible robotic hand harnesses large deformations for full-coverage human-like multimodal haptic perception

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.129960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:932d6ab0ccbc4bac6df507815101ae0d4356a68e34a14e4f9778a87a104b05bc

Observation e42408a2-f256-4ec3-943d-0c76a2e70be3 · outbound

This paper cites Language-conditioned affordance-pose detection in 3d point clouds.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Language-conditioned affordance-pose detection in 3d point clouds

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.178106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:4ddd11330fac1a80f71b3ea154089d95230a2ade07da14b9651f7b691245ae57

Observation 0a1a5556-9931-449e-aa6f-36b0dca8deb9 · outbound

This paper cites Uad: Unsupervised affordance distillation for generalization in robotic manipulation.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Uad: Unsupervised affordance distillation for generalization in robotic manipulation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.190008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:62992e7b9a31ad0313fe279c3f056ee5f706afc136ebce755741f3a01f70545d

Observation a4e492e0-bb0c-481c-b44b-27dc38e0c981 · outbound

This paper cites Affordancenet: An end-to-end deep learning approach for object affordance detection.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Affordancenet: An end-to-end deep learning approach for object affordance detection

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.172033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:163e1df57ed2468ded39a59f8aaf05e64f3f1713c29be83a78ae862a89723cd8

Observation dba22695-6937-4e36-a271-9c7681c59e48 · outbound

This paper cites Omnimanip: Towards general robotic manipulation via object-centric interaction primitives as spatial constraints.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Omnimanip: Towards general robotic manipulation via object-centric interaction primitives as spatial constraints

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.162648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:d8915b77d909fae8aee4dab82676f8faaeb530904e9031bce400d57fc4fc7d94

Observation e00a138e-3384-481c-a54d-b12839579ef9 · outbound

This paper cites A0: An affordance-aware hierarchical model for general robotic manipulation.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment A0: An affordance-aware hierarchical model for general robotic manipulation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.164870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:287bb79ef5a85cabc98b34c8f825022525f1731db2ccfebab68dcfba2a91e90c

Observation ae510d2a-34a0-4984-94a0-27c099ba9fbf · outbound

This paper cites Manipvqa: Injecting robotic affordance and physically grounded information into multi-modal large language models.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Manipvqa: Injecting robotic affordance and physically grounded information into multi-modal large language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.192041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:f51693b91d8bd3fa1e15cf7b2da830e34e59b502eb7c994273db2806db639b4c

Observation 3462b2f8-4cd0-4b26-abb0-76a06ad58603 · outbound

This paper cites Robots pre-train robots: Manipulation- centric robotic representation from large-scale robot datasets.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Robots pre-train robots: Manipulation- centric robotic representation from large-scale robot datasets

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.193745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:0f8f64b8af7fb88d79ed1d3167e04a89d53546ac2d6d8e97df0f0f0dfc4c128c

Observation a161325b-5166-4c75-9223-9bde3d11b005 · outbound

This paper cites Tars: Tactile affor- dance in robot synesthesia for dexterous manipulation.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Tars: Tactile affor- dance in robot synesthesia for dexterous manipulation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.131703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:e13164446fdb19c2d5c5cae8394442f4648f5d68e5eff8b6348ab4fa94000d0a

Observation 0ab42483-c3c2-48fa-aa66-0363f8c7992e · outbound

This paper cites Sa-dem: Dexterous ex- trinsic robotic manipulation of non-graspable objects via stiffness-aware dual-stage reinforcement learning.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Sa-dem: Dexterous ex- trinsic robotic manipulation of non-graspable objects via stiffness-aware dual-stage reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.104422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:cd39f7b2dc6769998b526479bc12ee072cc8088ea6772989ffeb0386cffbde2a

Observation 75f4f201-3e0b-41c0-a71e-84d5f40f79e4 · outbound

This paper cites Rt-1: Robotics transformer for real-world control at scale.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Rt-1: Robotics transformer for real-world control at scale

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.099880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:77182c27c4bdf1a47f1551a4f98b6003ace012eaa050b8489f379ba3d597f056

Observation 1d4a99c3-927b-4551-9858-c4bfcd7bd42e · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.126223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:4b8b8440a9d123bc6920719559d9eee64543e741b686c6b1ab32b459aa758f77

Observation d674d6fd-e408-4bb7-be12-5ca75bf7f1b0 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:53:17.687925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:7ca97e5a185e681d4365ad968b0a472b1583406571fe1e139b3b29b8b4415ed0

Observation 51a453dd-e221-4464-96f1-2f1c90c42aff · outbound

This paper cites OpenVLA: An open- source vision-language-action model.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment OpenVLA: An open- source vision-language-action model

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.179140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:448961d7102f30d3f2e61d393a9718a246b7a4dd3260eb4bf661a8cd6421bfbe

Observation 9cc0998f-da3c-4191-a37a-803b40ec33a8 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:53:17.726188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:c30a5b7d7f5e37f5d81ad70792c798653b5f8dc47b4746637ea8d171f9032c31

Observation 8b8f72f0-26ba-44ad-b8d8-8c8d7b61e4ff · outbound

This paper cites Vla-jepa: Enhancing vision-language-action model with latent world model.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Vla-jepa: Enhancing vision-language-action model with latent world model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.730425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:055bfb8d0013437cbb9f4b7db9e616a728c6fcf67cdba99389486210994b6714

Observation 59133584-fb53-4ed5-aca1-1e0ad34bd7db · outbound

This paper cites Reconvla: Reconstructive vision- language-action model as effective robot perceiver.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Reconvla: Reconstructive vision- language-action model as effective robot perceiver

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.177442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:d9f95f7d0c87d38012380f28bf5b588b9640301f03528c0489e788f534671572

Observation 432e511c-7e91-4b79-ab64-d558f7ca87dd · outbound

This paper cites Spatial forcing: Implicit spatial representation alignment for vision-language-action model.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Spatial forcing: Implicit spatial representation alignment for vision-language-action model

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.201587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:3afc28eb6467a260f41b63976c29b2c6e30d12c01c818f1ac1d00bb50cf20ac8

Observation 6ab08b49-dd82-4e33-bc51-c421305ab70a · outbound

This paper cites Rt-affordance: Affordances are versatile intermediate representations for robot manipulation.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Rt-affordance: Affordances are versatile intermediate representations for robot manipulation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.184320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:7e367642a111966a6c96414764ee0994200bb381575b1a7442d336869ec49b8a

Observation 0145e820-13d5-4025-9184-55029ff32c71 · outbound

This paper cites Moka: Open-world robotic manipulation through mark-based visual prompting.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Moka: Open-world robotic manipulation through mark-based visual prompting

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.211213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:f350cc785e6156138ca08f3fafd2179ade1402f7ddc3e772a4c08fa2e65ba369

Observation 8e10241c-45ca-45e2-bd69-8b179736809d · outbound

This paper cites Knowledge enhanced bottom-up affordance grounding for robotic interaction.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Knowledge enhanced bottom-up affordance grounding for robotic interaction

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.209407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:d66a380e48d5e446692ba4597ac3a9807dd559c93b29a6caa6fe2178028d0595

Observation 815ad580-1760-4fd9-83fc-2e4cd411505c · outbound

This paper cites Uncertainty-aware state space transformer for egocentric 3d hand trajectory forecasting.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Uncertainty-aware state space transformer for egocentric 3d hand trajectory forecasting

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.190574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:27b551831212e6fe8810896060ef750284f416e73eaeade374986adf0c96ef05

Observation a79f06d0-d910-4c5b-a070-bbb73373db87 · outbound

This paper cites arXiv preprint arXiv:2507.10672 , year=.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment arXiv preprint arXiv:2507.10672 , year=

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:53:17.707386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:50c0634bd15f687cafde548a1ea6712bbf0613b89e769512e3b952a862bae0a2

Observation f2c43b1b-1d17-4e03-ab93-cb4bbc6077ac · outbound

This paper cites Coa-vla: Improving vision-language- action models via visual-text chain-of-affordance.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Coa-vla: Improving vision-language- action models via visual-text chain-of-affordance

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.186883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:920bc411cbb1d67a9d9a64914bbf5219b6318cd19e01b6103a872951ba029ae6

Observation 3982747b-e186-4606-bb30-e8756490d2de · outbound

This paper cites PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:53:17.710898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:35ce5fbffb219a62a8b4360a13c4e66062b28e679cc8b7065cfe5edfd8d7a7ec

Observation 127fac3d-d8a6-4695-bb87-d8394e089365 · outbound

This paper cites GPT-4 Technical Report.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment GPT-4 Technical Report

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:53:17.693960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:6edc30fa2d2cebea8461ab0b2791f2317e9a49d6b46bc2e2d6e5f1ee7aaf2a7c

Observation 26c62ed2-df55-4b25-a3e6-da110d70e524 · outbound

This paper cites Qwen3-VL Technical Report.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Qwen3-VL Technical Report

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:53:17.703792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:28d98121627f9017041bf5407e4d08ed182702afa7a8c65dafb6ff81ff7b3a15

Observation 20807dcc-7c92-4e2d-88d1-eff5233743a0 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:53:17.700714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:7a5d8ead471c1ea71672d09fcff9c809c059925ce15d508536c9202ac4857524

Observation 213ba8ba-50af-4292-a7bf-99076b11954b · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collab- oration 0.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collab- oration 0

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.180900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:cf4a5edb994e0984fbc6bf6d9d4b94745944fc98022c37b9433284046b02eb2b

Observation a56228ed-2ed6-4908-a8be-bd4e9c9ad57b · outbound

This paper cites Cot-vla: Visual chain-of-thought reasoning for vision-language-action models.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Cot-vla: Visual chain-of-thought reasoning for vision-language-action models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.185019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:946e495ca8e623633a69c94032016fa0a10d45780556905d80ebfdc8aff264e0

Observation 2c17658d-5a34-4d27-9ddc-7639f6b5566d · outbound

This paper cites Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:53:17.718213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:d70b99c746e74f8b0113fafea967c0c85e772e7cab386f64de411f4188abc0b2

Observation e9335216-3146-4cdf-b019-fae644691dcc · outbound

This paper cites F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:53:17.714061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:5b0a8a011a4162ca45bfd6b2df76519701ab76438c7171606ff01958c2499d41

Observation bbf4e9fb-371b-497c-bc86-212e66694060 · outbound

This paper cites $\pi^{*}_{0.6}$: a VLA That Learns From Experience.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:53:17.733969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:84f52c0fa7251611ef2184aca7866061426ed1d1b769bc387782a196f245fc87

Observation d5bc91fc-6de5-4798-8560-472502e8c796 · outbound

This paper cites Ig-rft: An interaction-guided rl framework for vla models in long-horizon robotic manipulation.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Ig-rft: An interaction-guided rl framework for vla models in long-horizon robotic manipulation

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.722182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:52506a83808d595d7886c70a4d273e4a6d34db4c901f1e297c452a7b185e42b0

Observation d7845171-13e2-4158-8031-77b30695902b · outbound

This paper cites VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:53:17.670684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:35bdceba32142d6043ead210e08f28a53b9750f9192cb6b608f4fb4e1458c78d

Observation fa1698af-058c-4e4e-b48d-d132b2766042 · outbound

This paper cites The ecological approach to visual perception.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment The ecological approach to visual perception

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.196842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:056e998961ddbd6cdc6871ea35818aff5edd87ae04ab8d82452bfc4d0e217fc7

Observation 2c3ef289-13ec-48d2-8d79-ba01854eafb5 · outbound

This paper cites Affpose: An integrated rgb-based framework for simultaneous pose estimation and affordance detection in robotic tool manipulation.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Affpose: An integrated rgb-based framework for simultaneous pose estimation and affordance detection in robotic tool manipulation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.205791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:0b5405b64a8cf28d472c56605ba0ccb04d895df6e9e628239486e25e95834786

Observation 3da84687-8238-4a0c-8335-f4f2978a242b · outbound

This paper cites Affordancellm: Grounding affordance from vision language models.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Affordancellm: Grounding affordance from vision language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.188007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:bf9d1112f0bdd24c447c5580b6e42d82f9ba30da2b5884f70607dc62b591008a

Observation 85a8b5fd-b186-4dc2-8111-4528b4494520 · outbound

This paper cites Object affordance detection with relationship-aware network.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Object affordance detection with relationship-aware network

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.124475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:1788200d64ac87cab3e663cf64a21f328d91ddf69bc9d9adc1791cfd6f9bfc85

Observation 9f478216-4223-4a9d-a180-b49e0cfbf81f · outbound

This paper cites Learning from 10 Demos: Generalisable and Sample-Efficient Policy Learning with Oriented Affordance Frames.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Learning from 10 Demos: Generalisable and Sample-Efficient Policy Learning with Oriented Affordance Frames

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:53:17.677492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:2ddf0c6d52a22923c898226aa2d1d737c19de2c0034fbcf8306c8fda240f3a36

Observation 2b143792-06fc-4c29-8e2b-d5ebde73bc9a · outbound

This paper cites Closed-loop visuomotor control with generative expectation for robotic manipulation.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Closed-loop visuomotor control with generative expectation for robotic manipulation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.128088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:b8131dd0f6ceef10e064e249e3c2905c2454ee75f5bda59792dcd84c4112c4c6

Observation ffcaae19-44c8-42cb-a209-fc4839c0b3fb · outbound

This paper cites Manipgpt: Is affordance segmentation by large vision models enough for articulated object manipulation?.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Manipgpt: Is affordance segmentation by large vision models enough for articulated object manipulation?

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.162449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:9a23111afcb5b4afdb49bab2d5b42f588f51935cb4f9257df32992dd8ba859a8

Observation a9d2855f-7075-44ec-a710-dec7d8234f71 · outbound

This paper cites R3M: A Universal Visual Representation for Robot Manipulation.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment R3M: A Universal Visual Representation for Robot Manipulation

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:53:17.681129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:897d5366f4dcbcdafca9a3bdeb0a9483c18834e1f744ed78d2fc1749ccc30092

Observation a8dec64e-d821-4735-8a8f-b8f966d90d00 · outbound

This paper cites Representation alignment for generation: Training diffusion transformers is easier than you think.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Representation alignment for generation: Training diffusion transformers is easier than you think

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.174158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:279dfc78ee857508a255507e2afa237e81149b841558fb0e706c42d5a08bfaa3

Observation d1ac1003-d649-41a2-9ecc-b93ef8b26e84 · outbound

This paper cites 3drs: Mllms need 3d-aware representation supervision for scene understanding.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment 3drs: Mllms need 3d-aware representation supervision for scene understanding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.176092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:9d593e027d2487633656834e8245a4e4524ce6ec2f3174f3a60017cfcc09a8bc

Observation 59b041a6-3c7c-4314-a72a-cfb099456f91 · outbound

This paper cites Genhancer: Imperfect generative models are secretly strong vision-centric enhancers.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Genhancer: Imperfect generative models are secretly strong vision-centric enhancers

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.180032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:a655d5220311e7c5370164407648de6eee1c7d9e5ea5de416bfbb82cae2f0b42

Observation fb5616a3-2328-4a7a-acd4-b85ec585c9b6 · outbound

This paper cites Reconstructive visual instruction tuning.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Reconstructive visual instruction tuning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.194834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:1873952ec3b4e1ac96b109721be4732ed3ebd77e24282e446b238f5632d5c458

Observation f977aa6c-3148-43c6-97a9-1613c5ab54f6 · outbound

This paper cites Cross-modality alignment perception and multi- head self-attention mechanism for vision-language-action of humanoid robot.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Cross-modality alignment perception and multi- head self-attention mechanism for vision-language-action of humanoid robot

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.198939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:bda20ee3d1cf4305699fd93d020539e57862a5c6854e0c30aa408357235dad5f

Observation 8e927c4a-c817-4a2b-b8bc-c4ac2c54e454 · outbound

This paper cites Spatialvla: Exploring spatial repre- sentations for visual-language-action model.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Spatialvla: Exploring spatial repre- sentations for visual-language-action model

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.155405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:569267aff21553124cf79afaa4c53f7d20ef42cedbd317ad85bf3ade4ac96db5

Observation 55e7d085-89c9-49ce-9ebb-491aa79a0eb2 · outbound

This paper cites FLARE: Robot Learning with Implicit World Modeling.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment FLARE: Robot Learning with Implicit World Modeling

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:53:17.697262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:98e4a5eac317dddf7a3b3d1a125fd4a0e5d6fd5933c3f43488378e75d14fc421

Observation 46538472-d2c3-4216-bdca-8c6c7ef2b353 · outbound

This paper cites Flow matching for generative modeling.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Flow matching for generative modeling

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.151907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:7a36c9992ad5150a9bcfd61e7e9a6cf2a981c98eb9513855a17c4491c7256a70

Observation b153f260-a7f7-4d0a-869a-ac83214b19bb · outbound

This paper cites SAM 3: Segment anything with concepts.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment SAM 3: Segment anything with concepts

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.152167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:b0bae3e45ba466fd5761a32ec9fe459a4425904fcf31b41aa0e8b9495d40f815

Observation 7367755f-ca11-4dd2-bc85-63e443071924 · outbound

This paper cites Deciphering cross-modal alignment in large vision- language models via modality integration rate.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Deciphering cross-modal alignment in large vision- language models via modality integration rate

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.158038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:71d8e83ab92d5c76e8378c58feb365afde30d8a832aada80c6cb218ae3120b57

Observation fedddaaf-862b-4a52-a075-5ab7b42bc6b9 · outbound

This paper cites Learning affordance grounding from exocentric images.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Learning affordance grounding from exocentric images

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.166985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:ff9def16db4e5a7a67b75829785f05aade213dee4b97002b92401b3cf7ad8f7a

Observation 288c51a2-fa5f-4550-85fd-fd26ce3b5db5 · outbound

This paper cites Locate: Localize and transfer object parts for weakly supervised affordance grounding.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Locate: Localize and transfer object parts for weakly supervised affordance grounding

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.172268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:8660a7337f4842eacdfdabac2712b23bcb4f67bde35a924d70e3f05844a65039

Observation 5dc70278-2283-42ba-a76a-6d4c015f3619 · outbound

This paper cites What do different evaluation metrics tell us about saliency models?.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment What do different evaluation metrics tell us about saliency models?

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.207556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:4048a55da3ea46b4c08f5da250095f3dd9d4013b70f1184f7428d3ee137932d0

Observation 2d48d88a-0444-4c07-a334-145e7b4af50a · outbound

This paper cites Color indexing.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Color indexing

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.170013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:3bfe38940c7825fcb5007894381cd331c63151fecff8361e84a31268d77aa61e

Observation bb2a4a6d-75c0-469f-a7b4-980e48f40ccc · outbound

This paper cites Components of bottom-up gaze allocation in natural images.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Components of bottom-up gaze allocation in natural images

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.164145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:1f150ff26e6a8deb6548e9a31ee699c79f1c90567555b2f028ea43ff862963f7

Observation 9ad4f51e-eddb-41d4-99c8-d6813c24b5cc · outbound

This paper cites Understanding 3d object interaction from a single image.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Understanding 3d object interaction from a single image

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.195716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:0b3d5433bcb5c74762a32f4b286c90d4e0d426839aec6697adbb5729b6dc6228

Observation 10841c49-fa8b-461a-8d98-cba3d08ef548 · outbound

This paper cites One-shot open affordance learning with foundation models.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment One-shot open affordance learning with foundation models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.192756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:df20e2eff426901f8028f20a4bed2b1dcb01a03d4f737612cbc9f1566f719da2

Observation 03e23903-42be-4165-b5d9-b49c85e3ec21 · outbound

This paper cites AffordanceSAM: Segment Anything Once More in Affordance Grounding.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment AffordanceSAM: Segment Anything Once More in Affordance Grounding

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-13T01:20:01.599725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:7a31b6bb51ff14ff2e885ce98e897c2eb0fe3ba132f5f465e365badc717ffba5

Observation 95f9dfb1-24a1-40d7-b1d6-7928772f05ab · outbound

This paper cites Grounded human- object interaction hotspots from video.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Grounded human- object interaction hotspots from video

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.142475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:16ae1e17b397acd564b96ab566f21527d2057db4c9151d9591489ca3749e6465

Observation da60d217-8234-45df-a374-5b2e104f4b55 · outbound

This paper cites Intra: Interaction relationship-aware weakly supervised affordance grounding.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Intra: Interaction relationship-aware weakly supervised affordance grounding

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.203639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:58fcfd2002ce1328c9e583c938a6a57dbd01ff516062da085a9dfb674078f768

Observation 14467bd2-8949-4af2-87a9-f7d9651dc935 · outbound

This paper cites Resource-efficient affordance grounding with com- plementary depth and semantic prompts.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Resource-efficient affordance grounding with com- plementary depth and semantic prompts

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.139099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:2ae937a999720f6e8d92f1bf9606fbd425ccc31913cf51a03b2dad6022cd1484

Observation 444c3b59-7f1c-4494-a816-434a279fd112 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Lisa: Reasoning segmentation via large language model

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.142970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:7de645f9e84186f41bba562b29883ac932dad12268fcd7b7bdadf118b9b40f88

Observation d2900695-e64b-4ed3-be2a-fd3b5ecda3bc · outbound

This paper cites MMR: A large-scale benchmark dataset for multi-target and multi-granularity reasoning segmentation.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment MMR: A large-scale benchmark dataset for multi-target and multi-granularity reasoning segmentation

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.146825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:07ab6e7cc72c91f9009a64c689d98876c6b65843e894d5be3d46e1aa8dc2d89e

Observation a12663a9-6238-44a9-b0ad-96e20f6dc542 · outbound

This paper cites Affogato: Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Affogato: Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:17:13.047323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:b8bf6d05400b4a6efbbfde747bc9fc93970a22ae704bd4c3d927df6f5470dd29

Observation d2ee9b20-a786-49d4-b477-149a7b6c18e4 · outbound

This paper cites RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:53:17.684753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:1f2ef5cb00424bb0829e8aa4b4f3f1d992f4d6b58c8849310ac06608adcf547f

Observation 281fbaa8-e280-4702-9c90-90d952b320c6 · outbound

This paper cites Learning fine-grained bimanual manipulation with low-cost hardware.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Learning fine-grained bimanual manipulation with low-cost hardware

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.144895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:456a42b945fb9d394235ea5f9c223e84f0ad64ddcc731db0a0b008a10f959450

Observation cfbbd6f9-3a5d-4464-a8a2-b9d83e5a26c4 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Diffusion policy: Visuomotor policy learning via action diffusion

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.125435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:85ddcab42a4c15d23639bad19f7b77ddf159dacd06bcb4edd0d1c92098837beb

Observation 397de45f-204b-4998-8c52-20277e796062 · outbound

This paper cites 3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment 3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.136913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:063f148c9ff835220804f432945a37baaa50610008823b7457433336bbcb4ad3

Observation 45ab4b1e-79b5-4b88-8693-366b49c3da5d · outbound

This paper cites RDT-1b: a diffusion foundation model for bimanual manipulation.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment RDT-1b: a diffusion foundation model for bimanual manipulation

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.140760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:9078e59fd32299353fd3ef75e02cfe531d4bcc1010a165e4207849f825f1d014

Observation 93351783-2e91-4a60-81d5-39cb3fb18aea · outbound

This paper cites One-shot transfer of affordance regions? affcorrs!.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment One-shot transfer of affordance regions? affcorrs!

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.146298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:21e4a65819981d15b47edf7df307e91327d1312e381f4c4de2f643d91bf10b3a

Observation f73198b2-c764-435c-ba7a-3a9773eff58b · outbound

This paper cites Weakly supervised multimodal affordance grounding for egocentric images.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Weakly supervised multimodal affordance grounding for egocentric images

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.133498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:44fdc6097c72ba5980bb29de32d25d1eb3ec70203ea91481dc0940c2966f8693

Observation 18133da4-1426-44f0-8c0b-3e2c8ed6f92e · outbound

This paper cites Weakly-supervised affordance grounding guided by part-level semantic priors.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Weakly-supervised affordance grounding guided by part-level semantic priors

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.135126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:e4f217fa0689b34556cc81b519db693d999171629acefba5d3db6b7cfcec92a6

Observation 309184ed-5d07-4737-9b7b-7cfd51332b42 · outbound

This paper cites Reasoning mamba: Hypergraph-guided region relation calculating for weakly su- pervised affordance grounding.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Reasoning mamba: Hypergraph-guided region relation calculating for weakly su- pervised affordance grounding

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.148311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:80adf4987e5dedf39a9c5256acb7381574affc92561bb9dce4734e75c3116493

Observation 7051b13f-ea93-4d3f-80a3-dbeba47f1486 · outbound

This paper cites Visualizing data using t-sne.

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment Visualizing data using t-sne

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:53:29.170395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:49:02.418663Z digest=sha256:f653f465aa3ab1de6b986abb89955e8ccd1784814796b35af782c00654b83488

Pith citing papers

Observation 26ed6c4c-190f-40b2-acc4-f190ccbacca1 · inbound

LIRA: Local Cross-Layer Information Routing for Vision-Language-Action Decoding cites this paper.

LIRA: Local Cross-Layer Information Routing for Vision-Language-Action Decoding AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment

Reference 50

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T00:35:34.368582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-11T00:35:34.154516Z digest=sha256:ed344fff96a3df9310befe304f7a37dcec2b07fc6005f397818e941884bf54b9