Pith. sign in

Paper Citation Record · LEDGER

CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2412.20451.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.20451 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:26:03.787816Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T13:16:59.192241Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e40c096e-f28b-4109-bf45-bfaee6bd7b19 · inbound

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge cites this paper.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:03.787816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:03.787816Z digest=sha256:13e08707cd9ad6288e924944b939bde79f9ca0630fe2d5d5e287b9b8241ddf7c

Observation 1e4cf007-3d7a-4a75-b37c-c233baff2257 · inbound

FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies cites this paper.

FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T05:48:48.013404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:48:48.013404Z digest=sha256:1fcfe415cd9ed9bf2ef9bdf923d545e26eac1b938df8cf5691eb9e643facb8fd

Observation 2cb37b32-e59c-4e28-904d-aa6761045a37 · inbound

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies cites this paper.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:04.939069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:04.939069Z digest=sha256:cc2ed101a95aa80f270cfcb91ef52780d118cea484f2ed0ae3c4bee763c5e41d

Observation 46fd2d0a-a3d7-4d3c-8c44-198cfb7a3dba · inbound

VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models cites this paper.

VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:58:26.715669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T00:54:36.125258Z digest=sha256:452560b1ff83a3d260e460167031c4922431e12c4d27d1dfd85f80766b977960

Observation e0ea0d9d-76c4-4400-9481-c83e3acaa1d2 · inbound

TriRelVLA: Triadic Relational Structure for Generalizable Embodied Manipulation cites this paper.

TriRelVLA: Triadic Relational Structure for Generalizable Embodied Manipulation CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:36:09.184585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T14:54:48.058400Z digest=sha256:887fcfcba09b79fc060ffe89d43c8d4065d4794cb30c413d3d9d8680b6e011b7

Observation 97c78094-a882-4347-91c7-b03004757dd9 · inbound

General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling cites this paper.

General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance

Reference 144

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:33:27.807771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-29T13:33:03.368006Z digest=sha256:63b264603ba9cc500a5a75fc6e51270813cb017e388ae5dc085bd4b7b64803ca

Observation 9aaf89e3-443e-4bd7-882d-8f4c0dd255ae · inbound

GraspGen-X: Cross-Embodiment 6-DOF Diffusion-based Grasping cites this paper.

GraspGen-X: Cross-Embodiment 6-DOF Diffusion-based Grasping CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T21:16:13.226599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T17:26:12.044711Z digest=sha256:fa54f821126d55ee5d80b265f2d46711d6b38be18c3ef9c15881f06d541e3d01

Observation af54694b-a523-42e1-8949-d905413d30aa · inbound

AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding cites this paper.

AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:16:59.193734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T01:23:02.576098Z digest=sha256:eb791f55160a56a158767562f6e55cc1ff9af4a789b820c2ca258c414eb216de

Observation 3b8c218e-6a88-40b9-90ab-70045d15195f · inbound

Data Pyramid for Embodied Manipulation cites this paper.

Data Pyramid for Embodied Manipulation CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance

Reference 195

Resolution
unresolved
no resolver link, observed 2026-07-31T06:18:55.737801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:18:55.737801Z digest=sha256:eb475fe3b4130a827822233315cf65a646988f5f56f7552344f398b679949cc9