Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T18:01:16.779978Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 0 inbound Pith citation observations for arXiv:2608.05215.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T18:01:16.779978Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
70 of 70 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4afcadd5-0916-4933-95db-b78d56ec6edd · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Affordance detection of tool parts from geometric features,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 737c059f-3284-4117-bad3-93d266628e98 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Learning affordance grounding from exocentric im- ages,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 45ecd236-96a7-471d-9a0b-a4736986750f · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Object-based affordances detection with convo- lutional neural networks and dense conditional random fields,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 42711c5e-a75c-4848-8349-a382a2261f98 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Learning to act properly: Predicting and explain- ing affordances from images,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a1e962bb-a85c-4fce-9409-fcff03c19749 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances GLOVER++: Unleashing the Potential of Affordance Learning from Human Behaviors for Robotic Manipulation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62e92869-c436-4582-b2e1-320c4eea03ca · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances H2o: Two hands manipulating objects for first person interaction recognition,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6f046d92-7583-46e3-b762-7ef55f82d38e · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Hoi4d: A 4d egocentric dataset for category-level human-object interaction,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 415a5eaf-460c-4d5f-947b-2c03f368e661 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Scaling egocentric vision: The epic-kitchens dataset,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 25a8e8db-9fa3-4db0-8dc0-8cc1165cfd06 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances HD-EPIC: A Highly-Detailed Egocentric Video Dataset
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03f5bd15-b973-486e-b67a-30ca984cf170 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Ego4d: Around the world in 3,000 hours of egocentric video,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 256cdff0-b054-428c-b8fe-391d72767447 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fb84ae5-9966-4d9b-9ea9-393e4f6c26d4 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances LLaMA: Open and Efficient Foundation Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c31d479-3971-455c-ac56-ae9a4f27d3b1 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances GPT-4 Technical Report
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d069619-2cca-4b88-8899-aa069c6b5363 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Efficient training of artificial neural networks for autonomous navigation,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7453ea79-a5e4-4e0c-aace-a6499a1de816 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 199bafe5-7125-46a3-9edb-066e3711d317 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Masked Visual Pre-training for Motor Control
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f311e024-7c68-41b7-b879-9780b9ff46ba · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances R3M: A Universal Visual Representation for Robot Manipulation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecdd8d95-bbfd-49fd-a542-c2319d15605f · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b81d6aa-d3c8-4c2f-b0e0-b2360a35ca36 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Affordances from human videos as a versatile repre- sentation for robotics,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 037045af-ee10-450d-ae62-be357e4848d7 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances GLOVER: Generalizable Open-Vocabulary Affordance Reasoning for Task-Oriented Grasping
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47592aa3-039a-4610-9769-65ec639969e8 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Dexycb: A benchmark for capturing hand grasping of objects,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a049b818-b85a-429c-a7a2-3a051bbb495e · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Videodex: Learning dexterity from internet videos,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 185a3a09-3a05-491c-9e43-c17168325096 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac586404-0b6b-43e5-8dfe-a93647b520b7 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances LIV: Language-Image Representations and Rewards for Robotic Control
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a3bd658-5224-4d52-b996-4ef86f051fa4 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Hrp: Human affordances for robotic pre- training,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1a14e57d-602e-4b17-9612-b82b3bf0bf1c · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28494b12-e07b-4b42-9380-104fbf4387bc · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Weakly supervised affordance detection,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 82283a00-7753-47a8-9c02-746ed0f2ea23 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Locate: Localize and transfer object parts for weakly supervised affordance grounding,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 302c82a3-ba4f-432a-812d-0595356cd66e · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Affordancellm: Grounding affordance from vision language models,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7d9ddb00-788c-4c53-af49-f5898a6771a5 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2e4e690-726b-4e55-8dee-6bb6b08e5796 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Arctic: A dataset for dexterous bimanual hand-object manipulation,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2d0a5f96-7ad8-4499-bec2-947c838f8307 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Contactpose: A dataset of grasps with object contact and hand pose,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 77ba9438-3b92-4fc7-bc55-fdad3caf5c21 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Oakink: A large-scale knowledge repository for understanding hand-object interaction,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ae0e54f0-4f30-4e0f-bb85-d5738497ec7a · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances TACO: Benchmarking Generalizable Bimanual Tool-ACtion-Object Understanding
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7fa2ba1-2ad6-4042-a8a6-b099d5b6de3b · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances RT-1: Robotics Transformer for Real-World Control at Scale
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0834910-a2e2-44ea-b176-3987593bb2f3 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 860b4e08-c7b5-4619-9d48-352a770d01a5 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances OpenVLA: An Open-Source Vision-Language-Action Model
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73ee612d-2cd4-45f6-bc31-a5f081d21bec · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Octo: An Open-Source Generalist Robot Policy
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 740781d5-193d-4326-9eb2-ceca4192ae9f · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a6f6391-124f-473e-bc9c-ac5037d9c8fd · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Droid: A large-scale in-the-wild robot manipu- lation dataset,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 58cc99cf-5216-4b4e-9ed6-eb0907e52b96 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Embodied hands: Modeling and capturing hands and bodies together,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f73652f3-7549-4cf5-b5f7-eb0aab744698 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances On the continuity of rotation representations in neural networks,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation adc1acb2-2eac-4a54-aaae-e125d852bd85 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Understanding human hands in contact at internet scale,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a13e96ae-e07b-4b61-b85f-7cf605259994 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances SAM 2: Segment Anything in Images and Videos
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 337de9e0-c609-4eae-a8f2-0ee84ca9dd0f · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Vitpose: Simple vision transformer baselines for human pose estimation,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2d4fac88-8ece-43a2-b2b5-32513591a9f8 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e644f7fe-012b-4b8d-838f-6ced7cd0d4b5 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Partial Implementation of Max Flow and Min Cost Flow in Almost-Linear Time
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ddd8f7f0-4eae-4885-b123-7ec83f4881a5 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Egohos: Dataset and method for hand and object segmentation in egocentric videos,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cb7605e1-a217-4fb8-9ca8-f5a21f108487 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances ProPainter: Improving Propagation and Transformer for Video Inpainting
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation eecb27d4-4607-4903-aad2-264f6ab5d6c0 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp Details
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2799165f-53ac-4ec4-8364-369b457b7872 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Structure-from-motion revisited,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b735d328-27e1-415b-974e-7dfae58a4a13 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8759fa19-9196-41fa-aa92-4de9b38f996b · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Cosmological parameters estimated from velocity -- density comparisons: Calibrating 2M++
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fdc8a646-0858-4a01-9fd9-d9cbced9a379 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Lisa: Reasoning segmentation via large language model,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 79e1bd6b-7dc9-4cd2-bf84-99dc67c80cf2 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances DINOv2: Learning Robust Visual Features without Supervision
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f23d786b-694b-4c92-a8c0-f58f8cd3218f · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Qwen2.5-VL Technical Report
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfd00c33-f0af-4c47-b0ff-dac2de245f0c · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Handal: A dataset of real-world manipulable object categories with pose annotations, affordances, and reconstructions,
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c2c8c255-f4fe-4369-bd13-4008e10a089d · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Scenefun3d: Fine-grained functionality and affor- dance understanding in 3d scenes,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9d6cb10d-c1e1-4346-a1e3-342f08dc0d22 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Understanding 3d object interaction from a single image,
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c0fc3ebb-0e0f-44b9-ab11-b8af73969e03 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances RAM: Retrieval-Based Affordance Transfer for Generalizable Zero-Shot Robotic Manipulation
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10f87d64-38e4-41bd-83c1-e9b623a85cea · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Generalflow: Generalizable manipulation policy with flow matching,
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbce1271-1683-4f1c-ae77-aded962f06e7 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Graspnet-1billion: A large-scale benchmark for general object grasping,
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0ce24f75-d874-4b80-8241-2dc6602ccc45 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances R+x: Retrieval and execution from everyday human videos,
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8a16cb37-f10f-40e4-ac25-7ee8770c1759 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43ecab4d-afdc-401c-8985-05abc85f90dc · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Partmanip: Learning cross-category generalizable part manipulation policy from point cloud observations,
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4c206d36-488b-4132-8b50-3d28cd05b34c · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning,
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4f11054a-a739-4ca8-b5d3-3f2795a44ca0 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Maniskill2: A unified benchmark for generalizable manipulation skills,
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c3de1a73-28dc-4f25-afe4-a991ea4a3205 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Ag2manip: Learning novel manipulation skills with agent-agnostic visual and action representations,
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 97febfe7-476a-474b-bd2e-eb2a8fc6ef54 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Pointllm: Empowering large language models to understand point clouds,
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a4e6b661-3308-4060-b8a1-0d8ecb8ecdc7 · outbound
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances Generating 6dof object manipulation trajectories from action description in egocentric vision,
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
No inbound Pith citation observations are available.