Pith. sign in

Paper Citation Record · LEDGER

RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 73 inbound Pith citation observations for arXiv:2406.10721.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.10721 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 73 of 73 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:28:05.321365Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f1362d0b-4ad5-4183-9415-315b15b6b608 · inbound

A Survey on Vision-Language-Action Models for Embodied AI cites this paper.

A Survey on Vision-Language-Action Models for Embodied AI RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-24T01:25:54.470499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T01:25:10.150459Z digest=sha256:be9b498a17100a02817011a7282dbd035f09cbec7887e4b9898a770b4cbf1794

Observation a86b71ee-5115-4234-9f7b-5b5ec4fca855 · inbound

ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation cites this paper.

ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 112

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:25:17.993141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T08:25:17.847571Z digest=sha256:14b690dd7b6963c7310967ad56ae89bac36c1d56aeb9c5ebfb5405d1d7bc9f27

Observation e2de53a4-3b18-4735-991b-48d8882f0416 · inbound

On the Dual-Use Dilemma in Physical Reasoning and Force cites this paper.

On the Dual-Use Dilemma in Physical Reasoning and Force RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:05.321365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:05.321365Z digest=sha256:09a5de96e5c446e5ec4c92f309e0948d295dc70ca77c9a567fc275a8bfef35c5

Observation 8be2bea4-2890-496c-bca7-06d60f75503d · inbound

PartInstruct: Part-level Instruction Following for Fine-grained Robot Manipulation cites this paper.

PartInstruct: Part-level Instruction Following for Fine-grained Robot Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:54.956409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:54.956409Z digest=sha256:8aa8d80792063e955152e5e975276890b3cd7d91fc89d4534eb6c2327c413621

Observation effe07d9-5415-4556-af8c-7c6f59977368 · inbound

OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis cites this paper.

OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:48:02.636894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:48:02.636894Z digest=sha256:8242f2fcca5fa9330553ebeb5356c74264f27248629d3eafa986c2a15d842363

Observation cd915169-45e4-4cce-b1e1-7e8450643794 · inbound

VideoMolmo: Spatio-Temporal Grounding Meets Pointing cites this paper.

VideoMolmo: Spatio-Temporal Grounding Meets Pointing RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:30.956653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:30.956653Z digest=sha256:79811946c2c6073efa37946fe14da482950021add5ca6fa8f707248de1a11298

Observation 27e4b751-6c31-4737-b6b3-7cb5d2275871 · inbound

Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation cites this paper.

Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:56.073790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:56.073790Z digest=sha256:64b539032434c9dbf83b604ddf2deebe59560ada95890838eb7cd97e5d669616

Observation d130ec4d-70b9-4aee-89bd-8c518839ea01 · inbound

UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation cites this paper.

UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:54.748524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:54.748524Z digest=sha256:04161d5653236203d64b5692cafd4da33d4486026aa3e734d2629c20050d2b4d

Observation 822ca968-e8f6-4d95-9ddf-8cea73989b70 · inbound

GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation cites this paper.

GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T04:17:12.693223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:17:12.693223Z digest=sha256:64bce86328684ddf73ba3e8122007a5949f222be118da118334805abc9540b4f

Observation 97df0da4-788d-40b8-96b0-9f6b05992343 · inbound

Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation cites this paper.

Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:17:20.434960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:17:20.434960Z digest=sha256:d36db29bb825612bc20585670100ca3b1c83bd9ba1863a3d6913bff013f12494

Observation d2c9f0f7-54d6-4314-9f05-e85f520a24dd · inbound

CodeDiffuser: Attention-Enhanced Diffusion Policy via VLM-Generated Code for Instruction Ambiguity cites this paper.

CodeDiffuser: Attention-Enhanced Diffusion Policy via VLM-Generated Code for Instruction Ambiguity RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T23:43:16.136090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:43:16.136090Z digest=sha256:6fd4345392b6141acd70af0b11d2e84e55c996a72a809ea3eec3255b1128285d

Observation 919de7e8-553f-4b0f-a847-4d1c298511a7 · inbound

Hierarchical Vision-Language Planning for Multi-Step Humanoid Manipulation cites this paper.

Hierarchical Vision-Language Planning for Multi-Step Humanoid Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:54.651218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:54.651218Z digest=sha256:b5096cbda91da7018b09c12adfe1a043b273a6a2bd96d93c33dd8d1b885ed029

Observation 3fd2d4a8-16fb-4e9f-8587-06bfe8fdcf83 · inbound

RoboBrain 2.0 Technical Report cites this paper.

RoboBrain 2.0 Technical Report RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.110057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.110057Z digest=sha256:a7cb046a5cacd52459a3372b81f9d3a062067f19cf81c5a2a6d73380d1a3334f

Observation 2920025b-eab7-4974-8b61-b7e3e2db957a · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 137

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:42:41.639539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:263f2e4cb8378f58cd8e190e265424129d0f5c4f4f2b8c6c5bf1e6fbbcbab0a6

Observation 1c2ef52c-cfd3-4b4d-a73a-02f4766f676c · inbound

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos cites this paper.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 144

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:49.890593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:49.890593Z digest=sha256:60d2c3c0e4dd05bac525c171f87519818eab811f2c204d869f664bec37ad0494

Observation 799b436f-3e57-409b-abd4-28772b3e6b03 · inbound

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning cites this paper.

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:00.931026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T03:18:14.655384Z digest=sha256:b800da936327b28e5c86d681744eb8252f83e3a2a9409a732f8b96fde7607168

Observation 5a47d01f-2053-4c84-9239-7b43e393ea94 · inbound

PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic Manipulation cites this paper.

PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T23:07:31.209664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:07:31.209664Z digest=sha256:0b47bd1b248505b191e4e5f4007b72fdd38dc33aa80674030870ad45ecf59dfe

Observation 1883ef15-e4ae-4ef4-a52b-c9eb240603b1 · inbound

AimBot: A Simple Auxiliary Visual Cue to Enhance Spatial Awareness of Visuomotor Policies cites this paper.

AimBot: A Simple Auxiliary Visual Cue to Enhance Spatial Awareness of Visuomotor Policies RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T21:44:59.524590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:44:59.524590Z digest=sha256:f391598995a295a696232fbc37557eb2786449f0f39a0edc306f15d56ffa1c9b

Observation 84fa15c7-d9e8-47b4-a224-0d2076156adc · inbound

Robix: A Unified Model for Robot Interaction, Reasoning and Planning cites this paper.

Robix: A Unified Model for Robot Interaction, Reasoning and Planning RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-05T12:59:12.140493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:59:12.140493Z digest=sha256:557319895241a7b39bb974d7555a691713e60354e3803527f73d8e44b7918e21

Observation edd14926-7aa5-46c8-882f-ee5306c133e8 · inbound

Weakly-Supervised Learning of Dense Functional Correspondences cites this paper.

Weakly-Supervised Learning of Dense Functional Correspondences RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-05T10:37:52.706273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:37:52.706273Z digest=sha256:8dec2f7a1269ea7a7664fb435e377d2a062f56babbb8f0986473bbbb14f170a2

Observation 47f571ff-d220-4d07-a286-9e0699cbc58e · inbound

O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation cites this paper.

O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T23:55:43.219675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:55:43.219675Z digest=sha256:9937f100e548b978d1167f452f387bda9f0c32614313da8fe3d14580309e8ae0

Observation 46605936-c719-4104-b1f6-212c958e9f39 · inbound

FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models cites this paper.

FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:51.681730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:54:51.681730Z digest=sha256:fc9e215a1d0d6acaea50e178926f58a545f614616081072a02f13872b6a82bce

Observation faddb885-6fc6-4180-a6fc-477c51cb433b · inbound

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert cites this paper.

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T11:38:13.449967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:38:13.449967Z digest=sha256:3b023f1f41ffb41e61bde5054a19b58b47205c78631ff5e04119f21769f8f842

Observation c60c20fb-9034-4501-9924-567e7ebe634a · inbound

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy cites this paper.

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:09:40.060801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T20:09:39.677347Z digest=sha256:f265247d0b7891687c579de996e442fffd13237e4b84a9acf8ce018a3d268d66

Observation e2afd202-39a6-4391-aed0-81760dca456f · inbound

LACY: A Vision-Language Model-based Language-Action Cycle for Self-Improving Robotic Manipulation cites this paper.

LACY: A Vision-Language Model-based Language-Action Cycle for Self-Improving Robotic Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:35:27.890069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T07:34:55.240907Z digest=sha256:a7d76530b8b4950a97e8b0158316a81264b088064e3ca09fa375ae0d06e1a8d0

Observation 8e378c4c-a10a-4695-a5af-ca05f8850fde · inbound

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards cites this paper.

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:55.221065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:08:55.221065Z digest=sha256:dca0a12bd321bd4cbecd4b5e51c34a2a1e68201029d25291e489937a49ad1458

Observation 8ff12f55-fc60-4a0c-b496-23efb6822fd4 · inbound

MiMo-Embodied: X-Embodied Foundation Model Technical Report cites this paper.

MiMo-Embodied: X-Embodied Foundation Model Technical Report RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:42:05.663147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T20:40:54.096289Z digest=sha256:f9104f5dfb220e6aa2781dbda9fb07a8c86e2f3bed404f469b46cd8fe6c34736

Observation 2804e13d-11be-4ea6-8a0c-eacf649fd7f2 · inbound

SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL cites this paper.

SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-03T18:42:17.143144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:42:17.143144Z digest=sha256:ea91ada86a453a93bbd1a5a9cccbe1de780e6ec9deb989cd47915412afe13b02

Observation aa24a293-4403-4269-b033-1a30ab6cf798 · inbound

Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics cites this paper.

Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 113

Resolution
unresolved
no resolver link, observed 2026-08-03T16:27:34.236050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:27:34.236050Z digest=sha256:76a79e5ac38892af326c7b8aa7ea0df0c91a379eff913fea2ba02903b296ba79

Observation 9cd5c50c-af50-4cf1-8e1a-614f4bb1962a · inbound

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training cites this paper.

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T13:29:52.116644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:29:52.116644Z digest=sha256:d5d9308cee51132f9213cd9f35722d9538868ceca7140f9a85b4aeba8dbe576d

Observation 391defdb-eab7-48c5-806a-e3101db63784 · inbound

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models cites this paper.

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T12:30:33.879651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:30:33.879651Z digest=sha256:0eaba7c8503a4d183bc7be611ab7db2073a68b60b69d8264ce60f45ced0c4dfa

Observation 43e46557-d8fa-4ca2-8f09-9299bcc4f620 · inbound

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation cites this paper.

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 135

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:08:01.848923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T15:05:21.907878Z digest=sha256:a15b8b423095c5e55df99f517137429da5bba5d27f7363bbfd31ead992991b5a

Observation 9274b740-af13-46e0-9de9-c52b97d3638b · inbound

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation cites this paper.

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T17:15:42.488946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:15:42.488946Z digest=sha256:a163c008bb92cd29df2f71c3ce07a9da306e8c6c366276d7daeb08f77e3aae73

Observation 254beb79-cf5c-45dc-af76-9e0c1b5988d7 · inbound

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation cites this paper.

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T05:43:13.314491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:43:13.314491Z digest=sha256:a2faee3529bbc2e96a0e43961f2eeccf10760d9fc19a0edb9e50ccf9c6629abb

Observation de868af6-064d-48f7-8e51-4de660d3900e · inbound

LAST: Leveraging Tools as Hints to Enhance Spatial Reasoning for Multimodal Large Language Models cites this paper.

LAST: Leveraging Tools as Hints to Enhance Spatial Reasoning for Multimodal Large Language Models RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:00:47.718596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:26:23.660206Z digest=sha256:c95b814443bc4dad838b334dd59de8bf02758bb0bd08b226e8aff1d4afae8272

Observation 9e1b85eb-f5c9-4993-a106-14ab9545ac34 · inbound

AnySlot: Goal-Conditioned Vision-Language-Action Policies for Zero-Shot Slot-Level Placement cites this paper.

AnySlot: Goal-Conditioned Vision-Language-Action Policies for Zero-Shot Slot-Level Placement RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:21:00.596307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:41:28.529494Z digest=sha256:164df6f21e6f4098e8e27befe09498928776daabcccde849a7f534c492d19ec1

Observation 6e8dcd6a-1d33-4205-a770-2014147731f1 · inbound

Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models cites this paper.

Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:56:29.225311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T04:27:18.284698Z digest=sha256:b97a01d99b2f7bee7ceafa3641b05f0d81a107596ef1cc0f9fa0b566732374b8

Observation da05135e-e967-4443-9587-68188f6ea59c · inbound

Assessing VLM-Driven Semantic-Affordance Inference for Non-Humanoid Robot Morphologies cites this paper.

Assessing VLM-Driven Semantic-Affordance Inference for Non-Humanoid Robot Morphologies RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:16:16.400879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:02:16.082857Z digest=sha256:26e545a9f22cb86ecf03d74f4b9d8e879808e536bb7a22d323ca569a356e521c

Observation c3db3868-3eeb-4fb3-845e-da793759f331 · inbound

Exploring Spatial Intelligence from a Generative Perspective cites this paper.

Exploring Spatial Intelligence from a Generative Perspective RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:49:48.855770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:45:46.261005Z digest=sha256:b9629ca5337a00a68b38739ddc4b9ba4024fbb0b0c229c5a2b0ad536d642151e

Observation 5e866b2c-037e-410b-925d-368ac96618d4 · inbound

PRTS: A Primitive Reasoning and Tasking System via Contrastive Representations cites this paper.

PRTS: A Primitive Reasoning and Tasking System via Contrastive Representations RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:51:29.287485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T08:56:32.164424Z digest=sha256:8a63cc3f02bce74555e57ab277ef58a7fe7494ee6a29afaff95393f9e6475e09

Observation 12f7b1f8-18d8-401a-8b25-6097b7f9a629 · inbound

Robotic Desk Organization: A Multi-Primitive Approach to Manipulating Heterogeneous Objects via Environmental Constraints cites this paper.

Robotic Desk Organization: A Multi-Primitive Approach to Manipulating Heterogeneous Objects via Environmental Constraints RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:25:39.411846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:30:30.675027Z digest=sha256:cfd96e20ec1c07d1143b05f960509c0f17ea0bd6b70578b3b039abda056b241b

Observation 97c48c6d-4df8-4c1a-86fd-39b946f8bd03 · inbound

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models cites this paper.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:26.143646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:affae67f8a4b531ae9f7969e0264e7ddaf839e81324e4864fee43bc25f1914f9

Observation 90163eae-dc20-4ea3-8123-6f0d3136c618 · inbound

X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction cites this paper.

X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:52:17.036818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T04:50:07.792757Z digest=sha256:679c626dcec735605cef55c8019516ae82bfd825023ad127f95e061ca6b0d851

Observation 0fb92b2c-e9b2-4941-86a3-45e559b548f3 · inbound

SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding cites this paper.

SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-30T20:55:04.172064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T20:51:56.131205Z digest=sha256:1c545c5dd5262295e71b5b6ac060c2d63d4bb72b5bd6439f14f10f1679e2c9f0

Observation e03f028b-ed1b-44dd-b163-cecab73588cb · inbound

Humanoid Whole-Body Manipulation via Active Spatial Brain and Generalizable Action Cerebellum cites this paper.

Humanoid Whole-Body Manipulation via Active Spatial Brain and Generalizable Action Cerebellum RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T04:19:33.404871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T04:16:15.510851Z digest=sha256:2d16d621897da685b0b647f20d433c219836f7228182631e23d49bbb23484616

Observation 9d8fe995-50e9-4862-8ea5-8197db99b03d · inbound

Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance cites this paper.

Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:44:48.553382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T15:36:13.134340Z digest=sha256:8fe620a2c06417644976782aac94474ff406489f60638ed4519e61ea5ee07ab3

Observation e447a59a-8567-4e38-890b-11f3f06a3558 · inbound

Extending Embodied Question Answering from Perception to Decision cites this paper.

Extending Embodied Question Answering from Perception to Decision RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:33:59.010152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T21:30:40.182958Z digest=sha256:606f7094d72de8c224595c9233945cb7e8c77d547bde8dfed3415e107cebced3

Observation 16fc1a9d-0a96-405d-833a-d9b5e6981623 · inbound

GEM: Generative Supervision Helps Embodied Intelligence cites this paper.

GEM: Generative Supervision Helps Embodied Intelligence RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.820593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:dd03a5aff01397413eb876167de289331e01337876220d15326ac07882596952

Observation 88e654cc-355e-462e-8b12-7f2678198020 · inbound

Wall-OSS-0.5 Technical Report cites this paper.

Wall-OSS-0.5 Technical Report RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:26:00.522786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:35:00.258436Z digest=sha256:fdd193072b9870cad0626d825565cd8bc2101c83ac47c19dcd668ab193334d38

Observation f8729099-9003-4783-9eae-7ff1f06f14fa · inbound

Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration? cites this paper.

Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration? RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:12:24.790681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T17:08:59.111306Z digest=sha256:d31b1f966e872aaf96ce4a165c6a3326303a8c66de2279f42b0649d534afeced

Observation ce2fd4cf-eb14-47a6-b0e6-06898011e70d · inbound

GeoAlign: Beyond Semantics with State-Guided Spatial Alignment in VLA Models cites this paper.

GeoAlign: Beyond Semantics with State-Guided Spatial Alignment in VLA Models RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:26:29.221106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T10:04:10.627420Z digest=sha256:14ab3186ab9a95c3d4c54a530cf2828c69ce926fda1694e5bf7e389cdd4a5762

Observation a2989cb6-a0c0-4749-8013-d6925f1b1801 · inbound

AxisGuide: Grounding Robot Action Coordinate System in RGB Observations for Robust Visuomotor Manipulation cites this paper.

AxisGuide: Grounding Robot Action Coordinate System in RGB Observations for Robust Visuomotor Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:56:59.994289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T00:47:21.667833Z digest=sha256:f9c4390f9c16d5e605ece23b212f219e868a1c2f16cbb74502bfc1cadf1b930c

Observation 814a907f-5180-4204-9aa0-d918b211eebe · inbound

Two Bridges, One Pathway: From VLMs to Generalizable VLAs with Embodied Trajectory-Coupled Data cites this paper.

Two Bridges, One Pathway: From VLMs to Generalizable VLAs with Embodied Trajectory-Coupled Data RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:26.625464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T18:31:03.548002Z digest=sha256:7f7bf9d8867e1f088a202b031c7286c645c1cb09868685ae6b1e04ef7bab864d

Observation 7e2dfe1e-ff0c-453d-802c-5d0b80d94d4b · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T13:10:56.935354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T12:55:47.754632Z digest=sha256:85d4d9443b6f5c18da4d55b84dc946ac423d6359e231fb494b7f4dc43e26442f

Observation b3fff410-d358-4348-b8e6-ec8766235b83 · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 105

Resolution
unresolved
no resolver link, observed 2026-07-14T18:07:09.018997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T18:07:09.018997Z digest=sha256:aeac184260b0d8bad18ab9977672a65736ece893c3d1fb2e558654fa2975c1cf

Observation 1361bfc4-c107-491a-9f80-94ee524203d0 · inbound

SVoT: State-aware Visualization-of-Thought for Spatial Reasoning via Reinforcement Learning cites this paper.

SVoT: State-aware Visualization-of-Thought for Spatial Reasoning via Reinforcement Learning RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.935424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T09:59:02.899488Z digest=sha256:3d64e852b97ac82c4f2f26a644615948c1236ec2a48333dcc84c0bded825215e

Observation 58227ffd-37e0-4898-a3bd-91655f024069 · inbound

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models cites this paper.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:48:32.779436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:21e0e0ec24688603d0b329de9eb796e71813df7e41698578eca5a5b2795d873a

Observation 34ca37b2-dfa9-485e-933e-ebf12aacaae3 · inbound

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought cites this paper.

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:24:38.355937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T11:17:48.279808Z digest=sha256:306c9b1db2fea10f8b0608631adb4144a8a05d31bccb796c2351f43f1acafccb

Observation 491565ac-848c-4828-8859-870777043396 · inbound

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought cites this paper.

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T11:20:27.619258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:20:27.619258Z digest=sha256:2c7aa4449df9eabea20c57ebeb5d6ed8de55ec6661d82ecf027caa32c2b38e6c

Observation ca5517f8-f184-4605-aa09-f9fb89a051c2 · inbound

ZeroDex: Zero-Shot Long-Horizon Dexterous Manipulation via Multi-View 3D-Grounded VLM Reasoning cites this paper.

ZeroDex: Zero-Shot Long-Horizon Dexterous Manipulation via Multi-View 3D-Grounded VLM Reasoning RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:59:20.278059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:49:01.269513Z digest=sha256:ef2906a59b03da8d3dae89c8a9dc668be2369281cbec3965fa00db4b17dfa454

Observation 13ac652e-62d1-4a4f-ac54-99b95a57b51c · inbound

Vesta: A Generalist Embodied Reasoning Model cites this paper.

Vesta: A Generalist Embodied Reasoning Model RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 148

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:29:35.142839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T16:55:12.518255Z digest=sha256:c45d7b06c498332a496f2640cc9f319152bd30af8d36730e9985ea72bd44b081

Observation 1e56303d-cfbf-4606-852e-e181b9b08c8f · inbound

CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming cites this paper.

CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:49:42.515654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T10:54:49.774490Z digest=sha256:abd7d10beb409c0ccdfc8695868ec3e69f4c139ee38f15558bb187540ec6dfe7

Observation 1df1a1d4-2a42-43d7-9246-91b017b79155 · inbound

CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation cites this paper.

CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:49:56.523133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T01:25:21.796778Z digest=sha256:640a8dee3eff8689b63510a3d963c5245b471cd2c0ad7c5ea00a3fee5e975ad9

Observation 12c2c2d0-9b1c-4ad9-9380-92a2adc8ed2e · inbound

CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation cites this paper.

CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:05:49.688510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T00:54:50.393828Z digest=sha256:9ac2e2c6bd43b2eadc6e60ddd418b6d865eefd14eaaf7c3efa0601743851598d

Observation c3eeb3aa-87ef-4a5c-bf9a-778eb1f90d17 · inbound

TAP-VLA: Tactile Annotation Prompting for Vision Language Action Models cites this paper.

TAP-VLA: Tactile Annotation Prompting for Vision Language Action Models RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:24:40.641098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T09:10:55.016700Z digest=sha256:d36f315b0192e49c53f558b6a1dd2740fb83e7466b645ea874fc9acff0aadce7

Observation 15278caf-c70f-4e39-af8f-f2e00244729b · inbound

Enhancing Part-Level Point Grounding for Any Open-Source MLLMs cites this paper.

Enhancing Part-Level Point Grounding for Any Open-Source MLLMs RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:14:25.843437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T08:06:08.002125Z digest=sha256:ccef4da39d3c4bc4b62e89123b4a1ca073bf3d6e5b3eef75886cdb27e69c2904

Observation 34828b7b-75a3-440e-bbb2-385c40c67ae5 · inbound

OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping cites this paper.

OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-07-02T14:17:02.408465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T14:16:39.649823Z digest=sha256:df44736841fc2217b1ae42d31047e1868a48516757134f18e03b63251cfa5fc6

Observation 95305c74-c76c-43a8-b68f-eff34a19beb5 · inbound

EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos cites this paper.

EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-14T17:30:54.988498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:30:54.988498Z digest=sha256:506858867ecafa76a5becc09a531f067189a45ed5bc54ef9a3a9db4dc2008656

Observation fa6146c2-7270-49b2-a86e-7391c98c9509 · inbound

Action Map Policy: Learning 3D Closed-loop Manipulation via Pixel Classification cites this paper.

Action Map Policy: Learning 3D Closed-loop Manipulation via Pixel Classification RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T09:52:25.940408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T09:52:25.940408Z digest=sha256:c6ab5806afd6185cc278347f73eb2ec3c693c1d3be260b0474c9c502936ae79d

Observation 58d9205b-0127-4f60-bcdb-574a6b05700a · inbound

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination cites this paper.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.169515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.169515Z digest=sha256:b6f531dcff655b4130aaf61863c09a15c4b0ad60d5f70a973bee35b5141036b7

Observation 2accf09d-3eb9-4550-a3da-5f47d33ff0b6 · inbound

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation cites this paper.

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-01T14:39:42.590429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:39:42.590429Z digest=sha256:424803cf4e2a09c0aa7d1763ed95a54a591425072a5134217fb7f575a775ca0e

Observation 99762039-ea60-47ac-85e5-667f7911b654 · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 295

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:35.421540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:35.421540Z digest=sha256:b0a4b117c9358ef344fa9b20412b98c40c0b3171a71eed5937e078ed53963801

Observation c7340192-801b-4d8f-bdc4-a939dbc5ee48 · inbound

BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation cites this paper.

BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T10:44:18.160988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:44:18.160988Z digest=sha256:2c3445f109045e614e3224ca230ad16ca068e0f47d5b752552b1f6abcd9c4756