Pith. sign in

Paper Citation Record · LEDGER

RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 88 inbound Pith citation observations for arXiv:2406.10721.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.10721 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 88 of 88 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:37:03.502663Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f1362d0b-4ad5-4183-9415-315b15b6b608 · inbound

A Survey on Vision-Language-Action Models for Embodied AI cites this paper.

A Survey on Vision-Language-Action Models for Embodied AI RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-24T01:25:54.470499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-24T01:25:10.150459Z digest=sha256:afec776e38255ff7af407ed102b2245f7f9ea8d618e29b49cc71221f49c19b35

Observation a86b71ee-5115-4234-9f7b-5b5ec4fca855 · inbound

ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation cites this paper.

ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 112

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:25:17.993141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T08:25:17.847571Z digest=sha256:a05ec3f7ba04943ae9addac774044ba85f4c43ec78c49b0ceec67acf4bbb7023

Observation 3b68e0b3-9d45-4a05-a334-aaae2a20b675 · inbound

The One RING: a Robotic Indoor Navigation Generalist cites this paper.

The One RING: a Robotic Indoor Navigation Generalist RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T12:20:28.109167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:20:28.109167Z digest=sha256:2b251682e2792323e4927f90fac44269635da5995ec7c957078f26134d6ce32d

Observation b50fd15f-efef-406f-89b1-3a09f0a7777a · inbound

CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance cites this paper.

CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T23:26:37.471839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:26:37.471839Z digest=sha256:96db1d78a151daed342c8e9945cb327239d7e4cfbd856eb128ed951cda8bec3d

Observation c93a462f-87bc-4a49-a6b5-fe4b6a9b9cc0 · inbound

OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints cites this paper.

OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T21:50:41.392042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:50:41.392042Z digest=sha256:354d5e41938ae1f05a123409c5fb2bd82bdc3b957405fe8692738b80b27a0480

Observation 57a632c2-0a25-495a-9a35-fead3b46b510 · inbound

SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning cites this paper.

SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T19:27:41.288730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:27:41.288730Z digest=sha256:c7fa3ade48b2d9e707349fdcf2e014d385a51168e5b28081f7cf00482de86702

Observation 8ea0237c-f5f4-453d-96ec-09e0aa7d84b6 · inbound

A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards cites this paper.

A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T00:03:16.443621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:03:16.443621Z digest=sha256:c2110c8418d71bc0c40699058f864e2ecd7081883756425d0013350abf621f1a

Observation 3a4a4b18-029e-43cb-8e45-e868caeadc67 · inbound

Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes cites this paper.

Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-16T11:37:03.502663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:37:03.502663Z digest=sha256:1f7bd51b6158feeaf3207e9084bce08604bab8a4061b1003e94bb4c18c74be14

Observation 18a9aea2-11bc-47c2-a9a4-3db97b8b70a3 · inbound

CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation cites this paper.

CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T01:07:35.516610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:07:35.516610Z digest=sha256:c1c433be156ebd098e4c0bfaaa82b463afaff3df22d29b1f15cfd85c0acac88e

Observation 00338ff2-6760-4187-89cd-194e721a005e · inbound

RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration cites this paper.

RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:35.311965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:35.311965Z digest=sha256:8a9c502a7adeb4d1b1f20399381769d7d97ed03bdbcb9278d96af36f1b240910

Observation 76c23de5-1cb1-4420-aa71-f3125f94393a · inbound

Pixel Motion as Universal Representation for Robot Control cites this paper.

Pixel Motion as Universal Representation for Robot Control RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-15T22:12:55.137394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:12:55.137394Z digest=sha256:73d8eee5b68fa2f21cb6d01b5d0648cf8c91e15c0105f83a62fe72509ea23f1d

Observation c3307dc5-5f86-4d01-b4d5-6ed8da136d8f · inbound

Unfettered Forceful Skill Acquisition with Physical Reasoning and Coordinate Frame Labeling cites this paper.

Unfettered Forceful Skill Acquisition with Physical Reasoning and Coordinate Frame Labeling RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:31:00.408492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:31:00.408492Z digest=sha256:4fc7d8c4c94be2594634a644e26e84953ebeca142bd801284697b2cd553df498

Observation b1630d39-e1d2-42ae-9e1e-1d2b95845a11 · inbound

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing cites this paper.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.154135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.154135Z digest=sha256:55247be6f5a3c5922e618b274c3255a4e316ee7a455ea50c2a03beb369c66335

Observation 45cc28d1-1e8c-4088-a5bf-7406679fdd08 · inbound

GraspMolmo: Generalizable Task-Oriented Grasping via Large-Scale Synthetic Data Generation cites this paper.

GraspMolmo: Generalizable Task-Oriented Grasping via Large-Scale Synthetic Data Generation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:08.160963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:18:08.160963Z digest=sha256:fb4e946d4e79ea6c3a6279758bd597cc6f985fdafe42ab4a9d8cc48bb2460350

Observation e2de53a4-3b18-4735-991b-48d8882f0416 · inbound

On the Dual-Use Dilemma in Physical Reasoning and Force cites this paper.

On the Dual-Use Dilemma in Physical Reasoning and Force RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:05.321365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:05.321365Z digest=sha256:5082a123ba4be84e178bd57dac3d9a3964f2d2f14cffb8a9c004841f2f9ab583

Observation 8be2bea4-2890-496c-bca7-06d60f75503d · inbound

PartInstruct: Part-level Instruction Following for Fine-grained Robot Manipulation cites this paper.

PartInstruct: Part-level Instruction Following for Fine-grained Robot Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:54.956409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:54.956409Z digest=sha256:3e4e19ae52fd6e1d2b37074980355490162db6c5dc754da65e0952af7305af3c

Observation effe07d9-5415-4556-af8c-7c6f59977368 · inbound

OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis cites this paper.

OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:48:02.636894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:48:02.636894Z digest=sha256:64f139b06824bb2683cecfc05f29e8de5b780fd50d285424abfedfa82960f963

Observation cd915169-45e4-4cce-b1e1-7e8450643794 · inbound

VideoMolmo: Spatio-Temporal Grounding Meets Pointing cites this paper.

VideoMolmo: Spatio-Temporal Grounding Meets Pointing RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:30.956653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:30.956653Z digest=sha256:f61dac330c8330ec101121de9dfd2e8a1239308bc96bf93833ce22b8f3fa1279

Observation 27e4b751-6c31-4737-b6b3-7cb5d2275871 · inbound

Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation cites this paper.

Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:56.073790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:56.073790Z digest=sha256:38c521e8374fdc6df1466d95853cbfb0101084676427f7012a2703e0d5a8fb40

Observation d130ec4d-70b9-4aee-89bd-8c518839ea01 · inbound

UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation cites this paper.

UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:54.748524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:54.748524Z digest=sha256:6a18dd985f735e3c58a21b96e6832270de0e833f8be08823b7de98138507fbb4

Observation 822ca968-e8f6-4d95-9ddf-8cea73989b70 · inbound

GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation cites this paper.

GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T04:17:12.693223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:17:12.693223Z digest=sha256:dc625860d4cb314b6bb1e2a6a03fee61902a205bb660d59d20f57c3e876d19e4

Observation 97df0da4-788d-40b8-96b0-9f6b05992343 · inbound

Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation cites this paper.

Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:17:20.434960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:17:20.434960Z digest=sha256:2711aa80ef672c6ceda9fdee988a6c2e580cafaaf80d03057b4367bd4a9ebb8a

Observation 8b70650d-21a9-4356-b201-9eabd71f85ef · inbound

CodeDiffuser: Attention-Enhanced Diffusion Policy via VLM-Generated Code for Instruction Ambiguity cites this paper.

CodeDiffuser: Attention-Enhanced Diffusion Policy via VLM-Generated Code for Instruction Ambiguity RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T19:31:53.483663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:31:53.483663Z digest=sha256:a952b90b50a32d29862a59fd50c123538ac6695a9a42189ac80b8cd07980438d

Observation 8444b4ae-aabd-4e78-9011-0d699d113629 · inbound

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models cites this paper.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.612740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.612740Z digest=sha256:f8860546534c6fd079696c5ab6a406eda8aafcf48d01f867b206adfee03cbbe0

Observation 919de7e8-553f-4b0f-a847-4d1c298511a7 · inbound

Hierarchical Vision-Language Planning for Multi-Step Humanoid Manipulation cites this paper.

Hierarchical Vision-Language Planning for Multi-Step Humanoid Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:54.651218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:54.651218Z digest=sha256:c43d625f736cc9f6484e4404229b1d89170ada8272e7ad21af0592b3f7625173

Observation 3fd2d4a8-16fb-4e9f-8587-06bfe8fdcf83 · inbound

RoboBrain 2.0 Technical Report cites this paper.

RoboBrain 2.0 Technical Report RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.110057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.110057Z digest=sha256:9e59c99de09c2c3c05a29c514bc1c6a3da3b3ff511293255fc2b1cdcb0737570

Observation 2920025b-eab7-4974-8b61-b7e3e2db957a · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 137

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:42:41.639539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:37843363a3ce02fc91659dc8c2797199d945def0c5a2d8c6686170c0ba7bc4fb

Observation 1c2ef52c-cfd3-4b4d-a73a-02f4766f676c · inbound

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos cites this paper.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 144

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:49.890593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:49.890593Z digest=sha256:c353bbb0aa2f234570df635ab54da273333c53e5ec89a8512aa779da96c6b374

Observation 799b436f-3e57-409b-abd4-28772b3e6b03 · inbound

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning cites this paper.

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:00.931026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T03:18:14.655384Z digest=sha256:c2013f6e13c2940f587add312f6ce2eb48fd9cb36eaa21be51211710dd77394a

Observation 5a47d01f-2053-4c84-9239-7b43e393ea94 · inbound

PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic Manipulation cites this paper.

PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T23:07:31.209664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:07:31.209664Z digest=sha256:6f76c8893baf628cfdea28a3ac494c4635ebb4954644c189b427ae8534ab955c

Observation 1883ef15-e4ae-4ef4-a52b-c9eb240603b1 · inbound

AimBot: A Simple Auxiliary Visual Cue to Enhance Spatial Awareness of Visuomotor Policies cites this paper.

AimBot: A Simple Auxiliary Visual Cue to Enhance Spatial Awareness of Visuomotor Policies RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T21:44:59.524590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:44:59.524590Z digest=sha256:65c00ad7416f5801add5f505b14d91e5a63042436f8ca4575e0b4646b8d02fb2

Observation 84fa15c7-d9e8-47b4-a224-0d2076156adc · inbound

Robix: A Unified Model for Robot Interaction, Reasoning and Planning cites this paper.

Robix: A Unified Model for Robot Interaction, Reasoning and Planning RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-05T12:59:12.140493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:59:12.140493Z digest=sha256:3c583c3f0bdfc990ce817b818fe331196009727763cff1b4d8a1ee89f0ee5fed

Observation edd14926-7aa5-46c8-882f-ee5306c133e8 · inbound

Weakly-Supervised Learning of Dense Functional Correspondences cites this paper.

Weakly-Supervised Learning of Dense Functional Correspondences RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-05T10:37:52.706273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:37:52.706273Z digest=sha256:3a96daa2466d09d67780b46621bb9e35b7f27b668e91d990c73ea317ecf3ed44

Observation 47f571ff-d220-4d07-a286-9e0699cbc58e · inbound

O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation cites this paper.

O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T23:55:43.219675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:55:43.219675Z digest=sha256:a2f7b79d9bb345fb650aecc5b6938cbe6e64de4d21124f5913e85f5b9ce73f2d

Observation 46605936-c719-4104-b1f6-212c958e9f39 · inbound

FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models cites this paper.

FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:51.681730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:54:51.681730Z digest=sha256:1ee10f46a076cc112474fa6f97304e455e5344ca588ba49e73a441e2d85c4974

Observation faddb885-6fc6-4180-a6fc-477c51cb433b · inbound

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert cites this paper.

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T11:38:13.449967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:38:13.449967Z digest=sha256:537c976c18e0fbeb284e20e0f683c0046faecdf076565d964621dedd141e29bf

Observation c60c20fb-9034-4501-9924-567e7ebe634a · inbound

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy cites this paper.

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:09:40.060801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:09:39.677347Z digest=sha256:d31d48cb6cf9d1f1f29f10dc39b678cdc93fbb91ea726ba13654376b9d0947d9

Observation e2afd202-39a6-4391-aed0-81760dca456f · inbound

LACY: A Vision-Language Model-based Language-Action Cycle for Self-Improving Robotic Manipulation cites this paper.

LACY: A Vision-Language Model-based Language-Action Cycle for Self-Improving Robotic Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:35:27.890069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T07:34:55.240907Z digest=sha256:0d6fee5bc436b13c8ca5bc59460084d99a57e68bd403cecebef47be2bf9f4228

Observation 8e378c4c-a10a-4695-a5af-ca05f8850fde · inbound

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards cites this paper.

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:55.221065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:08:55.221065Z digest=sha256:365aaeb520b96220dee6e0230a18df288eaf781e20b5a66fb72bae81bc28ef50

Observation 8ff12f55-fc60-4a0c-b496-23efb6822fd4 · inbound

MiMo-Embodied: X-Embodied Foundation Model Technical Report cites this paper.

MiMo-Embodied: X-Embodied Foundation Model Technical Report RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:42:05.663147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T20:40:54.096289Z digest=sha256:5d4fd96b3751e2917b9607c000461905f33ae2c74b4512c6633132ece3962f3c

Observation 2804e13d-11be-4ea6-8a0c-eacf649fd7f2 · inbound

SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL cites this paper.

SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-03T18:42:17.143144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:42:17.143144Z digest=sha256:4a139dcfbe600da12409959f47acec30eb1f1a725531070eb10958654f9a450c

Observation aa24a293-4403-4269-b033-1a30ab6cf798 · inbound

Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics cites this paper.

Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 113

Resolution
unresolved
no resolver link, observed 2026-08-03T16:27:34.236050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:27:34.236050Z digest=sha256:c9a6708d6b6740bd84132118ae06ab7e16e502ad833d3c0fa886d593895ab258

Observation 9cd5c50c-af50-4cf1-8e1a-614f4bb1962a · inbound

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training cites this paper.

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T13:29:52.116644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:29:52.116644Z digest=sha256:fcb5f5f22ed6ff29002556c7162e5e59beed7679b544c33aabd9d74f14adc7c2

Observation 391defdb-eab7-48c5-806a-e3101db63784 · inbound

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models cites this paper.

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T12:30:33.879651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:30:33.879651Z digest=sha256:2f46126eccf6833cff7b60bd6e25a5d705c9d43bef25cdb81a33805bee441fa6

Observation 43e46557-d8fa-4ca2-8f09-9299bcc4f620 · inbound

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation cites this paper.

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 135

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:08:01.848923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T15:05:21.907878Z digest=sha256:48a47aabb6fbd49dc15167ad3f3ff11a5f85ccb79be1bd0ec832075cb845764e

Observation 9274b740-af13-46e0-9de9-c52b97d3638b · inbound

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation cites this paper.

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T17:15:42.488946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:15:42.488946Z digest=sha256:56c0673cb41b2ff83ac6937e4930bc5925c55860a16e9f8323fc962ef3ed7b33

Observation 254beb79-cf5c-45dc-af76-9e0c1b5988d7 · inbound

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation cites this paper.

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T05:43:13.314491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:43:13.314491Z digest=sha256:b5fb494b88bdfe8ca059a94b6528e38738b891ff895b7f2e38a579ee984275f7

Observation de868af6-064d-48f7-8e51-4de660d3900e · inbound

LAST: Leveraging Tools as Hints to Enhance Spatial Reasoning for Multimodal Large Language Models cites this paper.

LAST: Leveraging Tools as Hints to Enhance Spatial Reasoning for Multimodal Large Language Models RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:00:47.718596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T19:26:23.660206Z digest=sha256:8f7b1fa8e0968354d3518514fc798289ef2601ca5eb95f655a05db253d985a69

Observation 9e1b85eb-f5c9-4993-a106-14ab9545ac34 · inbound

AnySlot: Goal-Conditioned Vision-Language-Action Policies for Zero-Shot Slot-Level Placement cites this paper.

AnySlot: Goal-Conditioned Vision-Language-Action Policies for Zero-Shot Slot-Level Placement RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:21:00.596307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T16:41:28.529494Z digest=sha256:5846e4f33944b93a5ff9b21739ab5d22e12adad09f9a62b1840e09b0905371e6

Observation 6e8dcd6a-1d33-4205-a770-2014147731f1 · inbound

Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models cites this paper.

Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:56:29.225311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T04:27:18.284698Z digest=sha256:f7016ff8c82350dd9883146fc212d6cc0ca42b9d3010e5a5b8386c78ad213683

Observation da05135e-e967-4443-9587-68188f6ea59c · inbound

Assessing VLM-Driven Semantic-Affordance Inference for Non-Humanoid Robot Morphologies cites this paper.

Assessing VLM-Driven Semantic-Affordance Inference for Non-Humanoid Robot Morphologies RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:16:16.400879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T02:02:16.082857Z digest=sha256:1670b337404a7804022097e956dbe492dabfaf4fb3707284b9c640b3062cab9b

Observation c3db3868-3eeb-4fb3-845e-da793759f331 · inbound

Exploring Spatial Intelligence from a Generative Perspective cites this paper.

Exploring Spatial Intelligence from a Generative Perspective RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:49:48.855770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T00:45:46.261005Z digest=sha256:6ecc9b0c937ddd1ecd827a8ba4301abdf938ebb98b71a9a208990039e1a5dcee

Observation 5e866b2c-037e-410b-925d-368ac96618d4 · inbound

PRTS: A Primitive Reasoning and Tasking System via Contrastive Representations cites this paper.

PRTS: A Primitive Reasoning and Tasking System via Contrastive Representations RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:51:29.287485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-07T08:56:32.164424Z digest=sha256:812d733fb2941f5c5afe1a97004743e6a0b9bdee4f7e5a27e153bc10f92c39c4

Observation 12f7b1f8-18d8-401a-8b25-6097b7f9a629 · inbound

Robotic Desk Organization: A Multi-Primitive Approach to Manipulating Heterogeneous Objects via Environmental Constraints cites this paper.

Robotic Desk Organization: A Multi-Primitive Approach to Manipulating Heterogeneous Objects via Environmental Constraints RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:25:39.411846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T18:30:30.675027Z digest=sha256:7ff9e41e17b1e5d391167bbd552ad0a7ae6d4db48926581e8251e3b79b15a4c5

Observation 97c48c6d-4df8-4c1a-86fd-39b946f8bd03 · inbound

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models cites this paper.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:26.143646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:202a62a75e4db4e3f06aa58fef55a56eb7ef2366b72df30789301ed5799dbe0a

Observation 90163eae-dc20-4ea3-8123-6f0d3136c618 · inbound

X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction cites this paper.

X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:52:17.036818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T04:50:07.792757Z digest=sha256:d871075ee2f4ac455b47962451cf681ce51babc18af1b63c0f70c65414daf14a

Observation 0fb92b2c-e9b2-4941-86a3-45e559b548f3 · inbound

SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding cites this paper.

SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-30T20:55:04.172064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T20:51:56.131205Z digest=sha256:012d3911a6e04310ec6d4b94c34ffd5c840adbe535467b1dacfbb5b205a62bee

Observation e03f028b-ed1b-44dd-b163-cecab73588cb · inbound

Humanoid Whole-Body Manipulation via Active Spatial Brain and Generalizable Action Cerebellum cites this paper.

Humanoid Whole-Body Manipulation via Active Spatial Brain and Generalizable Action Cerebellum RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T04:19:33.404871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T04:16:15.510851Z digest=sha256:8eeb8452c697f19e40ec90305c422eddbfca487fd82a3c798e97af9e240dd9dd

Observation 9d8fe995-50e9-4862-8ea5-8197db99b03d · inbound

Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance cites this paper.

Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:44:48.553382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T15:36:13.134340Z digest=sha256:f853b61f10eeee66ec8198b8608abc8658d22796314205585e1f81db1d2fba17

Observation e447a59a-8567-4e38-890b-11f3f06a3558 · inbound

Extending Embodied Question Answering from Perception to Decision cites this paper.

Extending Embodied Question Answering from Perception to Decision RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:33:59.010152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T21:30:40.182958Z digest=sha256:20224a1ac82d26c586521cc57ab3d6bdfcc6f2db73814a4d6ea4d49699061431

Observation 16fc1a9d-0a96-405d-833a-d9b5e6981623 · inbound

GEM: Generative Supervision Helps Embodied Intelligence cites this paper.

GEM: Generative Supervision Helps Embodied Intelligence RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.820593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:5cb4f13ad760cc5f05b567465846643ae1524381d103cb9f8844b34cd0498dff

Observation 88e654cc-355e-462e-8b12-7f2678198020 · inbound

Wall-OSS-0.5 Technical Report cites this paper.

Wall-OSS-0.5 Technical Report RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:26:00.522786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T22:35:00.258436Z digest=sha256:6490317141b878044c832d61012037892963fd821f627128ff8b138750ab111d

Observation f8729099-9003-4783-9eae-7ff1f06f14fa · inbound

Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration? cites this paper.

Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration? RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:12:24.790681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T17:08:59.111306Z digest=sha256:1837b377fdbdb2fc24cab62688ed2a3e8750126ca77f96f1709e5e53ae738818

Observation ce2fd4cf-eb14-47a6-b0e6-06898011e70d · inbound

GeoAlign: Beyond Semantics with State-Guided Spatial Alignment in VLA Models cites this paper.

GeoAlign: Beyond Semantics with State-Guided Spatial Alignment in VLA Models RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:26:29.221106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T10:04:10.627420Z digest=sha256:833b0e734914ecc30c8661c24bd765e76d03959aa27aacb7bdc61ac42bbf3640

Observation a2989cb6-a0c0-4749-8013-d6925f1b1801 · inbound

AxisGuide: Grounding Robot Action Coordinate System in RGB Observations for Robust Visuomotor Manipulation cites this paper.

AxisGuide: Grounding Robot Action Coordinate System in RGB Observations for Robust Visuomotor Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:56:59.994289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T00:47:21.667833Z digest=sha256:097eb4493ae40ed4878db1a0324638369728cf3f0834a80b71948092c289550f

Observation 814a907f-5180-4204-9aa0-d918b211eebe · inbound

Two Bridges, One Pathway: From VLMs to Generalizable VLAs with Embodied Trajectory-Coupled Data cites this paper.

Two Bridges, One Pathway: From VLMs to Generalizable VLAs with Embodied Trajectory-Coupled Data RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:26.625464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T18:31:03.548002Z digest=sha256:23dd05e8169285493876460d506cd995002de2ffe18302abeed5efb7dbf31096

Observation 7e2dfe1e-ff0c-453d-802c-5d0b80d94d4b · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T13:10:56.935354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T12:55:47.754632Z digest=sha256:c6d4899229ddf6a2482a3081bd11c66d096c70f7e246e3c87b3e536fbddb7971

Observation b3fff410-d358-4348-b8e6-ec8766235b83 · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 105

Resolution
unresolved
no resolver link, observed 2026-07-14T18:07:09.018997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T18:07:09.018997Z digest=sha256:5ca5d6827d94754e449433ea6ec0ed74db15e2b15d9c26734b7ca8da4a010efd

Observation 1361bfc4-c107-491a-9f80-94ee524203d0 · inbound

SVoT: State-aware Visualization-of-Thought for Spatial Reasoning via Reinforcement Learning cites this paper.

SVoT: State-aware Visualization-of-Thought for Spatial Reasoning via Reinforcement Learning RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.935424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T09:59:02.899488Z digest=sha256:24bf528a866daee6cf04fa9244ad5ee9be7c69e670c65d84c1f573e0fccaa0c4

Observation 58227ffd-37e0-4898-a3bd-91655f024069 · inbound

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models cites this paper.

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:48:32.779436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T06:55:53.760595Z digest=sha256:c470a491d7096c04af8a8767c7419c498c3d49496aee30d8335329db0da010c3

Observation 34ca37b2-dfa9-485e-933e-ebf12aacaae3 · inbound

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought cites this paper.

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:24:38.355937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T11:17:48.279808Z digest=sha256:d7daa276eff117ae7110123bfb428876c0dad619e40e41f211387f55ce523110

Observation 491565ac-848c-4828-8859-870777043396 · inbound

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought cites this paper.

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T11:20:27.619258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:20:27.619258Z digest=sha256:4266f49f324a4d88bedf5ce80fe6ecdbbd9dc74688d038614f05e9d8d8db5786

Observation ca5517f8-f184-4605-aa09-f9fb89a051c2 · inbound

ZeroDex: Zero-Shot Long-Horizon Dexterous Manipulation via Multi-View 3D-Grounded VLM Reasoning cites this paper.

ZeroDex: Zero-Shot Long-Horizon Dexterous Manipulation via Multi-View 3D-Grounded VLM Reasoning RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:59:20.278059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T20:49:01.269513Z digest=sha256:ab7a28259d747f0edb6846cb618bb7eee75ac100e3ace36ca9ad144644ccc8c6

Observation 13ac652e-62d1-4a4f-ac54-99b95a57b51c · inbound

Vesta: A Generalist Embodied Reasoning Model cites this paper.

Vesta: A Generalist Embodied Reasoning Model RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 148

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:29:35.142839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T16:55:12.518255Z digest=sha256:5b17159bcf19f534c52da0f59ec7fce6fe2a586294536eb32968f09ad033c12f

Observation 1e56303d-cfbf-4606-852e-e181b9b08c8f · inbound

CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming cites this paper.

CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:49:42.515654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T10:54:49.774490Z digest=sha256:ee81b59ca244c937c1e517943dce53278aa22ff0c9accb9b3bb6ff8a509131da

Observation 1df1a1d4-2a42-43d7-9246-91b017b79155 · inbound

CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation cites this paper.

CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:49:56.523133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T01:25:21.796778Z digest=sha256:d1e4ee5ed99ea6ac75a6981a0befc8c43af2ff7ac4b0885b379571e11f8fe4dd

Observation 12c2c2d0-9b1c-4ad9-9380-92a2adc8ed2e · inbound

CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation cites this paper.

CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:05:49.688510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T00:54:50.393828Z digest=sha256:764bff75ea563833d9e6ea94a54757d5dde06e7bb179b3a09954ce5e20e5d2c3

Observation c3eeb3aa-87ef-4a5c-bf9a-778eb1f90d17 · inbound

TAP-VLA: Tactile Annotation Prompting for Vision Language Action Models cites this paper.

TAP-VLA: Tactile Annotation Prompting for Vision Language Action Models RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:24:40.641098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T09:10:55.016700Z digest=sha256:5d7c829d1b1ee887a3f133c446ae898adf2aaa4052f84fd4d60513ff277ff502

Observation 15278caf-c70f-4e39-af8f-f2e00244729b · inbound

Enhancing Part-Level Point Grounding for Any Open-Source MLLMs cites this paper.

Enhancing Part-Level Point Grounding for Any Open-Source MLLMs RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:14:25.843437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T08:06:08.002125Z digest=sha256:61085c35a1871c9d1bc7eb190be92011ee87cbfd8bfedbad806fe5088972e622

Observation 34828b7b-75a3-440e-bbb2-385c40c67ae5 · inbound

OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping cites this paper.

OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-07-02T14:17:02.408465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-02T14:16:39.649823Z digest=sha256:5e80bda98c476e90e5ada0e8301a7c158dc8382a4969d13b451a3bc0f8913d26

Observation 95305c74-c76c-43a8-b68f-eff34a19beb5 · inbound

EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos cites this paper.

EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-14T17:30:54.988498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:30:54.988498Z digest=sha256:c9a86b599aeef73410e5f843ba7da627d1b4f43653a1d57e84c96a21f810887f

Observation fa6146c2-7270-49b2-a86e-7391c98c9509 · inbound

Action Map Policy: Learning 3D Closed-loop Manipulation via Pixel Classification cites this paper.

Action Map Policy: Learning 3D Closed-loop Manipulation via Pixel Classification RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T09:52:25.940408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T09:52:25.940408Z digest=sha256:98c06084d5ee10beb58a40250fb2b34a76badeef000e7a93b8fabb53556a60b7

Observation 58d9205b-0127-4f60-bcdb-574a6b05700a · inbound

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination cites this paper.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.169515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.169515Z digest=sha256:9704765ac53787be61cc7fe660117366e0e6ce776a9ed7258b4722e9ce409b58

Observation 2accf09d-3eb9-4550-a3da-5f47d33ff0b6 · inbound

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation cites this paper.

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-01T14:39:42.590429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:39:42.590429Z digest=sha256:0d218e2d16793500778acfe6a9bc4ef5c60a4a89b3c81a5fb224f2d124591859

Observation 99762039-ea60-47ac-85e5-667f7911b654 · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 295

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:35.421540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:35.421540Z digest=sha256:7e420336855158f17c46c1c51f9b1014642c1f5599ddefcf2a2973bd4b543290

Observation c7340192-801b-4d8f-bdc4-a939dbc5ee48 · inbound

BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation cites this paper.

BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T10:44:18.160988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:44:18.160988Z digest=sha256:d240efc4ac473a3722146030cba0d4863db0dcb3910677914777d8766a408dab

Observation 6bb8a24e-72ac-4260-b6c5-6e26a666a9a2 · inbound

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence cites this paper.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:32.953272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:32.953272Z digest=sha256:a8beb891e1d9eff589f1fc9d6abdf0b96b25024efa5ec600deb0044a39a52dce

Observation b12fdf99-bef0-46b6-98e9-272d5d50a9c5 · inbound

G0.5: One Autoregressive Stream for Robot Reasoning and Action cites this paper.

G0.5: One Autoregressive Stream for Robot Reasoning and Action RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.083603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.083603Z digest=sha256:45a74f0cca8b49ee3aab5ac4f3adfa2c13e19c62d15dd782127d997d6cb0c4d0