Pith. sign in

Paper Citation Record · LEDGER

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making

As of 7 August 2026, this Paper Citation Record lists 100 of 137 outbound references and 0 inbound Pith citation observations for arXiv:2506.12374.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12374 v2

Coverage vector

measured 100 of 137 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:57:29.897079Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 137 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved99
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 45199b3e-0170-40a6-bbc6-d428e817f69f · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in Neural Information Processing Systems, 35:23716–23736, 2022.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Flamingo: a visual language model for few-shot learning.Advances in Neural Information Processing Systems, 35:23716–23736, 2022

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:25.081678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:25.081678Z digest=sha256:f44b9383d0bb777bd9184fbdf56128077cb541af3b884c98ada3b9dcb0a4c3e4

Observation e90b2662-f6d1-43e4-b255-88be81a0910b · outbound

This paper cites Learning transferable visual models from natural language supervision.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Learning transferable visual models from natural language supervision

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:25.150439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:25.150439Z digest=sha256:9163190429f7fcbcd1659d46cbcda583b3d3584982275c562d4f63dcae7648ad

Observation a76572f7-59b4-49f5-a18c-98f5025b2ad7 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Scaling up visual and vision-language representation learning with noisy text supervision

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:25.223630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:25.223630Z digest=sha256:18a9876f8371ef854421e4ad0438543f813465c458a5bc5fb9d0556a857f19fc

Observation 4ba83d3d-98be-424e-8a83-be2de65f9adf · outbound

This paper cites Zero-shot text-to-image generation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Zero-shot text-to-image generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:25.300306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:25.300306Z digest=sha256:5a26ab302ae49b43512028ca3116125317d8c72132d057e8c3ce9a6e7df05b98

Observation 8edbc6a0-d9e7-4f3f-a35a-387f5a3a0de7 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making PaLM-E: An Embodied Multimodal Language Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:25.396676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:25.396676Z digest=sha256:b71d2eeac9b03aa55d03035601af00af24502749fc5ff48b4849d847b64b3c4b

Observation 6518237b-bf54-4ede-81e0-f66736cc7e99 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:25.479391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:25.479391Z digest=sha256:e902e1622795f55f82a5accd3cfef2fc453bdb7b4eddea8255e0ff23cf5cd9b4

Observation b8a27dc8-e4ae-4833-92ba-aa3a621e239b · outbound

This paper cites GPT-4 Technical Report.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making GPT-4 Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:25.576785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:25.576785Z digest=sha256:30d95b49dea1e1b1db676114c2d0f10c6320d739e2e423eed9f8069ed5c3cac0

Observation c8333fc0-7cb3-41f0-8bf5-79688f81c4bd · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:25.641202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:25.641202Z digest=sha256:3b4660b07f13707cb2eb676e5f69fd83f8c8286569224a1053c17b5a14e59822

Observation cbf9ea69-12ff-4097-a862-f5a33c4881ad · outbound

This paper cites Open-vocabulary queryable scene representations for real world planning.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Open-vocabulary queryable scene representations for real world planning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:25.750801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:25.750801Z digest=sha256:71ae97b895e049777561674597903ecec5c4e4ef94c327106d36be0c1962f474

Observation b638acb7-e6f4-41bc-bb7b-a8ce4b092d30 · outbound

This paper cites Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:25.851964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:25.851964Z digest=sha256:7036deaf77f4b6c306e6061f2b1e44c7228eeda876b7b465789ceab4321d0b4f

Observation 1fd7d921-6d8c-46a1-8028-a7e487bca854 · outbound

This paper cites Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:25.947466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:25.947466Z digest=sha256:e48ed1635b75fd91a1a191f3d3ac61b48bb6f5901f7a9c0c49a6f30ad7de9499

Observation 0d1561d3-4d7c-4800-a5ac-e1472a3bf328 · outbound

This paper cites Language models as zero-shot planners: Extracting actionable knowledge for embodied agents.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Language models as zero-shot planners: Extracting actionable knowledge for embodied agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:26.062840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:26.062840Z digest=sha256:998e5730d531f8a45f657701ef58592921a778769a5a0f6f78d3b084dfe17f8a

Observation d401c5fd-abdf-45e9-9f76-541046427ecf · outbound

This paper cites Code as policies: Language model programs for embodied control.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Code as policies: Language model programs for embodied control

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:26.188599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:26.188599Z digest=sha256:e8b5eb829f9175b15dbc49e32427dc7a35d864e23ed5986c8bab9f3d5ba0f7a3

Observation 3ee45808-8a8d-4e6c-a6c0-1b2f11779223 · outbound

This paper cites VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:26.264561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:26.264561Z digest=sha256:3391d2ba30cde704a6eb715acb692c7e1f8aa9ab1367d1599fda59ce3ab9a9b4

Observation f898e174-1805-4323-9a7f-b500b241d761 · outbound

This paper cites PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:26.306305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:26.306305Z digest=sha256:3783ff480669186b709e921cdcb145f196cd6b7eda0ec81f177a9c6bcd9f3be2

Observation f1494ef9-d559-40ea-9c75-588bc06f24d5 · outbound

This paper cites ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:26.372509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:26.372509Z digest=sha256:8ee50134c6861c8e25477651b0865316c9cc97a09375a02ab5796447e521ab89

Observation 7699befe-e1bc-4ead-880f-0e1beb5288f5 · outbound

This paper cites RoboDexVLM: Visual Language Model-Enabled Task Planning and Motion Control for Dexterous Robot Manipulation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making RoboDexVLM: Visual Language Model-Enabled Task Planning and Motion Control for Dexterous Robot Manipulation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:26.482172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:26.482172Z digest=sha256:7b69c33ba597fe09bc07ff27c613c1580f478856a895868f40f21e2b2e0b9188

Observation 7c29c10a-ade1-4c72-91e6-45a7d04c8a88 · outbound

This paper cites RoboGround: Robotic Manipulation with Grounded Vision-Language Priors.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making RoboGround: Robotic Manipulation with Grounded Vision-Language Priors

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:57:31.256067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:26.612635Z digest=sha256:be0fd6b473a5347560aecbe6b5c8459a57e8705010a7c86b61f807f71b3baf48

Observation ba1a1b09-4043-4a96-8a6a-f2fe481d4284 · outbound

This paper cites Llm-grounder: Open-vocabulary 3d visual grounding with large lan- guage model as an agent.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Llm-grounder: Open-vocabulary 3d visual grounding with large lan- guage model as an agent

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:26.751072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:26.751072Z digest=sha256:dbb163434ffff7ee0a5ec24dc82a9dc35bd9dfe4fbd32b7f94314b45eaf74215

Observation 4a81f21e-71f4-46d6-a2cb-b4dd1f8683bb · outbound

This paper cites SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:26.863014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:26.863014Z digest=sha256:3832c6dae1e1908f3915dd2725dbece6545ecb65ff0aad38262dd15b41b21a48

Observation 0a19d958-9ac3-4651-8fac-421e7d39b69b · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.025972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.025972Z digest=sha256:65dea0e7c59c3695c0fed05d3287c344d079474efa74ea203d2817a2f2e3285c

Observation e85307f9-bf8d-42ae-aaf8-e44fdc65d69a · outbound

This paper cites SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.127004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.127004Z digest=sha256:c19f959a9046a215105cdafc26530194f9b5e13c171eead3471ec023525b845f

Observation 89ed4db9-1ac1-465e-9632-203f668c3226 · outbound

This paper cites CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.167072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.167072Z digest=sha256:97b504f259325ed9f53478a0bbd90175fef89fa38d9c54f166841cbc7a85c109

Observation 6e20543d-3d2d-4381-8c77-b14dd9b70932 · outbound

This paper cites Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.247850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.247850Z digest=sha256:cd5b136a39e769bd00fe5584badd37b38018a466ce42be0c28b8cbff1cb4e458

Observation af814dc3-26b1-45ef-83e8-940e705c8bfa · outbound

This paper cites Zero-shot visual reasoning by vision- language models: Benchmarking and analysis.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Zero-shot visual reasoning by vision- language models: Benchmarking and analysis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.372595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.372595Z digest=sha256:8b5e7a99946bd7c9061a77799ad539e5934acc310b2ea406de38f6bf94068333

Observation 7b350e45-f5a0-40d5-a504-12a5c2a2c2bc · outbound

This paper cites How to enable llm with 3d capacity? a survey of spatial reasoning in llm.arXiv preprint arXiv:2504.05786, 2025.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making How to enable llm with 3d capacity? a survey of spatial reasoning in llm.arXiv preprint arXiv:2504.05786, 2025

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.459312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.459312Z digest=sha256:22b27b12bf3c84924b735525fc2f3f3572eae01bcfb31e1499163adfe8683d7a

Observation d6bfe043-b776-4725-ac7b-a1fba9f5b982 · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.528961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.528961Z digest=sha256:95d4198558b89e594c5df7c84f4daea9d16f3e2bc115a1285080b4f2a93ffc44

Observation c8499945-4b1d-4d4d-89e3-eaac8adcd815 · outbound

This paper cites LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.599419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.599419Z digest=sha256:7f3c39b2238255ea0f34b94bfe87733b7604e52bd3afa95af2470bbe89f4486b

Observation 06b3bc4a-023e-42d1-9676-72c568e0c030 · outbound

This paper cites Agent3d-zero: An agent for zero-shot 3d understanding.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Agent3d-zero: An agent for zero-shot 3d understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.720163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.720163Z digest=sha256:f49b48b16251124b55577c0dc62871b7f399deff10200c67ebf7903971eb108b

Observation c990ba0e-9bf6-41c4-9c72-612f6e2a9d5f · outbound

This paper cites Shapellm: Universal 3d object understanding for embodied interaction.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Shapellm: Universal 3d object understanding for embodied interaction

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.842050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.842050Z digest=sha256:cb019a1e27d61ba38af81e7cae01c2908d2a79437f588b24b218409f7f8160e3

Observation d451d68d-d5ea-49c2-ac94-441f12a5004d · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:28.005755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:28.005755Z digest=sha256:714a0dd7abc8b09b1ca98a517a8b39c8b45a2a3eeb0ae846c02b5ae5d814d491

Observation e5a6f8f5-13d6-4362-af1f-7a0192560783 · outbound

This paper cites INF-LLaVA: Dual-perspective Perception for High-Resolution Multimodal Large Language Model.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making INF-LLaVA: Dual-perspective Perception for High-Resolution Multimodal Large Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:28.102223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:28.102223Z digest=sha256:b98f53b4bc1d48576238bd8b07ebea4cd40854299e184298ff942a90db9ca022

Observation 7764fc4b-8cf7-4717-926f-f3a11a814c8d · outbound

This paper cites G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:28.210537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:28.210537Z digest=sha256:e7d018d8e3972a3137b840888d5ae68b4d68edb312396e67bd9dce1432f51099

Observation 3c265a4b-d4ec-410e-bbe8-c6540ef8450c · outbound

This paper cites Manipvqa: Injecting robotic affordance and physically grounded information into multi-modal large language models.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Manipvqa: Injecting robotic affordance and physically grounded information into multi-modal large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:28.331337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:28.331337Z digest=sha256:f90f834ee6ca20a008083aff989d8821439ddf334a1e1a93712bfa28c4e93441

Observation 1ceeb23e-fa10-453e-b8d1-9f5377370db9 · outbound

This paper cites Robovqa: Multimodal long-horizon reasoning for robotics.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Robovqa: Multimodal long-horizon reasoning for robotics

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:28.492183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:28.492183Z digest=sha256:ba9b188392085445c83f26a4ad0072e2f5afa59a2b119777b6b92ad5a6f455e6

Observation b4f339b0-0fb5-4577-9c16-10da8048e90a · outbound

This paper cites MQA: Answering the Question via Robotic Manipulation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making MQA: Answering the Question via Robotic Manipulation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:28.618914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:28.618914Z digest=sha256:0cf57d0f6c1ca4bbd44532c085e91df29a0bb609c0fc3d1cd0e902f8ce8b1499

Observation d2251920-b18c-4576-b55b-d1fe88529cf5 · outbound

This paper cites Robotvqa—a scene-graph-and deep-learning-based visual question answering system for robot manipulation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Robotvqa—a scene-graph-and deep-learning-based visual question answering system for robot manipulation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:28.668841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:28.668841Z digest=sha256:f204626bc490126e40550d8b125133165e4047e5c5063405a7cf8383b2f32351

Observation f7b5ad7a-a19a-4970-88a9-af407adeb473 · outbound

This paper cites A visual questioning answering approach to enhance robot localization in indoor environments.Frontiers in Neurorobotics, 17:1290584, 2023.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making A visual questioning answering approach to enhance robot localization in indoor environments.Frontiers in Neurorobotics, 17:1290584, 2023

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:28.748826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:28.748826Z digest=sha256:f53e176c4d02b73545974f08261b731aa7721fe490ec21f6560e08a09fedd95f

Observation 007edbd3-b00c-48af-8651-72a8d2f5f8a5 · outbound

This paper cites VLMPC: Vision-Language Model Predictive Control for Robotic Manipulation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making VLMPC: Vision-Language Model Predictive Control for Robotic Manipulation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:28.855409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:28.855409Z digest=sha256:a1ca2cef409d9f7ff1ebc67ed43cc4d2483d69d8c64742acc1c13525f1870940

Observation 02853fee-6955-4fdc-a7a1-b06257ea453f · outbound

This paper cites Reflective Planning: Vision-Language Models for Multi-Stage Long-Horizon Robotic Manipulation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Reflective Planning: Vision-Language Models for Multi-Stage Long-Horizon Robotic Manipulation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:28.983421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:28.983421Z digest=sha256:c2e75b978825177496c8f5f4a417fa6053233bee753b31b2df4e492d7cbd2835

Observation 883cd5eb-1dda-4aad-99f3-c20eb39e9c0c · outbound

This paper cites Open-world task and motion planning via vision-language model inferred constraints.arXiv preprint arXiv:2411.08253, 2024.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Open-world task and motion planning via vision-language model inferred constraints.arXiv preprint arXiv:2411.08253, 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.049705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.049705Z digest=sha256:3fc2d29419b8210e0b43073a767a326eea7aecbd3d3305150bfd522224490798

Observation 5343a5cc-08c4-407a-a1ba-15e190d79ea4 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.064263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.064263Z digest=sha256:643c225cc1eb9bc046c899ba0439bfd8791514e81a6f6b9458c621867f69cb12

Observation 4e079adf-50b6-4d9a-8d6b-0816a3c10146 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making RT-1: Robotics Transformer for Real-World Control at Scale

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.147844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.147844Z digest=sha256:78e48c9630600d41b592148b92e70f9c897af0390d7ccf1d9c94c9b6673965c2

Observation 38cca183-d0f7-48ce-8f4d-94170850d9d4 · outbound

This paper cites RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.206479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.206479Z digest=sha256:8c2527a5ae1a5360cd0cf422a198253ae4ba22efd54e496a15423428fffb5798

Observation 6508197f-7505-44fb-ab49-954eae7e8b86 · outbound

This paper cites Scaling proprioceptive-visual learn- ing with heterogeneous pre-trained transformers.Advances in Neural Information Processing Systems, 37:124420–124450, 2024.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Scaling proprioceptive-visual learn- ing with heterogeneous pre-trained transformers.Advances in Neural Information Processing Systems, 37:124420–124450, 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.283597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.283597Z digest=sha256:d7c2201d37ea2573a89af0dca0e7dfed7f415fb00537ba3b36b03391b1335f92

Observation cda072aa-ae8e-446f-91f5-afd6562b6005 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.343756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.343756Z digest=sha256:e93e4aa8d8c5fbdb9001d2651dece7b819360e38a324d03c4bc7b598c8244d2d

Observation 56a818c4-2a61-47a4-9e76-3aa21b770964 · outbound

This paper cites Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.426099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.426099Z digest=sha256:107d727c1b39aa43d7ebdb4a0bee7f1b1b33e414be9a2700e61d1fbfd51ba42c

Observation ba9b3811-0d51-4698-b11e-46816b7ede77 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making OpenVLA: An Open-Source Vision-Language-Action Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.508963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.508963Z digest=sha256:29942e4f169fa4d77188cea636ee4bb5284022905295d31fce73c0a2ad4a3c3e

Observation 960a02d4-e7c6-4bec-8684-ef955625343a · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.572917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.572917Z digest=sha256:ed1b0b8c83e3e4f4d75a5b051787ce05280d566005d6dbd7eb38cfbcb6a54a38

Observation 3c9762e3-034f-4bc4-9d83-b3e51d650d2b · outbound

This paper cites Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.614110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.614110Z digest=sha256:9a5af1f368e4c950f0c34e6be843afd1504508767d93916bbaee7a48968b68ce

Observation 0be1e207-5932-42b9-b829-029e1581e81e · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.690756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.690756Z digest=sha256:81ec5e8f460fd77c0c1d788d6bd3008b6acb7297c18e2e9e8ce2186ae73ef4d5

Observation 9a508b3a-2b41-4f7d-8f98-022b07172a94 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.709441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.709441Z digest=sha256:e1267b40ce59205a66a6752eb69cb7f25820b913fc1c49494b7e1f6956d63920

Observation 44cfb571-0f28-4941-8826-e53bf5b252d0 · outbound

This paper cites Manipulate-Anything: Automating Real-World Robots using Vision-Language Models.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Manipulate-Anything: Automating Real-World Robots using Vision-Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.714057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.714057Z digest=sha256:9df23df8d0f3c3d72d4b9c825f30c55b4be38276d163e72148b55450d4ac6561

Observation d63ef4ae-4dc7-4a58-9977-ef06d5db41d7 · outbound

This paper cites Skillman—a skill-based robotic manipulation framework based on perception and reasoning.Robotics and Autonomous Systems, 134:103653, 2020.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Skillman—a skill-based robotic manipulation framework based on perception and reasoning.Robotics and Autonomous Systems, 134:103653, 2020

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.718582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.718582Z digest=sha256:c726a5adf08967afb1ca793e4dda9fc7a53df112f0502803b5ff12a3a8de83da

Observation 680fdbc5-9dcc-42c6-94f7-1772805f9740 · outbound

This paper cites A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.723051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.723051Z digest=sha256:2e9c73abe34b3bcc3f1da7b17b5bfe339ed15c5c68c4068402ebf65bc71f15b0

Observation 3218ec8b-1e33-4470-9fbc-28ae5e08c84a · outbound

This paper cites Sofar: Language-grounded orientation bridges spatial reasoning and object manipulation.arXiv preprint arXiv:2502.13143, 2025.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Sofar: Language-grounded orientation bridges spatial reasoning and object manipulation.arXiv preprint arXiv:2502.13143, 2025

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.727266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.727266Z digest=sha256:08a1473422ce2715d1a0cc33f5cffa424a54572cfcbdbe9bb8377afa22750da9

Observation 0dc6f32f-0d3d-46c8-a402-d3623c0da8b5 · outbound

This paper cites GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.731232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.731232Z digest=sha256:f1b7ec2542a98137d39e4ac93fca34de81aa6d069b8870cef3099d228ce12abe

Observation 3423f687-670e-46b2-9fb0-066bba60553d · outbound

This paper cites KUDA: Keypoints to Unify Dynamics Learning and Visual Prompting for Open-Vocabulary Robotic Manipulation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making KUDA: Keypoints to Unify Dynamics Learning and Visual Prompting for Open-Vocabulary Robotic Manipulation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.734890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.734890Z digest=sha256:1a0223878b9fde3e27d2df7801b10eca4567ff50092ba0dcf4db92981a7cb66e

Observation 2e2bcb92-9f27-4b04-b0d1-e5669a48fce2 · outbound

This paper cites RoboGSim: A Real2Sim2Real Robotic Gaussian Splatting Simulator.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making RoboGSim: A Real2Sim2Real Robotic Gaussian Splatting Simulator

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.738813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.738813Z digest=sha256:cb1e51e941f9d9b57c36c8c80453c849ca0dada2b93de708bae2c97f11c666a0

Observation da8c1c21-aa6f-476c-9a6d-610b252029e3 · outbound

This paper cites RL-GSBridge: 3D Gaussian Splatting Based Real2Sim2Real Method for Robotic Manipulation Learning.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making RL-GSBridge: 3D Gaussian Splatting Based Real2Sim2Real Method for Robotic Manipulation Learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.743022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.743022Z digest=sha256:ec43cc885ee0e3a474ebb94e55e602e0e10c36aef0034ab522828734d162397b

Observation f3ef6282-fc0d-4948-b435-6dd960cd8532 · outbound

This paper cites Discovery and Deployment of Emergent Robot Swarm Behaviors via Representation Learning and Real2Sim2Real Transfer.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Discovery and Deployment of Emergent Robot Swarm Behaviors via Representation Learning and Real2Sim2Real Transfer

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.746981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.746981Z digest=sha256:843a16493d81cffb4f2facbe59f80b9f28386f89560984987ec9c3d306d4c431

Observation ff4c716f-cc6d-4892-9889-7aff6367eb9d · outbound

This paper cites Reconciling Reality through Simulation: A Real-to-Sim-to-Real Approach for Robust Manipulation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Reconciling Reality through Simulation: A Real-to-Sim-to-Real Approach for Robust Manipulation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.750649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.750649Z digest=sha256:f75162bcaf7a899f65a593c7fd735f111423be2c490e8329580f688991f89d67

Observation 202612d6-a7d4-44a6-bd43-0bc2841f8da3 · outbound

This paper cites Rl-vigen: A reinforcement learning benchmark for visual generalization.Advances in Neural Information Processing Systems, 36:6720–6747, 2023.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Rl-vigen: A reinforcement learning benchmark for visual generalization.Advances in Neural Information Processing Systems, 36:6720–6747, 2023

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.754394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.754394Z digest=sha256:a571bb27b9428b0bde52e290658ff00f78ed3da94db60ddf500f785c2b7d44dc

Observation 81530520-73f9-46c8-91f4-6b365c713e3a · outbound

This paper cites Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.757961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.757961Z digest=sha256:46e45bf190a61415dc6f279418a6fa58e9e512aebde65c1ce8a0451ee9e3c995

Observation 496335a6-57fe-43fc-b268-e855daf8b95b · outbound

This paper cites Efficient real2sim2real of continuum robots using deep reinforcement learning with koopman operator.IEEE Transactions on Industrial Electronics, 2025.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Efficient real2sim2real of continuum robots using deep reinforcement learning with koopman operator.IEEE Transactions on Industrial Electronics, 2025

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.761706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.761706Z digest=sha256:b28344bf61ecb19a06a59b857ce3d0fadfdcbd1555886cff58ed35b397621ea2

Observation ffbbd005-5e6d-4c1a-a83c-00dca03b16e1 · outbound

This paper cites Real-time per- ception meets reactive motion generation.IEEE Robotics and Automation Letters, 3(3):1864– 1871, 2018.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Real-time per- ception meets reactive motion generation.IEEE Robotics and Automation Letters, 3(3):1864– 1871, 2018

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.765399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.765399Z digest=sha256:fee6f3a23736f08ef6130a394f177b34afcc42bcbdcba8b6d78e68f04d70fef8

Observation f3f35635-23cf-4e22-9e02-b007767edf48 · outbound

This paper cites You Only Demonstrate Once: Category-Level Manipulation from Single Visual Demonstration.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making You Only Demonstrate Once: Category-Level Manipulation from Single Visual Demonstration

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.768816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.768816Z digest=sha256:a32e0237e6cc6342287e077f4d5851a6d15a53580cbb5c1361e1783cd26970e2

Observation f574453c-c930-4e62-9c01-d6bda8e08740 · outbound

This paper cites One-2-3-45++: Fast single image to 3d objects with consistent multi-view generation and 3d diffusion.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making One-2-3-45++: Fast single image to 3d objects with consistent multi-view generation and 3d diffusion

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.773481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.773481Z digest=sha256:a24d47dfffe9b7496c00383242c5850d0389a56d2a7eaf0ae512db08f6e1839c

Observation adbc6daf-938c-4315-bfd0-5d33f89eb69c · outbound

This paper cites Sparp: Fast 3d object reconstruction and pose estimation from sparse views.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Sparp: Fast 3d object reconstruction and pose estimation from sparse views

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.777217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.777217Z digest=sha256:8acb51cc9930cb15eccb78e4abdcc8618b89f1fd3eccdd01cc4f6a05118d196f

Observation a35b8401-4e81-4ef2-bf10-f1589bf1bded · outbound

This paper cites Zero123++: a Single Image to Consistent Multi-view Diffusion Base Model.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Zero123++: a Single Image to Consistent Multi-view Diffusion Base Model

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.780816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.780816Z digest=sha256:92180a37278d7c7b91a3626675e9ca360cdb7eecb64646eed8b6f1c395245a66

Observation dc0a08ef-866b-441d-8562-1e5ae190eb2a · outbound

This paper cites Zero-1-to-3: Zero-shot one image to 3d object.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Zero-1-to-3: Zero-shot one image to 3d object

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.784567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.784567Z digest=sha256:abe7719401cf1c6560ddec1fb8ff6fd40487aa61f4ae0b4e238469f2e1e67180

Observation ef4dcf57-55e1-491d-864f-80c272ffa6e7 · outbound

This paper cites Get3d: A generative model of high quality 3d textured shapes learned from images.Advances In Neural Information Processing Systems, 35:31841–31854, 2022.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Get3d: A generative model of high quality 3d textured shapes learned from images.Advances In Neural Information Processing Systems, 35:31841–31854, 2022

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.788055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.788055Z digest=sha256:ca7f8cc45e0efc2e88157e4a69e9855722ab2f52f585fc89ee5811484e193695

Observation 48bc8d56-a630-4be9-957c-79ecffa15e31 · outbound

This paper cites A-sdf: Learning disentangled signed distance functions for articulated shape represen- tation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making A-sdf: Learning disentangled signed distance functions for articulated shape represen- tation

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.791944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.791944Z digest=sha256:08191134778b6ba0fd2b13f996a970e6037ec8020fc5a2f02c14070cf0aadd01

Observation 7580dbdc-6cdc-4ed1-9099-4bee1a27c0e8 · outbound

This paper cites Ditto: Building digital twins of articulated objects from interaction.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Ditto: Building digital twins of articulated objects from interaction

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.796122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.796122Z digest=sha256:4b30e774e89b32f29b38708f585a9de6951b53273723082d8102eb8032586a1d

Observation fa0071f9-0ff6-4537-83a7-1eb67104ae13 · outbound

This paper cites Structure from Action: Learning Interactions for Articulated Object 3D Structure Discovery.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Structure from Action: Learning Interactions for Articulated Object 3D Structure Discovery

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.799748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.799748Z digest=sha256:d85939618117d4f66f34f1a0c4e61c299af72b84c4bfdda17f7141c52ae3297e

Observation c419f85c-0346-4210-87fb-e8a8ed36e284 · outbound

This paper cites URDFormer: A Pipeline for Constructing Articulated Simulation Environments from Real-World Images.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making URDFormer: A Pipeline for Constructing Articulated Simulation Environments from Real-World Images

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.803567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.803567Z digest=sha256:fa06814e91390fa87158cf59dbcce2fcff43d4633484c0f8a3af7181a09cbcfb

Observation 20fce2ed-0716-448f-a67f-d57b6c645845 · outbound

This paper cites Paris: Part-level reconstruction and motion analysis for articulated objects.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Paris: Part-level reconstruction and motion analysis for articulated objects

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.807593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.807593Z digest=sha256:0ef588d0488246edd9af801499e7da671d73cce735dc46795963a76988024728

Observation 495fe785-5f7c-426a-92ef-5734e6c41c7b · outbound

This paper cites Cage: controllable articulation generation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Cage: controllable articulation generation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.811629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.811629Z digest=sha256:ad7ae6a092e9a4af6ee12e61f6d900ac7043e5f57a80d6b3fa0697646e218ef6

Observation 0fcd9752-97bc-425d-ad36-2a43b6db519d · outbound

This paper cites Sam-6d: Segment anything model meets zero-shot 6d object pose estimation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Sam-6d: Segment anything model meets zero-shot 6d object pose estimation

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.816117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.816117Z digest=sha256:6d2de946eb439cbb3f6575f45350a667175628ea880b4a76afb45a34e8d161fe

Observation 40a5fc99-2b26-40a1-876d-f1d74922802f · outbound

This paper cites Gigapose: Fast and robust novel object pose estimation via one correspondence.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Gigapose: Fast and robust novel object pose estimation via one correspondence

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.819976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.819976Z digest=sha256:39a766155350e0ddcad785fdc2f77871f9b1901030682b1fe401af0fe7d940ea

Observation c38baffb-291d-483c-bcbb-d7590ebed77e · outbound

This paper cites Any6D: Model-free 6D Pose Estimation of Novel Objects.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Any6D: Model-free 6D Pose Estimation of Novel Objects

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.823718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.823718Z digest=sha256:fa730a91fa33fbdad6306e3729db791d6c7b6154de8695dcc3d5a07de88e2be6

Observation 3d8d1f3f-b4e2-4abb-9faa-061c4988e054 · outbound

This paper cites Foundationpose: Unified 6d pose estimation and tracking of novel objects.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Foundationpose: Unified 6d pose estimation and tracking of novel objects

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.827851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.827851Z digest=sha256:fc1c6d99b40425a35cf65d12c21b23f061e8f59b21c9778e03ceb19672b25d4b

Observation ef06b616-a511-492c-a899-22c81d92154f · outbound

This paper cites Foundpose: Unseen object pose estimation with foundation features.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Foundpose: Unseen object pose estimation with foundation features

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.831798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.831798Z digest=sha256:e91fbbc0b013832c84093e4495c596b89e0c0c4cc3545fe3800e382ff7abe6d7

Observation 175bfc2c-8433-43e3-80fc-0818e56a3ebf · outbound

This paper cites A comprehensive survey on point cloud registration.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making A comprehensive survey on point cloud registration

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.836289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.836289Z digest=sha256:5c6598e7a60991636dba28b0099e94ac1c1bbc30d7407abb58e667de517a8427

Observation 64a7666d-48b9-4c6c-a827-03d561bd4fd3 · outbound

This paper cites Deep Learning-Based Point Cloud Registration: A Comprehensive Survey and Taxonomy.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Deep Learning-Based Point Cloud Registration: A Comprehensive Survey and Taxonomy

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.840574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.840574Z digest=sha256:0818ae806f08d31a12868a4278e00b530c0f7c564df538233e91cbd37e0e8c09

Observation 4b60822b-1233-45b1-a150-d639ee99a26c · outbound

This paper cites A tutorial review on point cloud registrations: principle, classification, comparison, and technology challenges.Mathematical Problems in Engineering, 2021(1):9953910, 2021.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making A tutorial review on point cloud registrations: principle, classification, comparison, and technology challenges.Mathematical Problems in Engineering, 2021(1):9953910, 2021

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.844471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.844471Z digest=sha256:9df84ba07d9e739999bda11d88bb56ca5f67db3fecb079cec01bb6d348cb5bcd

Observation 6ad01dc7-36e0-4509-a66b-54ff429948b6 · outbound

This paper cites A comprehensive survey of visual slam algorithms.Robotics, 11(1):24, 2022.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making A comprehensive survey of visual slam algorithms.Robotics, 11(1):24, 2022

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.848687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.848687Z digest=sha256:96929e1b024fbbd8291beba92d780623f4869df011e9e81bb5cc7cea45790098

Observation 64a42648-e64e-4826-9f97-f59431350a19 · outbound

This paper cites How NeRFs and 3D Gaussian Splatting are Reshaping SLAM: a Survey.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making How NeRFs and 3D Gaussian Splatting are Reshaping SLAM: a Survey

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.852481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.852481Z digest=sha256:72cf4f108a3def52149be4d4d020c23bb4eea36b64a735435e6281c63b976b51

Observation 15c6e1a5-f68c-4fa6-832e-18218550afdc · outbound

This paper cites A survey on active simultaneous localization and mapping: State of the art and new frontiers.IEEE Transactions on Robotics, 39(3):1686–1705, 2023.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making A survey on active simultaneous localization and mapping: State of the art and new frontiers.IEEE Transactions on Robotics, 39(3):1686–1705, 2023

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.856398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.856398Z digest=sha256:b30bb6350817bd0e50ca1ed74156d7f456b8b4a5d4f7d9b8a7a0ed5a372d9708

Observation 533d08e4-5d2e-4e78-ad27-85348be21404 · outbound

This paper cites Scalable Real2Sim: Physics-Aware Asset Generation Via Robotic Pick-and-Place Setups.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Scalable Real2Sim: Physics-Aware Asset Generation Via Robotic Pick-and-Place Setups

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.859946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.859946Z digest=sha256:87755e224fc34e4fe86333930694a262b34e31c80c09c121f2efaed0d904059d

Observation 5b6072d7-adf1-4bd5-a78c-54b04a71c4d0 · outbound

This paper cites PhysTwin: Physics-Informed Reconstruction and Simulation of Deformable Objects from Videos.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making PhysTwin: Physics-Informed Reconstruction and Simulation of Deformable Objects from Videos

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.863672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.863672Z digest=sha256:50f52c667bfdba7689048a86427b95147d9337c6260721e954a807f028898c15

Observation daec14fd-34fd-4737-8cad-94abc14cff19 · outbound

This paper cites Sim2real 2: Actively building explicit physics model for precise articulated object manipulation.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Sim2real 2: Actively building explicit physics model for precise articulated object manipulation

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.867354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.867354Z digest=sha256:7aa66c6c8045409e17ec56a6979845bd7421513f3a27267d7056da1402a6a60d

Observation de7bba1a-3309-4736-8a1a-7dda6b800c41 · outbound

This paper cites A real2sim2real method for robust object grasping with neural surface reconstruction.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making A real2sim2real method for robust object grasping with neural surface reconstruction

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.870998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.870998Z digest=sha256:076a1a42e402737ed2e9682810679e293c41e9e9099362a4edd982b94ed7fc80

Observation b42e82ff-3951-4d4f-ae39-dffb6af3f622 · outbound

This paper cites SPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single Images.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making SPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single Images

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.874470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.874470Z digest=sha256:4ae43c1106a2c1bace597313bd9dbea2f708ffb85c01a40c8a0b19313b8fe2f9

Observation f01ede5d-9fa8-446c-bdc7-6d8f221cebb1 · outbound

This paper cites Pointllm: Empowering large language models to understand point clouds.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Pointllm: Empowering large language models to understand point clouds

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.878448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.878448Z digest=sha256:90f5fb8a3f802b84e0b0d1ef9dcd28a38065cdfe760a567d7969bc9b35a833f7

Observation a181ef1d-f192-48db-ae2a-fa10fe1ae6ef · outbound

This paper cites Point-nerf: Point-based neural radiance fields.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Point-nerf: Point-based neural radiance fields

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.882105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.882105Z digest=sha256:77d4235f286bb2eb5b5f56aa2946401277a2053007097c925692f76457d84bcc

Observation df7441fc-37c6-4163-94d6-9fcbeddb9643 · outbound

This paper cites Pointclip: Point cloud understanding by clip.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Pointclip: Point cloud understanding by clip

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.885459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.885459Z digest=sha256:8917031156730e17a1bbf0b88196562da69ee1f5576dd1928f7fa09fbb0c5f15

Observation 61384a05-d953-4a85-8b61-2d7e14236333 · outbound

This paper cites Text2nerf: Text-driven 3d scene generation with neural radiance fields.IEEE Transactions on Visualization and Computer Graphics, 30(12):7749–7762, 2024.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Text2nerf: Text-driven 3d scene generation with neural radiance fields.IEEE Transactions on Visualization and Computer Graphics, 30(12):7749–7762, 2024

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.889110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.889110Z digest=sha256:79fad051712e63894c98099f2b3564c9d040bca85bbc2a2d5a6fdf0c38c08fe5

Observation 6086479a-47b4-4b7c-9399-15e7e233b747 · outbound

This paper cites Pointr: Diverse point cloud completion with geometry-aware transformers.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Pointr: Diverse point cloud completion with geometry-aware transformers

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.893362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.893362Z digest=sha256:c4de5d74e711a0a5607817284880ead7a6f51285d06674d36616a666933feea7

Observation 613b9ae0-f12c-4977-9baa-8f78b495a719 · outbound

This paper cites V oxel set transformer: A set-to-set approach to 3d object detection from point clouds.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making V oxel set transformer: A set-to-set approach to 3d object detection from point clouds

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.897079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.897079Z digest=sha256:1145281a9e58a89dfdc21510b1ed33474d5b18470e00792072c0690d1de56b2f

Pith citing papers

No inbound Pith citation observations are available.