Pith. sign in

Paper Citation Record · LEDGER

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models

As of 8 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 0 inbound Pith citation observations for arXiv:2506.19498.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19498 v1

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:12:09.155546Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

79 of 79 outbound references displayed

  • verified exact2
  • verified fuzzy20
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f7dcb54d-4c62-4aef-8593-774b4343b1d4 · outbound

This paper cites GPT-4 Technical Report.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.006778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.006778Z digest=sha256:0d2a82265eb62750fb91663b5f90f0de21910c645280ebb62c7c5f19f0407364

Observation 21aac602-01f4-46c6-9983-57593bd09a5e · outbound

This paper cites ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.047648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.047648Z digest=sha256:e9604b3aa44a26e53a3e4d38549a4d25afb07324876a0ec18caedebcfad8fde6

Observation a8fc755f-0a12-4267-98cd-080e70e43da0 · outbound

This paper cites Copa: General robotic manipulation through spatial constraints of parts with foundation models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Copa: General robotic manipulation through spatial constraints of parts with foundation models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.746596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:03.087048Z digest=sha256:4c1b5bf01bfb904eb795e06902e602b67187030614c001488b68e41a095a2383

Observation 76538e9f-30b3-4326-9c65-c30a45354c5b · outbound

This paper cites OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.142366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.142366Z digest=sha256:ea6a61ba0984db207f65ba5c09ef0249225b67ffe55ab381acc545bfc18b972d

Observation 3cf9391e-3c58-41cf-aaeb-7e7b2b496281 · outbound

This paper cites GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.239306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.239306Z digest=sha256:84e8d5c1ff4b38cfad2f392bd9d726391e84882bcedd57323ea28e7fc4459bc2

Observation b8cd82c8-14e0-4c5a-9b39-57dd03ce7e62 · outbound

This paper cites Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.288372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.288372Z digest=sha256:920b2c06928708a7ea5e4892918381f68d7b43cdf4a50b5e19320f8a6b72270c

Observation 976c396b-5b79-49c7-92ee-e309c1a1840f · outbound

This paper cites Guiding Long-Horizon Task and Motion Planning with Vision Language Models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Guiding Long-Horizon Task and Motion Planning with Vision Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.390138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.390138Z digest=sha256:e349f4bf474dbcaf2065c3ac1b6b6c8742de396840f538bd7dce37a5e34b1fbc

Observation 2760d490-e208-4e5d-ab2d-3edd5e53b8ba · outbound

This paper cites Open-world task and mo- tion planning via vision-language model inferred constraints.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Open-world task and mo- tion planning via vision-language model inferred constraints

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.448540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.448540Z digest=sha256:d4cdfded48b0badff1fa519656915b6f5939a89166a4d0a9a46ea3d56f631ad7

Observation 11c21ccd-35e1-438d-bb10-5eca6d4f2576 · outbound

This paper cites Physically grounded vision-language models for robotic manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Physically grounded vision-language models for robotic manipulation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.567384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.567384Z digest=sha256:c539cbf2097bb32e658a7af020fb394c353be265bdbbc0e01dd45ead2407cfa1

Observation dfc156ef-59c4-42dc-bcb1-abd204d73f86 · outbound

This paper cites Vlm see, robot do: Human demo video to robot action plan via vision language model.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Vlm see, robot do: Human demo video to robot action plan via vision language model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.613132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.613132Z digest=sha256:13003bd0c800b35cdc2eb7420952ea2d4859752569be20dabd3850ba6fde2c36

Observation 729f7f9d-31dd-4201-8883-8523075fde40 · outbound

This paper cites Sofar: Language-grounded orientation bridges spatial reasoning and object manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Sofar: Language-grounded orientation bridges spatial reasoning and object manipulation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.665705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.665705Z digest=sha256:1dee512062f76d360f3fda4d638d0b017faa29f10e2f60a943c3387bac3bae38

Observation 853f45df-a4e2-43cf-b81b-ee8f622a8357 · outbound

This paper cites VLMPC: Vision-Language Model Predictive Control for Robotic Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models VLMPC: Vision-Language Model Predictive Control for Robotic Manipulation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.746754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.746754Z digest=sha256:e59bfce0f9b7fc483fdf863995e227d929af919bb2eda0d07c9bd1897039f6c2

Observation a4dba97b-2dcf-4f91-b96b-a501ea8e86d5 · outbound

This paper cites KUDA: Keypoints to Unify Dynamics Learning and Visual Prompting for Open-Vocabulary Robotic Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models KUDA: Keypoints to Unify Dynamics Learning and Visual Prompting for Open-Vocabulary Robotic Manipulation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.801070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.801070Z digest=sha256:aeeab22f9923f6145ae1e244da03f348540d385f1160dbf75c76a699fa03c670

Observation 580133e2-e101-42c7-85c3-fa14c6b5f52e · outbound

This paper cites VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.885154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.885154Z digest=sha256:a1f805c3dde24ad55b7ce4629480a37feafaa5af39fa9b68595c4839c7695ad7

Observation 995d07c8-f01e-4779-be89-6c9c304d5a5a · outbound

This paper cites MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.958565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.958565Z digest=sha256:603b893616c9468251ee213071c48a06383e958d8e4777d5b7c3037adcc15450

Observation b1375379-f124-451c-aa10-bf2537dc0f4e · outbound

This paper cites A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.998541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.998541Z digest=sha256:7f2a1d574376219967bb7c3ad00876fcce24d8fc5011392dea1c8f3034393ba7

Observation 9dafe183-77f9-43a7-b1fb-f0083b08dbf6 · outbound

This paper cites Learning to interpret natural language commands through human-robot dialog.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Learning to interpret natural language commands through human-robot dialog

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.730699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:04.066584Z digest=sha256:b201ec9f89eff1f9285919448a8ec695381e44147c33a63480219536b43ebb7d

Observation aab21090-a360-4a35-9eac-4579e90a30e1 · outbound

This paper cites Grounding verbs of motion in natural language commands to robots.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Grounding verbs of motion in natural language commands to robots

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.719822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:04.142753Z digest=sha256:133bb04cc7b11bb43e8a04aa92d09e1108f873906349ed24cd50b3039a5684af

Observation 386300ee-6e29-40fe-9074-074470b7b85f · outbound

This paper cites Toward understanding natural language directions.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Toward understanding natural language directions

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.709353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:04.190234Z digest=sha256:b61f2240d78d50fa76662a11f78571b70b9e9b495647b33f92320f2740ab2290

Observation c3121184-f8a4-4d0c-b472-90cd1def4285 · outbound

This paper cites Understanding natural language commands for robotic navigation and mobile manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Understanding natural language commands for robotic navigation and mobile manipulation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.699711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:04.271113Z digest=sha256:2699b8658ed4b6903a50b44b705bfee7e0db575897981f1e2dc277d23cc6b187

Observation ef9b0120-ba87-41fe-8042-ff94900e6787 · outbound

This paper cites TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.326924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.326924Z digest=sha256:f689fdabbfd3747ad3eaa7bad46acf7286a98bca5a3d47494a57ba66f9120ee5

Observation 0f291216-b1a5-457e-a4a0-125447ba092c · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.421149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.421149Z digest=sha256:c92f60fafab54c6fa87f796f0d4f351d073d244bacefd1784254493333ab4d05

Observation eba4b8fc-77ba-492f-91d1-011bb97464b5 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.499840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.499840Z digest=sha256:24be03588a07245095a8d9e0b41545732d5d3d4778c5aae45f8641bc8aebdc5e

Observation 863da663-d828-452c-a62a-6c80cc8bc02e · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.566817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.566817Z digest=sha256:bfc16f863d4b6e8f5f7e401071a5dfbcbcc6471f48ea824e449edff8ba0a1a70

Observation 683d88e5-4fa9-4883-b4cf-83956b5d0b60 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.641939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.641939Z digest=sha256:95d328559d9c0072ec3d7686f630cbdf3bbbfc7a520bdb03b3e669db3feff1f6

Observation 59868033-286d-4dd9-8f01-30a6cca9c25a · outbound

This paper cites RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.680233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.680233Z digest=sha256:542498de31641f1ba817244ae8a3a09958c12f6f5943617b0082fd75c4f33512

Observation 4fccaf42-5f7c-4145-b561-acea518e2579 · outbound

This paper cites ALOHA 2: An Enhanced Low-Cost Hardware for Bimanual Teleoperation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models ALOHA 2: An Enhanced Low-Cost Hardware for Bimanual Teleoperation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.770431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.770431Z digest=sha256:36940f2d719561a932380885d8beb7ab0802f19895c1b03797765cc99ba319b9

Observation 173e3d0b-562e-4c94-860b-56bd7c878c87 · outbound

This paper cites Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.843906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.843906Z digest=sha256:7fb53672b53e2b049bde9ea87f840a130b535fdb0a971da38193828f153447c7

Observation bac650b5-f260-44c4-a0b8-30472031daa4 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.884890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.884890Z digest=sha256:2e330dcc7b1e96c2fefd78cca64ed1d7bbdc2afe154220ab3411bc3e27921c86

Observation 1bb37cf4-a892-44fe-9a4d-6c3904cccbdd · outbound

This paper cites 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.988783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.988783Z digest=sha256:d4a6b1c6dca508dca25cfc4a09558d98a59c103450428da671f3a3d26d013bf1

Observation e9f30bba-37a7-4640-9ca3-92b965550e7e · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Octo: An Open-Source Generalist Robot Policy

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.061262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.061262Z digest=sha256:c946bd2b1d11adce759114f38249a27c54c4c7da8ebc0ffe46f51334bfa8700d

Observation 74e98e85-7c9f-4150-862e-1c6478143aa8 · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.111915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.111915Z digest=sha256:e08bf21f60043a1b2cba3c93b976e7ec686293abe27654637c81b48df607bcb4

Observation 6d3f8af4-8d51-4a1a-a2ee-f9d3d15fa295 · outbound

This paper cites RoboFlamingo-Plus: Fusion of Depth and RGB Perception with Vision-Language Models for Enhanced Robotic Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models RoboFlamingo-Plus: Fusion of Depth and RGB Perception with Vision-Language Models for Enhanced Robotic Manipulation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.207574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.207574Z digest=sha256:fae903bc1d30005f2ba2b44fe2bc676bdb557aa75682169c1a43343970e493bb

Observation 0019b4a4-5b4e-4be1-a351-23cdb4da1fd4 · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Vision-Language Foundation Models as Effective Robot Imitators

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.276267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.276267Z digest=sha256:2324b2798762373a82c7a0698ab7c53b178c65510bbe73968061c655592bf80e

Observation 431fd637-b19d-4a66-8bf9-2668e5200083 · outbound

This paper cites HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.316266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.316266Z digest=sha256:64e4bb9e7f8f4fee7e7b2a9a995119588866039e8de4f66f0a4fbb3a353580a7

Observation 42b8a48b-fb42-4115-9dd3-18e67ce0d8f7 · outbound

This paper cites CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.418519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.418519Z digest=sha256:15e6e4ef2ebae41ad7495685ee4ea1b84af87af13a0776010038181de1033c69

Observation 5a081d2e-86fc-45f3-801d-ec30236c942d · outbound

This paper cites PointVLA: Injecting the 3D World into Vision-Language-Action Models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.478000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.478000Z digest=sha256:5910bba2d888e32e39355e33b1d00ae986da83b3b9e0117c5a13ac14dd173323

Observation 5053b701-0997-4d04-bd89-610c97133537 · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.527284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.527284Z digest=sha256:d804f041a32605197cc9a16f583e4d1431432dacdc5be806a4a6913231f78c60

Observation 792276d5-a035-444e-b92e-4bacef746649 · outbound

This paper cites Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.638223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.638223Z digest=sha256:e7506fbf884a165b5001f825aadf40662f33ea6acd6e98a5d4b89ba319665888

Observation 9e993ad7-d19a-4958-a9e1-6f1068b6278c · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.690842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.690842Z digest=sha256:e251b1fdbe0771affded6b7be5c49eea3fed078f255b8980d2baf8069b997fbe

Observation 5ebc3e1a-ec6a-40e4-8276-dbf108045ac7 · outbound

This paper cites EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.781091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.781091Z digest=sha256:425f1139e43015a4ac8ccefd755dc86dd10b7700cda277f1e3e487df2fc9db77

Observation c8e34603-f4c2-4037-8cca-a920b2473925 · outbound

This paper cites LLM+P: Empowering Large Language Models with Optimal Planning Proficiency.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models LLM+P: Empowering Large Language Models with Optimal Planning Proficiency

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.850037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.850037Z digest=sha256:4792a7946a0d4eb62c00d9d8815de9a8ddc3e188c0d2f4a0203e7c282c164942

Observation 948a6dba-6d7e-4e96-a0c6-8a10ee1607c8 · outbound

This paper cites Llm-planner: Few-shot grounded planning for embodied agents with large language models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Llm-planner: Few-shot grounded planning for embodied agents with large language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.893446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.893446Z digest=sha256:dbb0f3d86cf0aae109df9f77aade5b55b2d95da84af727442ce057e3310283d4

Observation e3447b32-c4e8-4a2d-b955-4ed8aa480917 · outbound

This paper cites Code as policies: Language model programs for embodied control.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Code as policies: Language model programs for embodied control

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.953173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.953173Z digest=sha256:fb4b7f1c4022f6a06b4bab6630a50b92a284099c16b7563efd2bbdc3188680aa

Observation e5024913-6ebf-4446-a471-4ceca1687d6b · outbound

This paper cites Progprompt: Generating situated robot task plans using large language models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Progprompt: Generating situated robot task plans using large language models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.677317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:06.080619Z digest=sha256:3eb33d0f8599ea3aa3590ee877abb665f98e1f940268021e823a6e9b96d6f820

Observation 782b5a8f-d56f-4d43-85d6-09706f9e0715 · outbound

This paper cites Chatgpt for robotics: Design principles and model abilities.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Chatgpt for robotics: Design principles and model abilities

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.667052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:06.222417Z digest=sha256:c5c1c2188454885f57ff5ea5832493ebdcbc2a3626a8a41373445d4efadfeaf4

Observation 549d0afc-d638-467f-9554-76eafc693ad5 · outbound

This paper cites Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:06.341930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:06.341930Z digest=sha256:670ec213f5ee7f382028299046b78607b65844482a87bd48ec10223a051637e7

Observation 33ee52c1-a2c9-4f89-9348-2d41e098b838 · outbound

This paper cites Foundation models defining a new era in vision: a survey and outlook.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Foundation models defining a new era in vision: a survey and outlook

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:06.449159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:06.449159Z digest=sha256:08e38ba514ee2f21784dc46b8c32e7ba628b86cb2dd6b4c7d00e504139bb6c9b

Observation 4f3a9ea5-f4d4-4d98-b2e7-3ce1934b8957 · outbound

This paper cites Yolov10: Real-time end-to-end object detection.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Yolov10: Real-time end-to-end object detection

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.513413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:06.591235Z digest=sha256:12301047f3ef23206ca765651a7f8940cb417122eb8bb2778091691bf42d8808

Observation 762187bf-5d7a-4126-8efa-83b428afc9f3 · outbound

This paper cites YOLOv12: Attention-Centric Real-Time Object Detectors.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models YOLOv12: Attention-Centric Real-Time Object Detectors

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:06.730254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:06.730254Z digest=sha256:07c0eb2d34a07b1e55a944e099aed29d4f8fc46c0e286a832dd1cfd4fffc2209

Observation 10cdafa9-2d9b-4743-913f-e11df4e5d3cf · outbound

This paper cites Yoloe: Real-time seeing anything.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Yoloe: Real-time seeing anything

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:06.841797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:06.841797Z digest=sha256:6bd0579a6118091fabc54aa10c55b5248ceb03706ccbc5ab8493ea9e4fb7f032

Observation 06330ba7-25e0-43ef-9875-729e32e584f9 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:06.986608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:06.986608Z digest=sha256:d81f00c4080f4a36d6931276ed0d64736574d78d59269d79a8a99263be0c625f

Observation cb635328-fa44-4d6b-81b1-06ef8efcb96f · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models SAM 2: Segment Anything in Images and Videos

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:07.109157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:07.109157Z digest=sha256:a385655b41f6f58c7bdbf28b4326c97e649158a04a739bf92a4888b837bc8a77

Observation 8b365f20-6304-4397-ab46-0b0f8a3e1c87 · outbound

This paper cites Fast Segment Anything.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Fast Segment Anything

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:07.220820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:07.220820Z digest=sha256:5108a5e59fad64cfebd99aae191977fd67433c952a36746e5b393d50f309e3c2

Observation e99c5c4b-e05b-42e4-aebb-d10330b70b74 · outbound

This paper cites Segment everything everywhere all at once.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Segment everything everywhere all at once

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:07.344271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:07.344271Z digest=sha256:b8104c68857af1bdb260591c71539c4f3eb78d51045cd901e3bb005a6fa37ed5

Observation 96c81ed4-59b5-4a9c-a761-ad09e2dac5ba · outbound

This paper cites kpam: Keypoint affordances for category-level robotic manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models kpam: Keypoint affordances for category-level robotic manipulation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.351801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:07.470139Z digest=sha256:d70113805c6a0693b8ce5c7c122afa07c05c3d1df3497e2c26ee9b6defbcd281

Observation c2077bc1-821e-4c72-8444-ed19d164cdc9 · outbound

This paper cites Any-point Trajectory Modeling for Policy Learning.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Any-point Trajectory Modeling for Policy Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:07.568520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:07.568520Z digest=sha256:a541335ce0d388cbf81dc643f9f2f116498d1d34e01039e5b4a0bf62bb7dc44c

Observation 8da321ad-6616-4dbe-a41e-20a59c43a8cc · outbound

This paper cites PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:07.661869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:07.661869Z digest=sha256:2bfbfdf2f415e627c51011fe06de66121f0c45703cf1fa03bd05b5bb6499eb58

Observation 8b477562-60e4-4212-88fb-6c53e5443b26 · outbound

This paper cites Foundationpose: Unified 6d pose estimation and tracking of novel objects.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Foundationpose: Unified 6d pose estimation and tracking of novel objects

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.113415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:07.861016Z digest=sha256:4396b25aaccfea275e147ee6ab06dacd44998a38b891beadeea2f1b90cd2240b

Observation ee1f49c4-f2d0-423e-a5db-1332efd13cba · outbound

This paper cites Sam-6d: Segment anything model meets zero-shot 6d object pose estimation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Sam-6d: Segment anything model meets zero-shot 6d object pose estimation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:11.842969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:07.911112Z digest=sha256:1d37bd07188ca50776fba81dd841c971312a26437974b4966894f7fd431644dc

Observation a807a592-c2bf-4450-a059-72ad9a73c895 · outbound

This paper cites Omni6dpose: A benchmark and model for universal 6d object pose estimation and tracking.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Omni6dpose: A benchmark and model for universal 6d object pose estimation and tracking

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:11.687810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:07.937920Z digest=sha256:bdcf7f72bb410e7fc0f1b01c56db7d61ae25c8b3948f9de96bedad699f2f7604

Observation b6cd504a-d784-4ee8-b767-f44c04bce151 · outbound

This paper cites Gen6d: Generalizable model-free 6-dof object pose estimation from rgb images.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Gen6d: Generalizable model-free 6-dof object pose estimation from rgb images

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:11.485692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:08.021674Z digest=sha256:a846c984e77f344f784572886a6a0a1e57ffe5c4fda33a3b463bf5fa76cb5812

Observation b9fca9de-91cb-4823-9fcd-ca78b24839be · outbound

This paper cites Onepose: One-shot object pose estimation without cad models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Onepose: One-shot object pose estimation without cad models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:11.329092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:08.096588Z digest=sha256:af9f96acdfbb70591d5a500467e7def3d37d48a04b68b74d2a4c67c56ea2096e

Observation fa9de27d-3787-4f29-a4f2-df1ca297a8f3 · outbound

This paper cites Onepose++: Keypoint-free one-shot object pose estimation without cad models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Onepose++: Keypoint-free one-shot object pose estimation without cad models

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:11.119342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:08.158463Z digest=sha256:55bfaa8005bbe1ce66249f6ac00206ff7b6ad54047d3ce01b3e5c9ec0fbeea46

Observation 4b407ffb-5d6c-4cca-973e-98c757567aaa · outbound

This paper cites GS-Pose: Generalizable Segmentation-based 6D Object Pose Estimation with 3D Gaussian Splatting.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models GS-Pose: Generalizable Segmentation-based 6D Object Pose Estimation with 3D Gaussian Splatting

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:12:09.497856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:08.250575Z digest=sha256:56677f328ec52296fc4840bd04d6ef8b8f59528cdb4e834022c0b6cefb816fea

Observation 746ef3ae-d489-4e7a-ba7a-ef9107518ad2 · outbound

This paper cites You Only Demonstrate Once: Category-Level Manipulation from Single Visual Demonstration.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models You Only Demonstrate Once: Category-Level Manipulation from Single Visual Demonstration

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.280921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.280921Z digest=sha256:0cb4e59e7e4f1a837dd4c6458ef16647c12ec4f8b51bde43980b8be56a4e4709

Observation f33928b8-72ce-43a3-a0f8-3f30a67c1475 · outbound

This paper cites SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.358863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.358863Z digest=sha256:dfdcd8d13fda24f4d32253c66911dc583b7a19210a501be39cf04452983efb22

Observation 6b33626f-3e6d-49ac-a012-b6b4d3b5ab12 · outbound

This paper cites VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.430220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.430220Z digest=sha256:7656b1cce53b1364a18891a47f8897bbfe09fcb24b12ca4a2bb36d7128a79646

Observation c570fe5d-f911-48a9-acee-dc952887e705 · outbound

This paper cites TSP3D: Text-guided Sparse Voxel Pruning for Efficient 3D Visual Grounding.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models TSP3D: Text-guided Sparse Voxel Pruning for Efficient 3D Visual Grounding

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:12:09.345486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:08.483436Z digest=sha256:5c0ac5e1d2ab5d022af8afdecb9a383a159e151e1bab33952d0896c5038675ac

Observation c5c84f78-1d05-4bdc-b997-dee6732cd87f · outbound

This paper cites Scanreason: Empowering 3d visual grounding with reasoning capabilities.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Scanreason: Empowering 3d visual grounding with reasoning capabilities

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:10.982349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:08.570568Z digest=sha256:ef65b37f20bdbb3249e92957e04a16cdcba7b0dee1f63e4d25f09cb48d923dcd

Observation 52372e6f-62d9-47c3-ba65-9551ac4a6178 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models DINOv2: Learning Robust Visual Features without Supervision

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.608592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.608592Z digest=sha256:7d9d52e4d0cbef51221cf7e172a9e21462eea973207dda0ac190ecf2c53aaf0b

Observation 3ce2b293-2739-4dba-8e62-d862d532f07d · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.686757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.686757Z digest=sha256:4efbacc58ee860ed8aaed0aaadc91c2c269bba0f8a565238318908aabb87b901

Observation 0a6daf92-271e-4862-b982-7d1c9c3cc231 · outbound

This paper cites SPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single Images.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models SPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single Images

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.750059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.750059Z digest=sha256:0e49e320409b15b3bc1ca5896ab2d0f2538c4c627db96ea1349ccb5df21ff1f9

Observation 53ca0041-7873-444e-ac7f-7068092a9c50 · outbound

This paper cites 9dtact: A compact vision-based tactile sensor for accurate 3d shape reconstruction and generalizable 6d force estimation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models 9dtact: A compact vision-based tactile sensor for accurate 3d shape reconstruction and generalizable 6d force estimation

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:10.785173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:08.827515Z digest=sha256:6028e687888d760ee3c31908e398aa8439d965340340a9fda5e28211c90dc2c7

Observation 76c4b5a8-b5b9-4cfe-99ab-e89056c77a1d · outbound

This paper cites Tac3D: A Novel Vision-based Tactile Sensor for Measuring Forces Distribution and Estimating Friction Coefficient Distribution.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Tac3D: A Novel Vision-based Tactile Sensor for Measuring Forces Distribution and Estimating Friction Coefficient Distribution

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.914602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.914602Z digest=sha256:ffa498f060bb1a57adc83e5ee5a0c121328ea1eeabd114df7fb6dc7b6142917e

Observation daadb4f5-6f63-488a-acb5-81076bc53dda · outbound

This paper cites Pointodyssey: A large-scale synthetic dataset for long-term point tracking.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Pointodyssey: A large-scale synthetic dataset for long-term point tracking

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:10.666620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:08.982750Z digest=sha256:7a1040b1d40016f7d37b11007b85a6e9390c667ef272da759a2970f2f7b104ef

Observation 191d4ab0-115c-4010-960f-05535a61a749 · outbound

This paper cites Cotracker: It is better to track together.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Cotracker: It is better to track together

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:09.050283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:09.050283Z digest=sha256:933d29e6f3cfa02994895c496608b350ba40c6bee477105fafe5a239e21ed60c

Observation ae492222-1f9a-42d4-81cf-495c70edff99 · outbound

This paper cites Robotap: Tracking arbitrary points for few-shot visual imitation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Robotap: Tracking arbitrary points for few-shot visual imitation

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:10.496351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:09.075532Z digest=sha256:b8a0e81bf73ca56f31bfa50a04bf96e48148b9cc1df333efa4e7338fc4c6c6d8

Observation 83a67dae-83d4-4528-88d5-e146ac0b2397 · outbound

This paper cites Recyclable.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Recyclable

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:10.345022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:09.155546Z digest=sha256:123ccc912fc053a5516bd231371c3edebadfa7d3dd598b68fc43b2a768b3f83d

Pith citing papers

No inbound Pith citation observations are available.