Pith. sign in

Paper Citation Record · LEDGER

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models

As of 8 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 0 inbound Pith citation observations for arXiv:2506.19498.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19498 v1

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:12:09.155546Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

79 of 79 outbound references displayed

  • verified exact2
  • verified fuzzy20
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f7dcb54d-4c62-4aef-8593-774b4343b1d4 · outbound

This paper cites GPT-4 Technical Report.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.006778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.006778Z digest=sha256:8a44cc5c5dae1e5d06462171e810091653689377f69172bd1d98fdb38166868b

Observation 21aac602-01f4-46c6-9983-57593bd09a5e · outbound

This paper cites ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.047648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.047648Z digest=sha256:027df1ac2897e3eb2e0705f3fe33e2dc677fb62056b1eab27df1ef9730c8b3f4

Observation a8fc755f-0a12-4267-98cd-080e70e43da0 · outbound

This paper cites Copa: General robotic manipulation through spatial constraints of parts with foundation models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Copa: General robotic manipulation through spatial constraints of parts with foundation models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.746596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:03.087048Z digest=sha256:a7c196e65f53706bf0da4423358c1402ea9a054fc1281b57939d2a019bc15372

Observation 76538e9f-30b3-4326-9c65-c30a45354c5b · outbound

This paper cites OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.142366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.142366Z digest=sha256:4e5812d6d6c3c224302365f750952d754b66b0bf218409d88df8973d1b1075b9

Observation 3cf9391e-3c58-41cf-aaeb-7e7b2b496281 · outbound

This paper cites GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.239306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.239306Z digest=sha256:e787dbfc7743a98ae293851a9b0dccb291e31246117c818fe1d4ab0a1773536b

Observation b8cd82c8-14e0-4c5a-9b39-57dd03ce7e62 · outbound

This paper cites Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.288372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.288372Z digest=sha256:3a14ea9c63f01a8ea19b8d15fdf0567a4b3c92eb3fdfcbf959786f3e7423c161

Observation 976c396b-5b79-49c7-92ee-e309c1a1840f · outbound

This paper cites Guiding Long-Horizon Task and Motion Planning with Vision Language Models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Guiding Long-Horizon Task and Motion Planning with Vision Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.390138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.390138Z digest=sha256:b68e74d12a76d22d699038f413738415b6286c77b95eff785acda2783f88ea03

Observation 2760d490-e208-4e5d-ab2d-3edd5e53b8ba · outbound

This paper cites Open-world task and mo- tion planning via vision-language model inferred constraints.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Open-world task and mo- tion planning via vision-language model inferred constraints

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.448540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.448540Z digest=sha256:a284b828a91d9684fa65664c90e669f028952e33f7b01524ff4db83bf88a7a0b

Observation 11c21ccd-35e1-438d-bb10-5eca6d4f2576 · outbound

This paper cites Physically grounded vision-language models for robotic manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Physically grounded vision-language models for robotic manipulation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.567384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.567384Z digest=sha256:5235d33b06ab5f1b073dc403aa55fd42bee2aef4d46991459d6abc4c0ed377d0

Observation dfc156ef-59c4-42dc-bcb1-abd204d73f86 · outbound

This paper cites Vlm see, robot do: Human demo video to robot action plan via vision language model.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Vlm see, robot do: Human demo video to robot action plan via vision language model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.613132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.613132Z digest=sha256:a77dcdbdc2708d924fdb04a4fe70deadd3af1032da4db89560832da46cb9c45f

Observation 729f7f9d-31dd-4201-8883-8523075fde40 · outbound

This paper cites Sofar: Language-grounded orientation bridges spatial reasoning and object manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Sofar: Language-grounded orientation bridges spatial reasoning and object manipulation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.665705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.665705Z digest=sha256:24bb8490307215f4a41e811d192a6be37e3db543454a88df481514aac97a0bdd

Observation 853f45df-a4e2-43cf-b81b-ee8f622a8357 · outbound

This paper cites VLMPC: Vision-Language Model Predictive Control for Robotic Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models VLMPC: Vision-Language Model Predictive Control for Robotic Manipulation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.746754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.746754Z digest=sha256:be2d6845c2bac133b67798a705c8dab21f57cf19994b2b6fba8ccb9fe2dd7f8c

Observation a4dba97b-2dcf-4f91-b96b-a501ea8e86d5 · outbound

This paper cites KUDA: Keypoints to Unify Dynamics Learning and Visual Prompting for Open-Vocabulary Robotic Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models KUDA: Keypoints to Unify Dynamics Learning and Visual Prompting for Open-Vocabulary Robotic Manipulation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.801070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.801070Z digest=sha256:664abf4e9b770e4dc4b5854f1e4ffb27d701c433448ac445800c9ed56a104c35

Observation 580133e2-e101-42c7-85c3-fa14c6b5f52e · outbound

This paper cites VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.885154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.885154Z digest=sha256:a043726f3de1bec0213695510f7814df15d8afb20c02b78f9a9166ef26df9f1d

Observation 995d07c8-f01e-4779-be89-6c9c304d5a5a · outbound

This paper cites MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.958565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.958565Z digest=sha256:101484592d48644e601373bbb8047a77697d72acf6102230541ce0745459c250

Observation b1375379-f124-451c-aa10-bf2537dc0f4e · outbound

This paper cites A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:03.998541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:03.998541Z digest=sha256:fbc6ee83e500a648a6a99e80347feb8a1b699bea894f6d13fe6811cc821f25be

Observation 9dafe183-77f9-43a7-b1fb-f0083b08dbf6 · outbound

This paper cites Learning to interpret natural language commands through human-robot dialog.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Learning to interpret natural language commands through human-robot dialog

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.730699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:04.066584Z digest=sha256:dde1fcf8cad2b006cdae34e926e3ab58ea2c83ebe621efed1cd280db27f1fca0

Observation aab21090-a360-4a35-9eac-4579e90a30e1 · outbound

This paper cites Grounding verbs of motion in natural language commands to robots.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Grounding verbs of motion in natural language commands to robots

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.719822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:04.142753Z digest=sha256:131914343f719e3acb8b5018ebff300f4c43ddaa9b6b9cc797537a5704a62400

Observation 386300ee-6e29-40fe-9074-074470b7b85f · outbound

This paper cites Toward understanding natural language directions.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Toward understanding natural language directions

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.709353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:04.190234Z digest=sha256:374051a13810bcc90482497d478ecbfe788909fe217849562926cd72b13531a2

Observation c3121184-f8a4-4d0c-b472-90cd1def4285 · outbound

This paper cites Understanding natural language commands for robotic navigation and mobile manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Understanding natural language commands for robotic navigation and mobile manipulation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.699711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:04.271113Z digest=sha256:4f3e3d94c63ca33061ec91af5c8792e2bb417f9582df472622d990daa39a2183

Observation ef9b0120-ba87-41fe-8042-ff94900e6787 · outbound

This paper cites TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.326924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.326924Z digest=sha256:dfbc0e541927b9679f73d051a7d2eef5e8de4b48eacaa8a41717d6d6d7b26513

Observation 0f291216-b1a5-457e-a4a0-125447ba092c · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.421149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.421149Z digest=sha256:85d9d2d9e44604cf4e4c95093a164172f11e527a3b27b0aebf14bb13d05a0b81

Observation eba4b8fc-77ba-492f-91d1-011bb97464b5 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.499840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.499840Z digest=sha256:2524dbd53ff038e12cdc7573d16d604dfa319957fde3e240fa18cc26a80c0162

Observation 863da663-d828-452c-a62a-6c80cc8bc02e · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.566817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.566817Z digest=sha256:2b80058737634d385b4ea1c8dee118e77d71e84c283d908e173f4450c2b8ce02

Observation 683d88e5-4fa9-4883-b4cf-83956b5d0b60 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.641939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.641939Z digest=sha256:4ec5242d59df9a13a57e0d163dda305d92d5464724665f57f3560d5a6acaefba

Observation 59868033-286d-4dd9-8f01-30a6cca9c25a · outbound

This paper cites RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.680233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.680233Z digest=sha256:50135575be9090bec7a715ee003d68984cc599c9fd42449a0f7861e50a821f43

Observation 4fccaf42-5f7c-4145-b561-acea518e2579 · outbound

This paper cites ALOHA 2: An Enhanced Low-Cost Hardware for Bimanual Teleoperation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models ALOHA 2: An Enhanced Low-Cost Hardware for Bimanual Teleoperation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.770431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.770431Z digest=sha256:58927cd8e13d8e5b8257ec0e210c5dabf1a0082fb2343b5b0d97e002ff9a09aa

Observation 173e3d0b-562e-4c94-860b-56bd7c878c87 · outbound

This paper cites Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.843906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.843906Z digest=sha256:f46b7a4ac0c04531dc358ef27aa02cda3316e7f0cc8519db44f39b5acf4ba79f

Observation bac650b5-f260-44c4-a0b8-30472031daa4 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.884890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.884890Z digest=sha256:377d0fecca207521515879ddb489b4ad85786cb943e7660be2b5653ede108058

Observation 1bb37cf4-a892-44fe-9a4d-6c3904cccbdd · outbound

This paper cites 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:04.988783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:04.988783Z digest=sha256:e068f6c707ce94c612ec342274d7723ed6c130424dfa5200413c6b181d5074ad

Observation e9f30bba-37a7-4640-9ca3-92b965550e7e · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Octo: An Open-Source Generalist Robot Policy

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.061262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.061262Z digest=sha256:362ee482bc9888b4c982bbab54c55c70a635a41d3657beeea9366b069da97eda

Observation 74e98e85-7c9f-4150-862e-1c6478143aa8 · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.111915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.111915Z digest=sha256:84c305ae9e3edab55484adc264e5029c31cf616ee65eabeccc16ffeac5e622be

Observation 6d3f8af4-8d51-4a1a-a2ee-f9d3d15fa295 · outbound

This paper cites RoboFlamingo-Plus: Fusion of Depth and RGB Perception with Vision-Language Models for Enhanced Robotic Manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models RoboFlamingo-Plus: Fusion of Depth and RGB Perception with Vision-Language Models for Enhanced Robotic Manipulation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.207574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.207574Z digest=sha256:98006e4ca69220c034ac94f00a13c5e37978d87217484f7f0f4025385978a4c1

Observation 0019b4a4-5b4e-4be1-a351-23cdb4da1fd4 · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Vision-Language Foundation Models as Effective Robot Imitators

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.276267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.276267Z digest=sha256:67d7f5085e8a2a092702af3d8b548a176270a864b2640807f10494916daafa81

Observation 431fd637-b19d-4a66-8bf9-2668e5200083 · outbound

This paper cites HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.316266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.316266Z digest=sha256:7faceca50ee2cae774f2a9b13c096ffa1f1a8e20922277876b010b4508bb2ffc

Observation 42b8a48b-fb42-4115-9dd3-18e67ce0d8f7 · outbound

This paper cites CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.418519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.418519Z digest=sha256:bbd35460cda075a1af7d3a9d175567b11b8254649f16ef00727203f37d49e4c4

Observation 5a081d2e-86fc-45f3-801d-ec30236c942d · outbound

This paper cites PointVLA: Injecting the 3D World into Vision-Language-Action Models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.478000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.478000Z digest=sha256:6b0bbb81848d965e2433739fa48baa2604eb521bb95f22dcb083994ae761b23c

Observation 5053b701-0997-4d04-bd89-610c97133537 · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.527284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.527284Z digest=sha256:ad911a476ecf0e3d224de51c81400fc6a4c20f9966c4c0b50926df9a03e28444

Observation 792276d5-a035-444e-b92e-4bacef746649 · outbound

This paper cites Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.638223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.638223Z digest=sha256:3286ab25d8e0bbd1a7ce6f0e98b55a7a36120b1c3645937152b8f8ad859de54a

Observation 9e993ad7-d19a-4958-a9e1-6f1068b6278c · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.690842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.690842Z digest=sha256:d6db50a1bfb878a8b75d02b923c0ad444f44d81b27726e73061d2d2726af574a

Observation 5ebc3e1a-ec6a-40e4-8276-dbf108045ac7 · outbound

This paper cites EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.781091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.781091Z digest=sha256:8ae4142256d1013f2d31f699423c1f8877264d84b169fb5ff6d7f24f797f8055

Observation c8e34603-f4c2-4037-8cca-a920b2473925 · outbound

This paper cites LLM+P: Empowering Large Language Models with Optimal Planning Proficiency.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models LLM+P: Empowering Large Language Models with Optimal Planning Proficiency

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.850037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.850037Z digest=sha256:3c4f78437d749be463d5db4d7124f58e13d33a22c0036b50ec9fd9753effe466

Observation 948a6dba-6d7e-4e96-a0c6-8a10ee1607c8 · outbound

This paper cites Llm-planner: Few-shot grounded planning for embodied agents with large language models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Llm-planner: Few-shot grounded planning for embodied agents with large language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.893446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.893446Z digest=sha256:9bb6ec3642cfe1fdbf9552ce79590ef5183111bf65a51a8fce95465ce64b0a19

Observation e3447b32-c4e8-4a2d-b955-4ed8aa480917 · outbound

This paper cites Code as policies: Language model programs for embodied control.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Code as policies: Language model programs for embodied control

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.953173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.953173Z digest=sha256:2b0ca5858987d5cca445ad72c889ef5277ba416e18d0c34dd641191d2f3d1e94

Observation e5024913-6ebf-4446-a471-4ceca1687d6b · outbound

This paper cites Progprompt: Generating situated robot task plans using large language models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Progprompt: Generating situated robot task plans using large language models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.677317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:06.080619Z digest=sha256:75c10aa8365ae61d875b905f0d9308915b2881f43bc2b9a5b032ccc7edb5faaa

Observation 782b5a8f-d56f-4d43-85d6-09706f9e0715 · outbound

This paper cites Chatgpt for robotics: Design principles and model abilities.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Chatgpt for robotics: Design principles and model abilities

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.667052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:06.222417Z digest=sha256:27a349e655638d2a590e4b134616edd8cd6f4eb5c973a4c665204de453eecc05

Observation 549d0afc-d638-467f-9554-76eafc693ad5 · outbound

This paper cites Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:06.341930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:06.341930Z digest=sha256:94c67e8e3bf74e3d28078a4eef1e0844110d48a87d55ba508fac491c5d61b44e

Observation 33ee52c1-a2c9-4f89-9348-2d41e098b838 · outbound

This paper cites Foundation models defining a new era in vision: a survey and outlook.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Foundation models defining a new era in vision: a survey and outlook

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:06.449159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:06.449159Z digest=sha256:f3b8da6176fc00548544cc408d7e0a2024098418241e90e11b8b872553a30373

Observation 4f3a9ea5-f4d4-4d98-b2e7-3ce1934b8957 · outbound

This paper cites Yolov10: Real-time end-to-end object detection.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Yolov10: Real-time end-to-end object detection

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.513413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:06.591235Z digest=sha256:b4b02d02739fad254320434c15b0ff0d94c2bdd1993f6dfde9c41c4680ecdf7b

Observation 762187bf-5d7a-4126-8efa-83b428afc9f3 · outbound

This paper cites YOLOv12: Attention-Centric Real-Time Object Detectors.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models YOLOv12: Attention-Centric Real-Time Object Detectors

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:06.730254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:06.730254Z digest=sha256:fbebe99ce4242052c9ebcdb19d0dc15f32d4d9d61d0f7214bb03d3e08593001e

Observation 10cdafa9-2d9b-4743-913f-e11df4e5d3cf · outbound

This paper cites Yoloe: Real-time seeing anything.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Yoloe: Real-time seeing anything

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:06.841797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:06.841797Z digest=sha256:63a4929373c7bdbe91fb8a63ccda82ff0a8c32cd87b1d323d0317c612349d685

Observation 06330ba7-25e0-43ef-9875-729e32e584f9 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:06.986608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:06.986608Z digest=sha256:4379ecc0a142d37a323256f2ce85ffc1ccce2e9d913a65eebf76d0145cc47d74

Observation cb635328-fa44-4d6b-81b1-06ef8efcb96f · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models SAM 2: Segment Anything in Images and Videos

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:07.109157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:07.109157Z digest=sha256:a8dc4995e1ca2689f2f74796cd10d7e6abee9af837de7650802159879a579628

Observation 8b365f20-6304-4397-ab46-0b0f8a3e1c87 · outbound

This paper cites Fast Segment Anything.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Fast Segment Anything

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:07.220820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:07.220820Z digest=sha256:b65056ce48856b10851c50d3210c09908672e89a076d48f1951eee4dc1a0bafe

Observation e99c5c4b-e05b-42e4-aebb-d10330b70b74 · outbound

This paper cites Segment everything everywhere all at once.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Segment everything everywhere all at once

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:07.344271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:07.344271Z digest=sha256:015696cee197164f1df775f400873925cc86d7e124463f1830391f1dfbfcd3f6

Observation 96c81ed4-59b5-4a9c-a761-ad09e2dac5ba · outbound

This paper cites kpam: Keypoint affordances for category-level robotic manipulation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models kpam: Keypoint affordances for category-level robotic manipulation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.351801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:07.470139Z digest=sha256:522657233b22a00cfbfce43ec2f59a657c0f6a79bdfb35e5d19d476b105de94f

Observation c2077bc1-821e-4c72-8444-ed19d164cdc9 · outbound

This paper cites Any-point Trajectory Modeling for Policy Learning.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Any-point Trajectory Modeling for Policy Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:07.568520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:07.568520Z digest=sha256:99fbdf4a027fbad6947838776c91e96cb608afb72fd40ebc691e15cd2ff42e64

Observation 8da321ad-6616-4dbe-a41e-20a59c43a8cc · outbound

This paper cites PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:07.661869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:07.661869Z digest=sha256:b13f9647e29ea5b83f01b188dc88cf8e28e45ecf5f3e5464224e9f2e56e86aa8

Observation 8b477562-60e4-4212-88fb-6c53e5443b26 · outbound

This paper cites Foundationpose: Unified 6d pose estimation and tracking of novel objects.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Foundationpose: Unified 6d pose estimation and tracking of novel objects

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:12.113415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:07.861016Z digest=sha256:6021b0283bcff3f31a7a4e8fdbd57066c341996f3a30b8c3e8f231a1f744a259

Observation ee1f49c4-f2d0-423e-a5db-1332efd13cba · outbound

This paper cites Sam-6d: Segment anything model meets zero-shot 6d object pose estimation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Sam-6d: Segment anything model meets zero-shot 6d object pose estimation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:11.842969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:07.911112Z digest=sha256:4dd20bfbeb61a94107e00d9e77e0a76db4d1f58cf8d3b1d9d7116044e2e95667

Observation a807a592-c2bf-4450-a059-72ad9a73c895 · outbound

This paper cites Omni6dpose: A benchmark and model for universal 6d object pose estimation and tracking.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Omni6dpose: A benchmark and model for universal 6d object pose estimation and tracking

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:11.687810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:07.937920Z digest=sha256:efab3d559a109e6ff7d9a7178f86bf75f8ab6054c863aa45e99badbfc9f8ecbc

Observation b6cd504a-d784-4ee8-b767-f44c04bce151 · outbound

This paper cites Gen6d: Generalizable model-free 6-dof object pose estimation from rgb images.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Gen6d: Generalizable model-free 6-dof object pose estimation from rgb images

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:11.485692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:08.021674Z digest=sha256:f488bf78c6f2766b6ad63d051c8dffb48c171dd6b779d1c115ddd494cfcee107

Observation b9fca9de-91cb-4823-9fcd-ca78b24839be · outbound

This paper cites Onepose: One-shot object pose estimation without cad models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Onepose: One-shot object pose estimation without cad models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:11.329092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:08.096588Z digest=sha256:370705c86ce3cec54742d83eff0d65f089dbcd199eaee13250fc43372f554a46

Observation fa9de27d-3787-4f29-a4f2-df1ca297a8f3 · outbound

This paper cites Onepose++: Keypoint-free one-shot object pose estimation without cad models.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Onepose++: Keypoint-free one-shot object pose estimation without cad models

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:11.119342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:08.158463Z digest=sha256:c9f6aa737ec77884099bafd6e464b0596fca94c76aa3a057e53b8dfc382a8063

Observation 4b407ffb-5d6c-4cca-973e-98c757567aaa · outbound

This paper cites GS-Pose: Generalizable Segmentation-based 6D Object Pose Estimation with 3D Gaussian Splatting.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models GS-Pose: Generalizable Segmentation-based 6D Object Pose Estimation with 3D Gaussian Splatting

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:12:09.497856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:08.250575Z digest=sha256:06e57064b5cd238c65d672a7db470f86ed21a116cb3a41b5a39b84504006705a

Observation 746ef3ae-d489-4e7a-ba7a-ef9107518ad2 · outbound

This paper cites You Only Demonstrate Once: Category-Level Manipulation from Single Visual Demonstration.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models You Only Demonstrate Once: Category-Level Manipulation from Single Visual Demonstration

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.280921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.280921Z digest=sha256:a33af44696f2c6609bdd35d2f19ff1f43df64724ba7ef87f2b9c665d82b84cca

Observation f33928b8-72ce-43a3-a0f8-3f30a67c1475 · outbound

This paper cites SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.358863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.358863Z digest=sha256:d41d4fd7f882628d0c45f88ee2b9f39d7493713da370944bdb68b8b8ec3f1100

Observation 6b33626f-3e6d-49ac-a012-b6b4d3b5ab12 · outbound

This paper cites VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.430220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.430220Z digest=sha256:809d6f92aa3e543c332f3a7c357b12bc313768d41b5caf5d936f4ff1c26d1f4b

Observation c570fe5d-f911-48a9-acee-dc952887e705 · outbound

This paper cites TSP3D: Text-guided Sparse Voxel Pruning for Efficient 3D Visual Grounding.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models TSP3D: Text-guided Sparse Voxel Pruning for Efficient 3D Visual Grounding

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:12:09.345486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:08.483436Z digest=sha256:b626e1b79af747cb710d0ba5b05322a8d05d04b2af33ab88172fecc1d3883d84

Observation c5c84f78-1d05-4bdc-b997-dee6732cd87f · outbound

This paper cites Scanreason: Empowering 3d visual grounding with reasoning capabilities.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Scanreason: Empowering 3d visual grounding with reasoning capabilities

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:10.982349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:08.570568Z digest=sha256:40f05b6d144b552eadc87abbc74b17c67303285657b80a1a09f29a46e0069318

Observation 52372e6f-62d9-47c3-ba65-9551ac4a6178 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models DINOv2: Learning Robust Visual Features without Supervision

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.608592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.608592Z digest=sha256:f42b8d9b4392197fa14f983f86f156b413c355612ee7abdde42a1b4d0c6bac61

Observation 3ce2b293-2739-4dba-8e62-d862d532f07d · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.686757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.686757Z digest=sha256:6c8203d8b9f64e6822ed8689b65bddc0a1b993b208b987e561af28c93350c190

Observation 0a6daf92-271e-4862-b982-7d1c9c3cc231 · outbound

This paper cites SPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single Images.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models SPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single Images

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.750059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.750059Z digest=sha256:1bb77b6426d87893a45d906ee54a3abc2446340694f2ebc284afa41e1c0ba629

Observation 53ca0041-7873-444e-ac7f-7068092a9c50 · outbound

This paper cites 9dtact: A compact vision-based tactile sensor for accurate 3d shape reconstruction and generalizable 6d force estimation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models 9dtact: A compact vision-based tactile sensor for accurate 3d shape reconstruction and generalizable 6d force estimation

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:10.785173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:08.827515Z digest=sha256:033bd9f4df5d8d6438c0a85476b0204450d04991f5349c9d2eaf3713080ee2e8

Observation 76c4b5a8-b5b9-4cfe-99ab-e89056c77a1d · outbound

This paper cites Tac3D: A Novel Vision-based Tactile Sensor for Measuring Forces Distribution and Estimating Friction Coefficient Distribution.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Tac3D: A Novel Vision-based Tactile Sensor for Measuring Forces Distribution and Estimating Friction Coefficient Distribution

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.914602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.914602Z digest=sha256:a3b2f69352e8f875f71e660df6d467c43aa0806acf73af05ab29d47545991f70

Observation daadb4f5-6f63-488a-acb5-81076bc53dda · outbound

This paper cites Pointodyssey: A large-scale synthetic dataset for long-term point tracking.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Pointodyssey: A large-scale synthetic dataset for long-term point tracking

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:10.666620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:08.982750Z digest=sha256:e04492e21b017a66e964f6ff978fe3e66da359a23cfd9fb51fb61351c5297cac

Observation 191d4ab0-115c-4010-960f-05535a61a749 · outbound

This paper cites Cotracker: It is better to track together.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Cotracker: It is better to track together

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:09.050283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:09.050283Z digest=sha256:f70a5492cdf9a7c8f95ef2d6d26b7feed809444c489430c4ff6e840982195230

Observation ae492222-1f9a-42d4-81cf-495c70edff99 · outbound

This paper cites Robotap: Tracking arbitrary points for few-shot visual imitation.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Robotap: Tracking arbitrary points for few-shot visual imitation

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:10.496351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:09.075532Z digest=sha256:218f4b99346b1962799ce5be9684e03ad90ca6c0890cba6d838f3ac0ff8f3a2d

Observation 83a67dae-83d4-4528-88d5-e146ac0b2397 · outbound

This paper cites Recyclable.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Recyclable

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:10.345022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:09.155546Z digest=sha256:bcff31e68f6f1a64b20b80df3f182a9789acc16220a6e2b4d01a8f1ee4772b2a

Pith citing papers

No inbound Pith citation observations are available.