Pith. sign in

Paper Citation Record · LEDGER

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping

As of 7 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 2 inbound Pith citation observations for arXiv:2506.06535.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06535 v3

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:59:30.182455Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T06:18:55.243330Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T20:00:10.735031Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact1
  • verified fuzzy47
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 95d3747e-a435-4af7-819f-b34f2fe1c027 · outbound

This paper cites End-to-end train- able deep neural network for robotic grasp detection and semantic segmentation from rgb.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping End-to-end train- able deep neural network for robotic grasp detection and semantic segmentation from rgb

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:38.563723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:22.608398Z digest=sha256:0222a92b1ea3d331f653244f872f23ee5c6c16fdb4824ab31f45872366eeb629

Observation d5230111-0944-4869-aaa7-51d05ad76cda · outbound

This paper cites Hifi-cs: Towards open vocabulary visual grounding for robotic grasping using vision-language mod- els, 2024.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Hifi-cs: Towards open vocabulary visual grounding for robotic grasping using vision-language mod- els, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:38.399923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:22.722777Z digest=sha256:4b654250401cdd37fb01d921d7a08fe1eaf67838982f89bb4a49b5d860ea2379

Observation 1d6d9234-71cd-4c17-b891-6ab7bd47d82f · outbound

This paper cites RT-2: Vision-language-action models transfer web knowledge to robotic control.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping RT-2: Vision-language-action models transfer web knowledge to robotic control

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:38.284066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:22.879648Z digest=sha256:cdac715f2bcc09265fad512ebced0b7daccc39f595d61bc5388a40eceb2ad3c8

Observation 9b9661ce-8956-4d31-a4b7-792933afa43c · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:23.007621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:23.007621Z digest=sha256:13d846ca266a92d114264d9751e2b5c07370037e14e22a6b7afc1b7439072dd3

Observation 8804535b-9083-4b43-ae69-5573142dff35 · outbound

This paper cites Trust the PRoc3s: Solving long-horizon robotics problems with LLMs and constraint satisfaction.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Trust the PRoc3s: Solving long-horizon robotics problems with LLMs and constraint satisfaction

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:38.116592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:23.129115Z digest=sha256:b39db39d51c184b420ba4176da31a694a639117ba4b5f598c8d28c1f41cac4a3

Observation c9c66032-d282-4326-80ec-a5260fcb86e0 · outbound

This paper cites Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:38.002526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:23.245485Z digest=sha256:e731d624090850bed10582caf896d9366c60448f676fcb96c0ad34a41b619ef6

Observation 72636aec-7308-464e-8e8b-98f7078e2290 · outbound

This paper cites Jacquard: A large scale dataset for robotic grasp detection.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Jacquard: A large scale dataset for robotic grasp detection

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:37.873290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:23.383235Z digest=sha256:49602d92ca8a94b51e03f4200caa745442ee0081a21f68f8b4edc684a36d4c86

Observation 4449ba98-7be0-468f-8750-45f386955cb7 · outbound

This paper cites GraspMolmo: Generalizable Task-Oriented Grasping via Large-Scale Synthetic Data Generation.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping GraspMolmo: Generalizable Task-Oriented Grasping via Large-Scale Synthetic Data Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:23.566111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:23.566111Z digest=sha256:ea35ae3d053a014f2452b11bf77246ec9bf8b8725c0a961cc936370b71f33fc7

Observation c4544b63-71b8-4fb6-a098-b045da646004 · outbound

This paper cites Acronym: A large-scale grasp dataset based on simulation.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Acronym: A large-scale grasp dataset based on simulation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:37.729654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:23.734957Z digest=sha256:b2e3b3898e7136496dbc0e1248f474f1e69c1768bce039028a1b635c106964c0

Observation 0058bcd9-a88b-4582-98c9-09c158d40814 · outbound

This paper cites GraspNet-1Billion: A large-scale benchmark for general object grasping.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping GraspNet-1Billion: A large-scale benchmark for general object grasping

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:37.592702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:23.863902Z digest=sha256:997e293bc5aec6cf354d2a025bff7d2977706b1bda50b6a4869f19819790538d

Observation 80e38701-26c3-4df1-8cc3-b79c95fa4ef2 · outbound

This paper cites Physically grounded vision-language models for robotic manipulation.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Physically grounded vision-language models for robotic manipulation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:37.421739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:23.974250Z digest=sha256:36eb2919097ed8dee6ca1aed92381652d4f79048592c39dec3b23ef4b52dcfee

Observation 1deaa556-6bb4-422e-8f02-e061f21cff2c · outbound

This paper cites Rvt2: Learning precise manipulation from few demonstrations.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Rvt2: Learning precise manipulation from few demonstrations

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:37.245120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:24.079791Z digest=sha256:9ffde8042069e8ed06ebfbfa504b4ec4ae2c5502da22f017cb8a621db1b4d0ff

Observation 4377422f-956a-490b-9c78-932e5a1f2b3c · outbound

This paper cites Language-grounded dy- namic scene graphs for interactive object search with mobile manipulation.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Language-grounded dy- namic scene graphs for interactive object search with mobile manipulation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:37.075236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:24.256736Z digest=sha256:4bbc99caae96652178bb626ab90fcd72a258ee130af194a2398ac986d3d0feff

Observation 98abafa1-c906-46d5-95e6-45a8c9ce0719 · outbound

This paper cites Inner monologue: Embodied reason- ing through planning with language models.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Inner monologue: Embodied reason- ing through planning with language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:36.886801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:24.387555Z digest=sha256:20bde64b7204ccaa02afa596eda3a8cea73e47c797deed7662e3e1dc824a93bd

Observation e3a2a822-747f-408d-88b2-8dec101af597 · outbound

This paper cites Effi- cient grasping from rgbd images: Learning using a new rect- angle representation.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Effi- cient grasping from rgbd images: Learning using a new rect- angle representation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:36.701046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:24.502989Z digest=sha256:b99af0efa865314037631833bea795530623d95c507cbf77107bf2c13e6550a6

Observation ac6541b7-d66c-424f-9970-9ab555c89a33 · outbound

This paper cites Prismatic vlms: Investigating the design space of visually-conditioned language models.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Prismatic vlms: Investigating the design space of visually-conditioned language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:36.504139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:24.611566Z digest=sha256:ef496a2b89d17d268ec680828c7510e410be51224ece03966c0596f431a4074d

Observation efc502b1-c795-4751-9c80-ca80a1a1160b · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping OpenVLA: An Open-Source Vision-Language-Action Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:24.715654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:24.715654Z digest=sha256:5565abdb4e4a0de303ff47b7a6ebf13b052636d4696b5e0028d3a0027c6622d2

Observation 64924a4c-7d2e-4595-8831-24af304a44b5 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:24.756115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:24.756115Z digest=sha256:d8492ef7a194ea08b4e3f664bc3a30a0b9aad0dd180873b71b2aa6d75670e52d

Observation 8acab195-1467-4563-a6c8-7b772ae19fca · outbound

This paper cites Segment anything.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Segment anything

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:36.320177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:24.859095Z digest=sha256:5c62408cff7aefd51edae2cdc78b329c69e25591c96e71ba46de925fee1736d8

Observation 6a96d6ae-0ffe-4182-b7f6-2a1a65e858ec · outbound

This paper cites Antipodal robotic grasping using generative residual convolutional neu- ral network.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Antipodal robotic grasping using generative residual convolutional neu- ral network

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:36.135015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:24.973178Z digest=sha256:d515f716584b7d6d84fe84dd8fb1c3bb69f660953d962fc149eae578d4239be8

Observation 8d15859f-698d-4b35-8eca-405fcc3b0c96 · outbound

This paper cites Ovgnet: A unified visual-linguistic framework for open-vocabulary robotic grasping.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Ovgnet: A unified visual-linguistic framework for open-vocabulary robotic grasping

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:35.950064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:25.044852Z digest=sha256:3c0194ffa5475be8524781182456bc22c9d52e8b2bc3459776e1d4367c09d1f4

Observation a8950cf3-5ee1-4482-b786-7007bb0e4524 · outbound

This paper cites Ovgnet: A unified visual-linguistic framework for open-vocabulary robotic grasping.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Ovgnet: A unified visual-linguistic framework for open-vocabulary robotic grasping

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:35.713433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:25.210319Z digest=sha256:88ffd32d2349265cbfa8429157b918f0d1a52b1d6594a66bf2eb44e353229d68

Observation 31909571-58c9-41a6-b582-3ab879eff332 · outbound

This paper cites Vision-language foun- dation models as effective robot imitators.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Vision-language foun- dation models as effective robot imitators

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:35.488906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:25.336778Z digest=sha256:2c8984ce2efca98e9399c8af6e14c0f248e0deaeb3c48f26c87e2ad7766dd1b5

Observation 15acdb14-72b2-4a3b-9598-749606a6de8d · outbound

This paper cites an unresolved cited work.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:59:35.267215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:25.480480Z digest=sha256:3789313f84e8abf2734bde96dbd09c40e857585ec3dd31e14f536ea2e2d030d7

Observation 0c3965e4-eaeb-458b-8cd7-3bf30b0f60a7 · outbound

This paper cites OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:25.594969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:25.594969Z digest=sha256:f6ed812e2e1b22311950509617ba1328d5bbb7d4dcf8c53402adea3d5228b702

Observation 4e0eb91f-eaf0-42a7-ac46-1c547ca68b6e · outbound

This paper cites Grounding dino: Marry- ing dino with grounded pre-training for open-set object de- tection.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Grounding dino: Marry- ing dino with grounded pre-training for open-set object de- tection

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:35.045677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:25.698166Z digest=sha256:63a11c9f4142a4ba5e6c73b2b34614f0e410a870b283f1a9f34ac8d05265a6a8

Observation 04bbb5e5-bf10-45f8-b658-e3ea8faf660c · outbound

This paper cites Gao, Xi Vin- cent Wang, and Lihui Wang.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Gao, Xi Vin- cent Wang, and Lihui Wang

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:34.791071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:25.822106Z digest=sha256:4edf7d283f92fd990c6a3027ccf0b115d3a481b7dbe42d01f38649224d88625b

Observation 916c6a14-fd01-4e8d-bafa-bbb4aa0e0f7f · outbound

This paper cites Deepseek-vl: Towards real-world vision- language understanding, 2024.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Deepseek-vl: Towards real-world vision- language understanding, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:34.642108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:25.948312Z digest=sha256:43080d8a00193efe5f645f7e1b8348699e3902be21c7aa19f4d80677401f5def

Observation 2d5e4f87-565e-4bc3-8ba6-19c355082e09 · outbound

This paper cites Hybrid physical metric for 6-dof grasp pose detection.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Hybrid physical metric for 6-dof grasp pose detection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:34.476271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:26.088842Z digest=sha256:701ff86183c14c82e37de54c3fd0e9604e8f59b242b0253c31027d338eccdc5d

Observation bde316c3-db5a-4528-8d4e-2d604cd6a3e7 · outbound

This paper cites Vl-grasp: a 6-dof interactive grasp pol- icy for language-oriented objects in cluttered indoor scenes.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Vl-grasp: a 6-dof interactive grasp pol- icy for language-oriented objects in cluttered indoor scenes

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:34.294109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:26.196122Z digest=sha256:bd00fd17e5190a2dae24e029cd1a8b9366aa25f2282480f41669bd6ed8ca1ccd

Observation 7f678d07-53ca-4bb0-b1bd-5ad1f2bd9083 · outbound

This paper cites GraspMamba: A Mamba-based Language-driven Grasp Detection Framework with Hierarchical Feature Learning.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping GraspMamba: A Mamba-based Language-driven Grasp Detection Framework with Hierarchical Feature Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:26.297329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:26.297329Z digest=sha256:8485709cf55bdbecab2cdee7be64f3b037ee8b1366ddd49c4d849cb037fa48e1

Observation 1ff06100-d67a-4dd8-920d-72a05779c937 · outbound

This paper cites Lightweight language-driven grasp detection using conditional consis- tency model.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Lightweight language-driven grasp detection using conditional consis- tency model

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:34.143028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:26.411855Z digest=sha256:4bc32070cdca62c387c89531d4dd10db105c1c4b953ab93daae7777d839e05e5

Observation e4d91950-9375-4366-97b8-dc1d6c167b02 · outbound

This paper cites Language-driven 6-dof grasp detection using negative prompt guidance.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Language-driven 6-dof grasp detection using negative prompt guidance

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:33.962764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:26.526176Z digest=sha256:e8860590caf608a7176c76d570f833377f35a6710806b87638b16f5a6224f431

Observation 4382b3af-21e9-4b5c-9450-43b3bd7c59be · outbound

This paper cites GraspSAM: When Segment Anything Model Meets Grasp Detection.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping GraspSAM: When Segment Anything Model Meets Grasp Detection

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:26.631101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:26.631101Z digest=sha256:f93f9f969a169433783b64278a23e8e9c2217f47f4287b10786c242539b490a4

Observation 84da3673-2a32-4842-8123-6cac36c08c2b · outbound

This paper cites 3D-MVP: 3D Multiview Pretraining for Robotic Manipulation.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping 3D-MVP: 3D Multiview Pretraining for Robotic Manipulation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:26.711288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:26.711288Z digest=sha256:3b2410dab66b1bd964f28cc489770540717e03634a013f4ae411afa4b7d5e0fc

Observation ad675a5d-fbb1-496c-a3a1-c096f658dc72 · outbound

This paper cites Sayplan: Ground- ing large language models using 3d scene graphs for scalable task planning.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Sayplan: Ground- ing large language models using 3d scene graphs for scalable task planning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:26.875867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:26.875867Z digest=sha256:cf19b25efba8658fec9e7052ae4cba13cba24d6b13c8812b28d02622fcd01501

Observation 874fc6f0-5cf7-43b7-a473-b25a0c20b72b · outbound

This paper cites SAM 2: Segment anything in images and videos.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping SAM 2: Segment anything in images and videos

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:33.836594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:26.961017Z digest=sha256:7f3395f4789d03b9da763912ecbcce5ad8c551624467e7d801c8cbe0b767e288

Observation ba039e28-f965-4619-a1ee-96657594658f · outbound

This paper cites Sadler, Wei-Lun Chao, and Yu Su.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Sadler, Wei-Lun Chao, and Yu Su

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:27.082005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:27.082005Z digest=sha256:4f2ea5a615a012c3792858eb7ed881e974954e474091429e67f5c66cf6a7c9e3

Observation 90556185-0a7f-4436-8081-5fa6887a1a1d · outbound

This paper cites Grasping in the wild: Learning 6dof closed- loop grasping from low-cost demonstrations.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Grasping in the wild: Learning 6dof closed- loop grasping from low-cost demonstrations

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:33.685490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:27.203026Z digest=sha256:8ce7cabffb48314e5b20ad4c7325587e8bdc10a8c795f4b9de6189f2d2aceeca

Observation aeb9fe25-0073-4a06-bcf6-a018aabd0836 · outbound

This paper cites Contact-graspnet: Efficient 6-dof grasp gen- eration in cluttered scenes.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Contact-graspnet: Efficient 6-dof grasp gen- eration in cluttered scenes

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:33.460345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:27.307342Z digest=sha256:0dbf39b54e45522c8095727428944c00a2e4c711d926272ede4ebf70f6c3a7a7

Observation 1ae8b650-b462-4bc5-ae8c-32693fb1adb5 · outbound

This paper cites Foundationgrasp: Generalizable task-oriented grasping with foundation models.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Foundationgrasp: Generalizable task-oriented grasping with foundation models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:33.269395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:27.457409Z digest=sha256:4992e1ddb1b888fd806ad0f3e07b71c83e3604675957b8c46aafd93988a229f5

Observation 01eadda2-9a00-4ffc-834c-36c28cbcae26 · outbound

This paper cites Mlp-mixer: An all-mlp ar- chitecture for vision.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Mlp-mixer: An all-mlp ar- chitecture for vision

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:33.161225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:27.584816Z digest=sha256:7cdf6f43f570a417e5d5a24f3ca7ecb186cfa1da2e8b5802a956154e4b771f91

Observation ffc1c13a-1cc4-4606-bcef-97a24d047f52 · outbound

This paper cites Towards open- world grasping with large vision-language models.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Towards open- world grasping with large vision-language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:32.918336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:27.686069Z digest=sha256:05c6c7a2ec4b7397324109abeb1f50058ba466dfea04ed383ba3833c6623ca11

Observation 6be11674-9bb7-44c0-8bee-7a94c658a422 · outbound

This paper cites Language-guided robot grasping: Clip-based referring grasp synthesis in clutter.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Language-guided robot grasping: Clip-based referring grasp synthesis in clutter

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:32.698032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:27.827578Z digest=sha256:2da222e4f4aafb53d46112cd4bb21cbf9cc422865ed15a316b83ec7fe4e7fa3d

Observation 43d95ba8-604b-4d8d-80a3-acabd87a550a · outbound

This paper cites Language-driven grasp de- tection with mask-guided attention.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Language-driven grasp de- tection with mask-guided attention

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:32.596186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:27.949297Z digest=sha256:f4f90cc15e0b6053e53d40ee2658636d0f8948ed80c19e82f6058444e471df0a

Observation 6e81fb7e-166a-43db-88ca-6573b096d7f6 · outbound

This paper cites an unresolved cited work.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:59:32.492609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:28.038044Z digest=sha256:514720fcfd06bf545c0227ef76a593921a4ab2dc8d9c43a39584b123599534d9

Observation 641a213e-9eda-4e76-a227-60a72c8335e4 · outbound

This paper cites an unresolved cited work.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:59:32.368506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:28.161716Z digest=sha256:62f4e4b363d3202c0bd52364a122b713256073af6184734c39d02e1485b336e3

Observation 1af2a83d-7d1c-480e-95d7-9343529fd530 · outbound

This paper cites Gpt-4v(ision) for robotics: Multimodal task planning from human demonstration.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Gpt-4v(ision) for robotics: Multimodal task planning from human demonstration

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:32.271887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:28.290693Z digest=sha256:0630d20ff51a0fd33a7c98c5f4074a355c885bb0382ae60416bc0591e7ea993c

Observation 1a56d105-9007-4492-b106-b077948e6287 · outbound

This paper cites Grasp as you say: Language-guided dexterous grasp genera- tion.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Grasp as you say: Language-guided dexterous grasp genera- tion

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:32.171500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:28.362358Z digest=sha256:36a425b7946a6e76f1297d6b64501053f802d29f739c73a3d240ff63c13784c9

Observation 294416c8-393b-4a14-829c-747a29c4fe17 · outbound

This paper cites Tidybot: Personal- ized robot assistance with large language models.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Tidybot: Personal- ized robot assistance with large language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:32.105335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:28.524037Z digest=sha256:c0d38e730314dff939937aa36f1afee89a82f3210615025d7a79e195225b829e

Observation 3ef0dfe0-0ff4-4228-8fc4-dc9fe8b725c4 · outbound

This paper cites A joint modeling of vision-language-action for target-oriented grasping in clutter.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping A joint modeling of vision-language-action for target-oriented grasping in clutter

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:32.069134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:28.668792Z digest=sha256:a1ebc38398ca55672e43f300f1d1f18371f3ce7094c650b62f2b9a596d31bca5

Observation 6f45701b-fde2-4643-8fcf-3e20fe3915a1 · outbound

This paper cites Instance-wise Grasp Synthesis for Robotic Grasping.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Instance-wise Grasp Synthesis for Robotic Grasping

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:59:30.485178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:28.814172Z digest=sha256:e05d66b0c8d937a551ddf60c9945320edda62269a3c93029f20236f023ee616f

Observation e1fa16b9-38dc-463b-b2b2-f403246a2b57 · outbound

This paper cites Universal instance perception as object discovery and retrieval.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Universal instance perception as object discovery and retrieval

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:32.030130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:28.954588Z digest=sha256:e53a6f8769af379098ccc9de6c0cf81c30a887f9993a4fafd6ed9e11d5aebcc0

Observation 9f23039d-fbc9-44a0-934d-a1738bf971a6 · outbound

This paper cites Ground4act: Leveraging visual-language model for collaborative pushing and grasping in clutter.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Ground4act: Leveraging visual-language model for collaborative pushing and grasping in clutter

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:31.998882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:29.021121Z digest=sha256:7ad059d73d769c9d9a4245ab020672aba47c5029af4ee25989e67f632b5833f8

Observation 2470b129-8cad-4f0c-a086-f2b7fc481499 · outbound

This paper cites A Parameter-Efficient Tuning Framework for Language-guided Object Grounding and Robot Grasping.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping A Parameter-Efficient Tuning Framework for Language-guided Object Grounding and Robot Grasping

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:29.189968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:29.189968Z digest=sha256:56a74ccc4019a11a6f37bc9870b87c9208d74f42186ca4d00741cf08f9642a70

Observation ae0cf24b-1a09-40af-8434-edc2065200ef · outbound

This paper cites Se-resunet: A novel robotic grasp detection method.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Se-resunet: A novel robotic grasp detection method

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:31.973274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:29.307333Z digest=sha256:c1f631ac35abe79040ca968da817b4b246a4568b562b79fe2c049f4595e4526a

Observation 7396091c-f77e-448b-adbd-dbc59406fb45 · outbound

This paper cites GLiNER: Generalist model for named entity recognition using bidirectional transformer.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping GLiNER: Generalist model for named entity recognition using bidirectional transformer

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:31.887794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:29.380199Z digest=sha256:819e3e1a966dc1c6c7a1725d9f9f318d1679865650d56414a1d63dab51acdde4

Observation 0434e761-0c96-4fda-9833-7318bb7ef5d8 · outbound

This paper cites Sigmoid loss for language image pre-training.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Sigmoid loss for language image pre-training

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:31.742129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:29.522314Z digest=sha256:927d2e5746d790028837355597bc6b74fc17d4f5b337c30b3780c7feddf9e69e

Observation 98014d88-fc38-407f-8456-6851306eaac8 · outbound

This paper cites Roi-based robotic grasp detection for object overlapping scenes.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Roi-based robotic grasp detection for object overlapping scenes

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:31.517765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:29.630514Z digest=sha256:edcabebd2ee887964fc4aa485cd84079627821a49d2f83fd890c7ba8f4d52c28

Observation 6c734f9e-ea84-4ac8-a2e5-80d9356fc318 · outbound

This paper cites EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:29.719733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:29.719733Z digest=sha256:4d5c4827c697acf14d24bb7db83c9814e41f92c92905cb6808d627600dd83eee

Observation 37f2c474-a2c0-495b-ac14-5ce4b1aac357 · outbound

This paper cites Language-guided cat- egory push–grasp synergy learning in clutter by efficiently perceiving object manipulation space.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Language-guided cat- egory push–grasp synergy learning in clutter by efficiently perceiving object manipulation space

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:31.296924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:29.848669Z digest=sha256:2ea3449de45439c81cfac2a81f50830645a3f7eac2c6ed96ed3ce4199dfe48f6

Observation ea22c6f7-ecd9-47cf-af37-89e8b7bf70da · outbound

This paper cites Vlmpc: Vision-language model pre- dictive control for robotic manipulation.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Vlmpc: Vision-language model pre- dictive control for robotic manipulation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:31.051808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:29.952184Z digest=sha256:ea44b94e66163139ac0cacd8e4b2a0396d4e3a0b2099195eec451576c78839a9

Observation e4fa7e13-3078-47fa-8767-b217232b7b8e · outbound

This paper cites Grasping detection network with uncertainty estima- tion for confidence-driven semi-supervised domain adapta- tion.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping Grasping detection network with uncertainty estima- tion for confidence-driven semi-supervised domain adapta- tion

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:59:30.848781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:59:30.182455Z digest=sha256:e8bb937e1736794763fa1a04a4664575875a34c266658e2bf3cf224c5985d0e7

Pith citing papers

Observation 953c6cb1-efc4-4e48-9d4e-ea5a098e3913 · inbound

BOP-ASK: Object-Interaction Reasoning for Vision-Language Models cites this paper.

BOP-ASK: Object-Interaction Reasoning for Vision-Language Models MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:00:10.738442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T19:58:19.309634Z digest=sha256:88821d671e5d05969b5c1339e888fa01e773a75c4e1f57a755da83a2a643a0a7

Observation 405b64ad-7cef-4384-bd3d-931f7f1737c9 · inbound

Data Pyramid for Embodied Manipulation cites this paper.

Data Pyramid for Embodied Manipulation MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T06:18:55.243330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:18:55.243330Z digest=sha256:257c151c973db7ade61bd01521733302abe02257db7ef89bb327d3674937dbea