Pith. sign in

Paper Citation Record · LEDGER

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task

As of 9 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2605.00963.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.00963 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-09T18:37:32.865795Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact5
  • verified fuzzy22
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 183148c1-d87f-4995-936b-ad28412c9836 · outbound

This paper cites An advanced medical robotic system augment- ing healthcare capabilities-robotic nursing assistant.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task An advanced medical robotic system augment- ing healthcare capabilities-robotic nursing assistant

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.203037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:70ee4237adb7c13b79a535daf88ada10e3c87549f6d2af41cc90a656cec137df

Observation 4c39b7fb-c675-4a5e-b346-a92b0075a63a · outbound

This paper cites A human-robot interac- tion applicution based on augmented reality (ar) for industrial robot grasping process.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task A human-robot interac- tion applicution based on augmented reality (ar) for industrial robot grasping process

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.192863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:ba5f081e336ccd6a4469f9f7b3b4e8181181a313b4b90093439c3f59aa323389

Observation 5f767e74-d661-4836-97ee-93e2fa331a0d · outbound

This paper cites An educational robot system of visual question answering for preschoolers.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task An educational robot system of visual question answering for preschoolers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.196266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:2df84701f40a4858c48222ca1b98183878c82cca48f802c0414338c42c0971d7

Observation 1b97e164-789b-43b4-bb42-d7ad824ecdbb · outbound

This paper cites Home robot service by ceiling ultrasonic locator and microphone ar- ray.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Home robot service by ceiling ultrasonic locator and microphone ar- ray

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.199644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:bba5bae4b2152d74df95fc5108383893ec7913906ecd0d47822c404c9862b02a

Observation 4cb39b36-ef97-4a59-97be-d4ef7086790e · outbound

This paper cites The human intention: a taxonomy attempt and its applications to robotics.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task The human intention: a taxonomy attempt and its applications to robotics

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.209649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:c9615531ff072b22bb9b8de229825e9af463acadf0cf0f73fab2857a69b653a6

Observation dad20ba3-845b-467a-8972-462403d80010 · outbound

This paper cites Anticipatory robot control for efficient human-robot collaboration.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Anticipatory robot control for efficient human-robot collaboration

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.212910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:b4e8cfc2fb07122a86aa854eb42500e2b426978e850783fae50f2cabc13becac

Observation 2497d62a-9f5e-4f98-8efc-6ffa89c4794e · outbound

This paper cites Pointing gestures for human-robot interaction with the humanoid robot digit.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Pointing gestures for human-robot interaction with the humanoid robot digit

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.175250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:cc58d4f8957553e70e0d0107d31ec4d41c583be18cb199f04ada2e593425151b

Observation b4cf76bc-5a15-4b13-bf98-bd7e7f99f3ed · outbound

This paper cites Autonomous laparoscopic robotic suturing with a novel actuated suturing tool and 3d endoscope.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Autonomous laparoscopic robotic suturing with a novel actuated suturing tool and 3d endoscope

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.178563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:188c60aab1764799f6490ad2af1fc3c578b59cd97825a8ce5839c1b511ce790d

Observation 57376c0c-cd56-4d10-b263-a654bf42da18 · outbound

This paper cites Perception– intention–action cycle in human–robot collaborative tasks: the col- laborative lightweight object transportation use-case.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Perception– intention–action cycle in human–robot collaborative tasks: the col- laborative lightweight object transportation use-case

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.182051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:b082e6981672e6b59e520fc4030d3bbc3b0b9de21e6452bd90fe7a7738bbc1d3

Observation 72930273-a607-48c6-8b7a-fa1bf6619a00 · outbound

This paper cites Exploring transformers and visual transformers for force prediction in human-robot collaborative transportation tasks.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Exploring transformers and visual transformers for force prediction in human-robot collaborative transportation tasks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.185616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:e562a7843a5176bacc835e3db4483627a2aa8da6a05280739fa3cf90a66c6c95

Observation 0bae2e8b-63e6-4330-afbc-a4e404d56cb2 · outbound

This paper cites Force and velocity predic- tion in human-robot collaborative transportation tasks through video retentive networks.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Force and velocity predic- tion in human-robot collaborative transportation tasks through video retentive networks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.167802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:afcb1a4ef6475986f948ff0689a5d9591061b8e9b2c8a70e627118248627e72a

Observation f2372dc3-f71f-4b7e-afbe-ae0f8c7f8912 · outbound

This paper cites Language and sketching: An llm-driven interactive multimodal multitask robot navigation framework.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Language and sketching: An llm-driven interactive multimodal multitask robot navigation framework

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.160638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:ac5396b7ecf2a8d1f979904cda5aa93d05ff52a3d0fa924069002dd5a5996268

Observation 22cf6f80-e16a-4957-ba16-d774eb3240a8 · outbound

This paper cites Interactive navigation in environments with traversable obstacles using large language and vision-language models.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Interactive navigation in environments with traversable obstacles using large language and vision-language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.164292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:d062b2ed95f4b59d6e17f352a3d1c088f84340116bdeb1743ea123244e097a5e

Observation 8ed54e66-2d45-4925-a56b-e6218c2ba471 · outbound

This paper cites Physically grounded vision-language models for robotic manipulation.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Physically grounded vision-language models for robotic manipulation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.171128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:872d96b88e496bc0ef9424c079fdbc92fc91ebe20c192dd931432c969e0732fe

Observation 030fc9cd-05b7-426a-b924-e44f6f4c3c36 · outbound

This paper cites When the inference meets the explicitness or why multimodality can make us forget about the perfect predictor.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task When the inference meets the explicitness or why multimodality can make us forget about the perfect predictor

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.189369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:d1ee3eb28426dd0d63e58ce32427645be9d11ea36fd68062f384eb758d7f3595

Observation f8fd0236-cd83-4c0c-ae9e-ac6ae973bd37 · outbound

This paper cites Anticipation and proactivity. unraveling both concepts in human-robot interaction through a han- dover example.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Anticipation and proactivity. unraveling both concepts in human-robot interaction through a han- dover example

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.206291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:d82ba015df7dcb1a6839dd2c08bc84038f69b43e3722c3674c41838aef28fce8

Observation f2799689-edbc-4285-8b07-4686b201b390 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task LLaMA: Open and Efficient Foundation Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:11:08.303610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:e75b100e5616713c7af9cfe6ce26bd544b3f3e7548522b53cb0960297c247e71

Observation 3b09c363-9f43-462d-82f8-163f0228952d · outbound

This paper cites GPT-4 Technical Report.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task GPT-4 Technical Report

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:11:08.332019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:f3c8ba65fb2c9be450592e3d099604af8e9b4400ef27c5af81c700d2cb550296

Observation de45a080-e331-48cc-8793-67ce3e8e4c70 · outbound

This paper cites Leveraging Large Language Models in Human-Robot Interaction: A Critical Analysis of Potential and Pitfalls.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Leveraging Large Language Models in Human-Robot Interaction: A Critical Analysis of Potential and Pitfalls

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:11:08.312260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:7e57eb591ec9b74dcd3bdffbdabf596b809b36b79f168455cae201d62cc0a2dc

Observation 2d6f11dc-a3a9-4e3e-9ad5-b08864f22db5 · outbound

This paper cites Florence-2: Advancing a unified representation for a variety of vision tasks.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Florence-2: Advancing a unified representation for a variety of vision tasks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.216378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:1bcc4979f5545463e352bb90e30cd0e53a4d08b1771e486f7f83ec1abeacb492

Observation 648ebf6d-0598-46ac-96bf-602e3686f29a · outbound

This paper cites Robust speech recognition via large-scale weak super- vision.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Robust speech recognition via large-scale weak super- vision

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.219864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:973c1b37b81c959b10dbdf7f7a5e447229925323686d7eba447e3d4d9f0e6a23

Observation 1e947acc-fc63-438d-920b-2950bf62b7eb · outbound

This paper cites AST: Audio Spectrogram Transformer.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task AST: Audio Spectrogram Transformer

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:11:08.290066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:01287f6747982d36630de014562cfb4370cf95f600d70653aa64d0d89b653d16

Observation b3a39b44-fbaa-42af-a3df-96d3d1081215 · outbound

This paper cites Fuzzy logic systems for engineering: a tutorial.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Fuzzy logic systems for engineering: a tutorial

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.149913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:2fc98e10b00d6bbb3bc69e8cb7ab3285520ee70b1401d42aa9a678ec0e393d57

Observation 3c8abc05-8cd9-455e-82ed-4a69708aef0f · outbound

This paper cites Interval type-2 fuzzy logic systems: theory and design.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Interval type-2 fuzzy logic systems: theory and design

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.153373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:5634ffd9084b07ff1eb8a7cd7082028cac4c0c0a29fce9cd86d0ad118502458c

Observation 8491ac5d-1db2-4bc8-8f56-ca3ac55856c1 · outbound

This paper cites Fuzzy logic introduction.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task Fuzzy logic introduction

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.156900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:713d00124372e1f795113a1b82a7f0b965d89988fdac063eaf879d8c0d695e21

Observation b35d1674-42f5-4558-8681-a1df0b965274 · outbound

This paper cites A type-2 fuzzy logic controller for autonomous mobile robots.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task A type-2 fuzzy logic controller for autonomous mobile robots

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T22:57:14.146744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:e03d1b2956d57fe3332af83ab64c1d189be04fbe18d139d1c70597d475af370e

Observation d5daa156-1f5f-444d-8b37-f23969b1d5f3 · outbound

This paper cites An approach to combining video and speech with large language models in human-robot interaction.

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task An approach to combining video and speech with large language models in human-robot interaction

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:11:08.325436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:37:32.865795Z digest=sha256:26a71ddae5f6984c6da97e4411a48360ced0a7f7d26cfbc1fb294cc4f13ea624

Pith citing papers

No inbound Pith citation observations are available.