Pith. sign in

Paper Citation Record · LEDGER

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech

As of 5 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 4 inbound Pith citation observations for arXiv:2503.16492.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.16492 v3

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-23T01:10:55.850274Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T19:39:07.075356Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T04:57:38.604855Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact2
  • verified fuzzy43
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a90f630d-ffd0-4942-a234-0f17297e9db1 · outbound

This paper cites Social robots in therapy and care.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Social robots in therapy and care

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.498346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:f533f6c29d2a76a84909f5cec2967358246a0bd76ac5a48f16a06f9e3e605373

Observation 2c5c96ee-84a7-45b9-b6a9-82fa5872d50b · outbound

This paper cites Jubileo: An open- source robot and framework for research in human-robot social interac- tion.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Jubileo: An open- source robot and framework for research in human-robot social interac- tion

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:12:22.128820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:f8f6be1cb82693139d10968a9dd6d8e93331d573a99a237785b96da6d7a15e80

Observation f34bb0eb-1fa4-4202-b70c-34b47f84741b · outbound

This paper cites Human-aware physical human–robot collaborative transportation and manipulation with multiple aerial robots.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Human-aware physical human–robot collaborative transportation and manipulation with multiple aerial robots

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:12:22.119264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:b0dcfcebebb42aa0c1fa41f6d236adb13cf5f4a9966064debbdbd97deabe78fe

Observation d98a3c25-71e5-4f90-b9c5-728f256fc094 · outbound

This paper cites Communicating human intent to a robotic companion by multi-type gesture sentences.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Communicating human intent to a robotic companion by multi-type gesture sentences

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:12:22.134707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:b4431025991056191f2dd4f7f5cd71a0ae3580c5d65472d214ace12b781216c0

Observation 5b6af122-e4a8-40b3-a9be-a587238033ad · outbound

This paper cites Nvp-hri: Zero shot natural voice and posture-based human–robot interaction via large language model.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Nvp-hri: Zero shot natural voice and posture-based human–robot interaction via large language model

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:12:22.122045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:aa8909aa55ae2bf4508daacb97323c05351987a6ba918d728220dd484db6c37c

Observation e82bc16b-485c-4eec-99fc-960afc86de6d · outbound

This paper cites Robot reading human gaze: Why eye tracking is better than head tracking for human-robot col- laboration.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Robot reading human gaze: Why eye tracking is better than head tracking for human-robot col- laboration

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:12:22.125007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:51920156abf819c16dc1e02ca5ed8a78993166763a1b9e1bd0dc7f1ad12e6e26

Observation 8b016c46-0c14-406b-9e6a-28c479563a7d · outbound

This paper cites A gaze-speech system in mixed reality for human-robot interaction.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech A gaze-speech system in mixed reality for human-robot interaction

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:12:22.115991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:89eef15edc70f5b922dca737927acaa03faf4b2edff071315c929d54cf540bc0

Observation de8dca7e-1ef4-48da-8c8b-9c8aebb24d00 · outbound

This paper cites Human–robot interaction through eye tracking for artistic drawing.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Human–robot interaction through eye tracking for artistic drawing

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:12:22.110198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:e83e03904efc849ed23b2e6c97412c002c042d82b126dc3623aafbb9a3b1a431

Observation 5c486e0b-366d-4f44-8c52-95c382ebf73b · outbound

This paper cites Is it possible to recognize a speaker without listening? unraveling conversation dynamics in multi-party interactions using continuous eye gaze.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Is it possible to recognize a speaker without listening? unraveling conversation dynamics in multi-party interactions using continuous eye gaze

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.501692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:2a85542eb47e56f6c5df5dc4470502d7386a78292ed1c3bd811f97d54756b05d

Observation 8466c8fb-5ca1-4abc-a755-aedd77d4e226 · outbound

This paper cites Eye-gaze control of a wheelchair mounted 6dof assistive robot for activities of daily living.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Eye-gaze control of a wheelchair mounted 6dof assistive robot for activities of daily living

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:12:22.137422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:1184bc620bbb876e5da36f79d1108a2e78e3902e4e4996d1342c4de001b4c677

Observation 935e96a0-07e1-43b4-abba-59deb811e27b · outbound

This paper cites Free-view, 3d gaze-guided, assistive robotic system for activities of daily living.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Free-view, 3d gaze-guided, assistive robotic system for activities of daily living

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:12:22.132103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:8aee8c70fb53e34a8065af5bc4572da9fe749af3442637ee5b91bad2af5a839a

Observation 2ec08df2-47da-4ebf-b7ba-6311edd74354 · outbound

This paper cites Microsaccade-inspired event camera for robotics.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Microsaccade-inspired event camera for robotics

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:12:22.113054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:224bbe3de502370e73766bb8131d8adc4889be71676002c952805ca3b05f2301

Observation 84a328ec-92e3-407c-bad0-6116ac31cee4 · outbound

This paper cites Project Aria: A New Tool for Egocentric Multi-Modal AI Research.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Project Aria: A New Tool for Egocentric Multi-Modal AI Research

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-23T01:12:20.528544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:f1f7145d215c3430805ca4e63e9c2c4e47dc6f7d24ffe3b6fef304cb60384a6d

Observation 9599ee3f-bdf7-433d-b998-3f8e424e18fd · outbound

This paper cites Robust gaze- based intention prediction for real-world scenarios.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Robust gaze- based intention prediction for real-world scenarios

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.620189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:4138214397556f6cf93525368382ccfd76eff9c19f45c676a13b1e914a73f495

Observation ea247167-bc67-49e9-afed-9df33b676395 · outbound

This paper cites Getting to know your robot customers: Automated analysis of user identity and demographics for robots in the wild.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Getting to know your robot customers: Automated analysis of user identity and demographics for robots in the wild

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.617076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:a9f10b0a447cdcf629cdbdbfabc907b8933a07d5ce390a6875a8c0b510c952be

Observation 910f1d15-3185-4d1a-b07e-28c592d36a60 · outbound

This paper cites A personalized comfort space with variable shape based on environmental information for robot navigation in homes.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech A personalized comfort space with variable shape based on environmental information for robot navigation in homes

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.613531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:f7220ee2df691b28eb084830dde7ed7a4943939a1ebc45762d8965dde8dc4f41

Observation afe87ad1-6531-44e1-a26d-41153c03a66c · outbound

This paper cites Improving the collision tolerance of high-speed industrial robots via impact-aware path planning and series clutched actuation.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Improving the collision tolerance of high-speed industrial robots via impact-aware path planning and series clutched actuation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.610101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:cba7351eb6e2d0c63319ace91db865aef04ab595dec2fa50da8de87b93cef451

Observation a57f45cf-4fa9-494c-8990-9392a22f9248 · outbound

This paper cites In situ calibration of six- axis force–torque sensors for industrial robots with tilting base.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech In situ calibration of six- axis force–torque sensors for industrial robots with tilting base

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.606522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:93d632463c8ac5d7ad7c90244c537a22c1992748d54c8a8fdb57d2bb98cf9b63

Observation e243694a-600a-4025-af69-5088701a14c1 · outbound

This paper cites Gesture-informed robot assistance via foundation models.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Gesture-informed robot assistance via foundation models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.603049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:30ec14bbe724ca94eb8132d03f20fd9101a7cc2ad5b99a479cc8119bd48f4595

Observation 940510b7-d6c0-4d94-92e9-14a5af3548a0 · outbound

This paper cites Interactive multimodal robot dialog using pointing gesture recognition.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Interactive multimodal robot dialog using pointing gesture recognition

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.599715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:38df63d628439f98b9009a59647eab20604d22a6b8ad79e41e372fd8616fe0d3

Observation 496eb716-9ef3-4144-b255-ac71be72844e · outbound

This paper cites Code as policies: Language model programs for embodied control.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Code as policies: Language model programs for embodied control

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.596254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:90ec22d4bd2cadc3987090c964eafb2e0e398f10bbec2e974efdc7897149ccf7

Observation 0b0e1534-af26-4086-bd52-341690d9cfd4 · outbound

This paper cites Progprompt: Generating situated robot task plans using large language models.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Progprompt: Generating situated robot task plans using large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.592956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:7f04cce5cbeee8ebe1bdca647ff9970c7e0ace3de7145729e632ab627274b7fe

Observation de495b96-5a16-4344-bf7c-f97d58b016be · outbound

This paper cites Semi-autonomous robotic arm reaching with hybrid gaze–brain machine interface.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Semi-autonomous robotic arm reaching with hybrid gaze–brain machine interface

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.589241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:d0521b0b6cf28ed9f80582514c6151091f7e8668f328e19aa9e960f3275646cf

Observation 0c579db8-c1f5-44ad-abe7-36cdd32dcd11 · outbound

This paper cites Investigating the usability of collabo- rative robot control through hands-free operation using eye gaze and augmented reality.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Investigating the usability of collabo- rative robot control through hands-free operation using eye gaze and augmented reality

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.585779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:8399cdf5732af8e51f1d7dc984fb9b2621bcae0c9a3f9e5e116931ee244838eb

Observation c424de04-57ed-4f11-bde6-5bdfb2588bf8 · outbound

This paper cites Human gaze following for human-robot interaction.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Human gaze following for human-robot interaction

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.582187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:0690465c6e0a6d0ee483b9a7e945d8d139d5925be6746de348af3a27a7d82d86

Observation f52a08cc-f2d2-4f1a-9642-e0fb93061f79 · outbound

This paper cites Gaze-based attention recognition for human-robot collaboration.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Gaze-based attention recognition for human-robot collaboration

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.578972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:ee9d1485ead3b2ae040f212cc7fd98ee010bb94d32c53775ef70e46b7fb3c43a

Observation 5751a08f-385d-46d0-b29c-f6e50498ead3 · outbound

This paper cites A novel human-in-the-loop multimodal intention fusion method for human-robot interaction.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech A novel human-in-the-loop multimodal intention fusion method for human-robot interaction

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.575616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:8d5421e70896b15a2bdf14572172eea8afe2a18c717acf4c4e3748521b23910b

Observation 8578005a-437a-419a-96ab-d558537fcaee · outbound

This paper cites Alchemist: Llm-aided end-user development of robot applications.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Alchemist: Llm-aided end-user development of robot applications

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.572180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:0ad8a852f64b89b856369617b5b270231a22d216c588ce1b80b480ada72f9b85

Observation ecbd186f-2683-4c15-8ba5-4f0afafe9617 · outbound

This paper cites Lami: Large language models for multi-modal human-robot interaction.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Lami: Large language models for multi-modal human-robot interaction

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.568621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:82ac9a97baff474f1e2c74820a81d9d05eb0b6ef4a373044c9eee3c1db778975

Observation bb177b39-f85f-4513-92e2-f0658bc631a8 · outbound

This paper cites Robust speech recognition via large-scale weak supervi- sion.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Robust speech recognition via large-scale weak supervi- sion

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.565092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:ecf70775bad0233db8b4fd06e2ffcca44f37db90542602cb7686532ff35de5f8

Observation a72528e2-30a4-4431-aa57-a0b3192fc074 · outbound

This paper cites Grounding DINO: marrying DINO with grounded pre-training for open-set object detection.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Grounding DINO: marrying DINO with grounded pre-training for open-set object detection

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.561791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:ad4ce7b8fc230c35ffabcbb386117fa8bcb472fc6b5a7f373a801d4157a541ee

Observation ce1f1fe4-5005-4905-ad46-a85942036ea9 · outbound

This paper cites Sam 2: Segment anything in images and videos.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Sam 2: Segment anything in images and videos

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.558174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:a52f3968ad846d599f2541a445722dd615c67519f36e6dc7da3411cb853df89a

Observation 9fda874c-469f-4bb3-a6f7-507ac3e52dcc · outbound

This paper cites ORB-SLAM3: An accurate open-source library for visual, visual-inertial and multi-map SLAM.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech ORB-SLAM3: An accurate open-source library for visual, visual-inertial and multi-map SLAM

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.554943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:edbf9345e40ac0e172eb0478af1f7879fc08f4f590b65bf870fc86edd92c9507

Observation 2c5ad54f-e294-4151-b0f0-2c4d1dcff7a8 · outbound

This paper cites Super- glue: Learning feature matching with graph neural networks.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Super- glue: Learning feature matching with graph neural networks

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.551583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:f0e532a210b05982a0394624e33435152b2f5a05b9eec0c668991f36c48d554e

Observation 1f99c207-0fb2-4a07-bf49-3715be11343f · outbound

This paper cites Design and implementation of a haptic measurement glove to create realistic human-telerobot interactions.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Design and implementation of a haptic measurement glove to create realistic human-telerobot interactions

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.548125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:b8d4066d89a0c7263c62796948adfb68fa4191ad744992f1363c9e997fa855d9

Observation 7b540175-6910-4a4e-a4f2-d200020897b2 · outbound

This paper cites Don’t yell at your robot: Physical correction as the collaborative interface for language model powered robots.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Don’t yell at your robot: Physical correction as the collaborative interface for language model powered robots

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.544376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:321088222499e8301920b899cfc92d680ece945acfb16e2115813b4077bcb04d

Observation 178de768-70af-43cc-90b0-4aaf91676fdc · outbound

This paper cites Lora: Low-rank adaptation of large language models.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Lora: Low-rank adaptation of large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.540619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:5998ae83668f43fdf8991fd962bfe8b6758f8dbf5ce40fc02aa24b1e868c49be

Observation acb66482-e5a4-4b3c-9f17-d229533c0146 · outbound

This paper cites Model compression.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Model compression

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.537242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:a5bbb5686b2238642c54e2ef3d94cd8399c6920fe82a4ae02c36732985108aff

Observation 814818ae-8916-4fba-aca1-5f51523fa296 · outbound

This paper cites Retrieval- augmented generation for knowledge-intensive nlp tasks.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Retrieval- augmented generation for knowledge-intensive nlp tasks

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.534150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:4c5c05871b35187e05cb30320fad6a832aec728b9d15c6ae14387e731725acf2

Observation b1c3cf0a-1399-45fd-85ed-783303144749 · outbound

This paper cites Human gaze improves vision transformers by token masking.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Human gaze improves vision transformers by token masking

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.530595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:9a405106233b8390c39240ea3b79ab1a84924002cd3cf30d9745196038e7de47

Observation c56d6b22-0910-4d2f-940e-506fbbffe23a · outbound

This paper cites Egolife: Towards egocentric life assistant.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Egolife: Towards egocentric life assistant

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.527137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:f5d91d94a4fedf9ad580784f0e1783bdeb8671335d4fa89e78752c0e7fbe25aa

Observation a4db9305-d286-4410-9fa0-ff7150e5c81d · outbound

This paper cites Interactive multimodal robot dialog using pointing gesture recognition.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Interactive multimodal robot dialog using pointing gesture recognition

Reference 42

Resolution
verified exact
doi, observed 2026-05-23T01:12:19.807631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:b77bfd8340147ae5acfbf38f0d19974fa8357d723276737ecc4d75c85c06c64d

Observation 539233a1-3390-4a59-a9d1-24b70220ad98 · outbound

This paper cites The speech recognition error in our system primarily due to misinterpretation of similar-sounding words.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech The speech recognition error in our system primarily due to misinterpretation of similar-sounding words

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.523573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:4353dd3fa8e06767dbbc12f4dac247f5f6ce6007b280b602e69cb05067160e69

Observation 1bf8564e-aeb3-4089-85b9-e3865399e044 · outbound

This paper cites an unresolved cited work.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:15:18.519970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:2999628cd2cf878277adf85656244071b8368d347c1d61bede208fbf1f797111

Observation eff27f64-7fbf-4972-a976-62c64a22edbc · outbound

This paper cites Fea- ture matching using superglue becomes unreliable in the presence of weak object textures, repetitive patterns, or partial occlusions, leading to incorrect correspondences.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Fea- ture matching using superglue becomes unreliable in the presence of weak object textures, repetitive patterns, or partial occlusions, leading to incorrect correspondences

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.516832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:12fcf6aeb2168068631f056e5ed58a0a89d6631c66c4379d91f7dcb63a608c46

Observation afcba079-64b8-4e3e-ae7d-5b6cbcf93483 · outbound

This paper cites Since FAM-HRI requires the LLM’s response to strictly follow a prede- fined prompt format, any deviation renders the output unusable by subsequent system modules.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Since FAM-HRI requires the LLM’s response to strictly follow a prede- fined prompt format, any deviation renders the output unusable by subsequent system modules

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:15:18.513854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:f5b46eeb333469e2a28cc1cbe856a2259cd64bac96c7a540bd71845abc0cbe07

Observation ee2e4734-419e-463b-9803-3a8b807a24d2 · outbound

This paper cites an unresolved cited work.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:15:18.510143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:23466bc17545da5036b85fbb36313cca48120d33aca9f646dc839085db313c84

Observation cc3d4649-3672-4fd2-80a2-53745af9c675 · outbound

This paper cites an unresolved cited work.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:15:18.507120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:9336513bae7e54ba3b07fea032921455bc5308772c1c10ffa961845fc2745874

Observation 11f0fb02-9e1d-4821-9e3e-14599f2dd333 · outbound

This paper cites an unresolved cited work.

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:15:18.504367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:10:55.850274Z digest=sha256:a47c20b75c2cdf40690d9d0c6582bdb96ebceeb772905454fa01385d39105be9

Pith citing papers

Observation f9ad8420-686a-4da6-9528-a95aa196d81e · inbound

Mind Meets Space: Rethinking Agentic Spatial Intelligence from a Neuroscience-inspired Perspective cites this paper.

Mind Meets Space: Rethinking Agentic Spatial Intelligence from a Neuroscience-inspired Perspective FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-04T19:39:07.075356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:39:07.075356Z digest=sha256:50cf46d30482e2e885610af4470f7377304f56f9e71563c6a7b1f7209325abcf

Observation b9f2c4a1-f946-4ca1-afb4-b51790544ce6 · inbound

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving cites this paper.

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-22T12:21:31.264373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T12:17:54.325055Z digest=sha256:c0dc0b7c871c947a47b0420ff10e3659fee95c70eaf78701189686185ef2a7a3

Observation 0601019d-8154-4643-a96d-f50bc647a580 · inbound

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving cites this paper.

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T17:07:42.809307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:07:42.809307Z digest=sha256:bfabfee805b133e166cb179ed5ce33ca424608f56a9c4567870b65025e214131

Observation 8c624a78-df58-47cc-82e1-a232ae2d3f6a · inbound

Hierarchical Policies from Verbal and Egocentric Human Signals for Natural Human-Robot Interaction cites this paper.

Hierarchical Policies from Verbal and Egocentric Human Signals for Natural Human-Robot Interaction FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-03T04:57:38.606095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:31:02.640975Z digest=sha256:a117f0188b11c1d77f84d46790109ecde73280681f7fc435dfa16cb3ef5b69c9