Pith. sign in

REVIEW 2 cited by

Comparing Apples to Oranges: LLM-powered Multimodal Intention Prediction in an Object Categorization Task

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.08424 v3 pith:A2NVGEUQ submitted 2024-04-12 cs.RO cs.AIcs.HC

classification cs.ROcs.AIcs.HC
keywords userintentioncuespredictiontaskcategorizationhumanllms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Human intention-based systems enable robots to perceive and interpret user actions to interact with humans and adapt to their behavior proactively. Therefore, intention prediction is pivotal in creating a natural interaction with social robots in human-designed environments. In this paper, we examine using Large Language Models (LLMs) to infer human intention in a collaborative object categorization task with a physical robot. We propose a novel multimodal approach that integrates user non-verbal cues, like hand gestures, body poses, and facial expressions, with environment states and user verbal cues to predict user intentions in a hierarchical architecture. Our evaluation of five LLMs shows the potential for reasoning about verbal and non-verbal user cues, leveraging their context-understanding and real-world knowledge to support intention prediction while collaborating on a task with a social robot. Video: https://youtu.be/tBJHfAuzohI

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Casper: Inferring Diverse Intents for Assistive Teleoperation with Vision Language Models

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A VLM-powered assistive teleoperation system infers diverse user intents from teleoperation snippets and executes them with a skill library, outperforming baselines on real-world mobile manipulation tasks.

  2. Gaze-supported Large Language Model Framework for Bi-directional Human-Robot Interaction

    cs.RO 2025-07 conditional novelty 4.0 of 10

    A gaze- and speech-driven LLM framework for assistive robots matches a scripted interaction pipeline on task performance while slightly increasing user-perceived confidence, at higher energy cost.

Pith tools