Pith. sign in

REVIEW 1 cited by

Here's What I've Learned: Asking Questions that Reveal Reward Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2107.01995 v1 pith:FDWOZSOE submitted 2021-07-02 cs.HC cs.RO

classification cs.HCcs.RO
keywords questionshumanrobotrobotsinformativelearnedlearningreveal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Robots can learn from humans by asking questions. In these questions the robot demonstrates a few different behaviors and asks the human for their favorite. But how should robots choose which questions to ask? Today's robots optimize for informative questions that actively probe the human's preferences as efficiently as possible. But while informative questions make sense from the robot's perspective, human onlookers often find them arbitrary and misleading. In this paper we formalize active preference-based learning from the human's perspective. We hypothesize that -- from the human's point-of-view -- the robot's questions reveal what the robot has and has not learned. Our insight enables robots to use questions to make their learning process transparent to the human operator. We develop and test a model that robots can leverage to relate the questions they ask to the information these questions reveal. We then introduce a trade-off between informative and revealing questions that considers both human and robot perspectives: a robot that optimizes for this trade-off actively gathers information from the human while simultaneously keeping the human up to date with what it has learned. We evaluate our approach across simulations, online surveys, and in-person user studies. Videos of our user studies and results are available here: https://youtu.be/tC6y_jHN7Vw.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. What Do You Think I Think? Accounting for Human Beliefs Using Second-Order Theory of Mind

    cs.HC 2026-05 unverdicted novelty 5.0 of 10

    A second-order ToM agent using I-POMDP models human erroneous beliefs and cognitive biases to generate adaptive feedback that improves interaction informativeness.

Pith tools