Pith. sign in

REVIEW 1 cited by

Improving Multimodal Interactive Agents with Reinforcement Learning from Human Feedback

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.11602 v1 pith:ANAIA3O7 submitted 2022-11-21 cs.LG cs.HCcs.MA

classification cs.LGcs.HCcs.MA
keywords agentshumanlearningfeedbackreinforcementrewardagentbehaviour
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

An important goal in artificial intelligence is to create agents that can both interact naturally with humans and learn from their feedback. Here we demonstrate how to use reinforcement learning from human feedback (RLHF) to improve upon simulated, embodied agents trained to a base level of competency with imitation learning. First, we collected data of humans interacting with agents in a simulated 3D world. We then asked annotators to record moments where they believed that agents either progressed toward or regressed from their human-instructed goal. Using this annotation data we leveraged a novel method - which we call "Inter-temporal Bradley-Terry" (IBT) modelling - to build a reward model that captures human judgments. Agents trained to optimise rewards delivered from IBT reward models improved with respect to all of our metrics, including subsequent human judgment during live interactions with agents. Altogether our results demonstrate how one can successfully leverage human judgments to improve agent behaviour, allowing us to use reinforcement learning in complex, embodied domains without programmatic reward functions. Videos of agent behaviour may be found at https://youtu.be/v_Z9F2_eKk4.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Conditional Multi-Stage Failure Recovery for Embodied Agents

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A conditional four-stage chain-prompting method for failure recovery improves success on the TEACH embodied-agent benchmark from 24.9% to 36.5% with the same plan and executor.

Pith tools