Pith. sign in

Navigation Turing Test (NTT): Learning to Evaluate Human-Like Navigation

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

A key challenge on the path to developing agents that learn complex human-like behavior is the need to quickly and accurately quantify human-likeness. While human assessments of such behavior can be highly accurate, speed and scalability are limited. We address these limitations through a novel automated Navigation Turing Test (ANTT) that learns to predict human judgments of human-likeness. We demonstrate the effectiveness of our automated NTT on a navigation task in a complex 3D environment. We investigate six classification models to shed light on the types of architectures best suited to this task, and validate them against data collected through a human NTT. Our best models achieve high accuracy when distinguishing true human and agent behavior. At the same time, we show that predicting finer-grained human assessment of agents' progress towards human-like behavior remains unsolved. Our work takes an important step towards agents that more effectively learn complex human-like behavior.

fields

cs.LG 1

years

2024 1

verdicts

CONDITIONAL 1

representative citing papers

Effective Reward Specification in Deep Reinforcement Learning

cs.LG · 2024-12-10 · conditional · novelty 4.0

A thesis presenting four methods (ASAF, TeamReg, CoachReg, constrained RL, goal-conditioned GFlowNets) that improve reward specification for deep RL through demonstrations, policy regularization, behavior constraints, and multi-objective conditioning.

citing papers explorer

Showing 1 of 1 citing paper.

  • Effective Reward Specification in Deep Reinforcement Learning cs.LG · 2024-12-10 · conditional · none · ref 83 · internal anchor

    A thesis presenting four methods (ASAF, TeamReg, CoachReg, constrained RL, goal-conditioned GFlowNets) that improve reward specification for deep RL through demonstrations, policy regularization, behavior constraints, and multi-objective conditioning.